Files
EvoScientist/docs/superpowers/specs/2026-07-06-webui-sse-checkpoint-fallback-design.md
m4 e0acc6155e
Build / build (push) Has been cancelled
Docker / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
feat: improve WebUI run recovery
2026-07-10 17:35:44 +08:00

9.3 KiB
Raw Permalink Blame History

WebUI SSE Truncation Recovery via Checkpoint Fallback (α)

Date: 2026-07-06 Status: Design — pending implementation plan Scope: @evoscientist/webui front-end (separate npm repo) + one optional convenience route in this Python repo

1. Background & Root Cause

A WebUI user saw a 127 s research run render truncated output, ending mid-token at O-8282d7e. Investigation established:

  • The server-side run completed cleanly; the thread checkpoint in ~/.evoscientist/sessions.db holds the full 6818-char final answer (status: idle, error: None).
  • The disconnect is browser-side: langgraph dev sends an SSE heartbeat every 5 s, send_timeout is None, and the server supports resume via last-event-id (langgraph_api/api/runs.py:623). But /runs/stream is POST, so the browser reads it via fetch(); native EventSource auto-reconnect does not apply, and @evoscientist/webui does not re-POST to resume. When the connection drops mid-run, the front-end is left showing the last partial token and never fetches the completed answer.

Core insight: the server already has the complete answer in the checkpoint. The fix is to make the front-end fall back to that checkpoint when the stream ends abnormally. This is small, surgical, and directly addresses the symptom.

2. Goal

The user always sees the complete final answer in the WebUI, even if the SSE stream drops mid-run.

3. Non-Goals

  • Seamless stream resume (tracking last-event-id / seq and replaying buffered events). That is option β — better UX, ~3–5× the code, deferred.
  • The TUI pivot / launch-mode decoupling (option γ). Separable; only relevant if the port-conflict / orphan-process issues need addressing.
  • Patching langgraph_api upstream.

4. Detection Logic — What Counts as "Truncated"

The langgraph SSE protocol emits terminal events when a run finishes (success → end, failure → error; exact names to be confirmed against the @langchain/langgraph-sdk version during implementation). The front-end's streaming reader loop currently processes events until the fetch() ReadableStream closes.

Truncation = the stream closed (reader returned) without a terminal event having been received. Causes: browser tab throttled, network blip, proxy idle timeout. On truncation, the in-progress assistant bubble is left showing a prefix of the real answer.

5. Components

F1 — Truncation detector (front-end)

In the streaming reader loop, track a terminalSeen flag. Set it when a terminal event arrives. When the reader returns, if !terminalSeen, mark the run truncated and trigger F2.

File: @evoscientist/webui — the run-stream consumer (the module that wraps the langgraph-ts SDK's runs.stream and renders into the message list).

F2 — Checkpoint fallback fetch (front-end)

On truncated, fetch the complete final assistant message and replace the in-progress bubble's content with it (not append — the partial tokens are a prefix of the same message).

Two implementation choices, pick one:

  • F2a (zero server change): call the standard langgraph endpoint GET /threads/{thread_id}/state, walk values.messages backwards to the last AIMessage, join its content blocks. No new route, but the front-end parses raw state and reimplements "find latest assistant text" logic.
  • F2b (recommended, uses S1 below): call GET /api/threads/{id}/final-answer → {content, completed_at, complete}. Server owns the parsing; front-end stays dumb.

Render a subtle affordance (e.g., a dim "⚠ stream dropped — recovered from checkpoint" line above the message) so the user knows a disconnect happened and the displayed text is authoritative, not stale streaming.

F3 — Retry until the run settles (front-end)

The browser may drop while the run is still executing server-side. A single fallback fetch at drop-time could return a not-yet-final message. So F2 retries with backoff until either:

  • complete == true (run finished — render and stop), or
  • a max wait elapses (default 240 s — comfortably exceeds the observed 127 s run even if the drop happens at t=0), then render whatever is latest and stop retrying.

Backoff: poll at 1 s, 2 s, 4 s, then every 5 s up to the max. Each poll replaces the bubble with the latest content, so if the run finishes mid-retry the user sees it update to the final form.

Enables F2b. Mounted on the existing custom ASGI app already wired via langgraph.json → EvoScientist.langgraph_dev.http:app (currently only serves /api/models).

Contract:

GET /api/threads/{thread_id}/final-answer
200 → { "content": str,            # last AIMessage text, blocks joined
        "completed_at": iso8601,    # run end timestamp, or null
        "complete": bool }          # true iff run reached terminal state
404 → thread not found

complete derivation (run reached a terminal state): thread.status == "idle" or latest checkpoint's next tuple is empty.

File: EvoScientist/langgraph_dev/http.py — add get_final_answer route beside the existing get_models. Implementation reads the thread state via the shared checkpointer (AsyncSqliteSaver at ~/.evoscientist/sessions.db, already used by sessions.py). If direct checkpointer access is awkward from inside the mounted app, fall back to an internal langgraph_sdk.get_client() call to GET /threads/{id}/state on localhost — the contract stays identical.

Why prefer F2b+S1 over F2a: the "find latest AIMessage + join content blocks

  • decide complete?" logic is non-trivial (content is a list of typed blocks, including reasoning blocks that should be excluded from the rendered answer — see the checkpoint: block[0] was reasoning, block[1] was text). Centralizing it server-side keeps the front-end thin and makes the same logic reusable by the TUI's future recovery path if γ is ever pursued.

6. Data Flow

browser fetch(/runs/stream, POST) ──SSE──▶ langgraph dev
        │                                        │
        │  (mid-run, connection drops)           │ run continues to completion
        ▼                                        │
 reader returns, no terminal event ──▶ truncated │
        │                                        │
        │  GET /api/threads/{id}/final-answer    │
        ▼                                        ▼
 http.py::get_final_answer ──read checkpoint──▶ sessions.db
        │
        ▼
 {content, complete} ──▶ if !complete, retry (F3); else render (F2)

7. Error Handling

Case Behavior
Stream ends normally (end/error received) F1 sets terminalSeen; no fallback; existing render path unchanged
Stream drops, run already finished server-side First fallback fetch returns complete=true; render immediately
Stream drops, run still executing F3 retries with backoff; each retry renders latest; stops when complete or 120 s max
Thread unknown / deleted 404; front-end shows "stream interrupted and the session could not be recovered"
Convenience route unreachable (older server without S1) Front-end falls back to F2a (raw state endpoint) — degrade gracefully

8. Testing

Front-end (@evoscientist/webui):

  • Unit: feed the reader a synthetic event stream that closes without a terminal event → assert truncated == true; feed one with end → false.
  • Integration: mock fetch to drop mid-stream; assert the fallback fetch fires and the bubble's content is replaced with the checkpoint value.

Server (this repo, if S1): new tests/test_http_final_answer.py

  • Completed thread → 200, complete=true, content matches the last AIMessage text (not the reasoning block).
  • Mid-run thread (forced busy / non-empty next) → 200, complete=false.
  • Unknown thread → 404.
  • Content with mixed reasoning + text blocks → only text in content.

9. Cross-Repo Coordination & Rollout

  • Order: land S1 in this Python repo first (route + tests, behind no flag — it's a pure addition). Then F1/F2/F3 in @evoscientist/webui.
  • Versioning: the WebUI launches via npx @evoscientist/webui@latest, so the front-end fix reaches users on their next launch automatically once published. No coordinated upgrade required on the Python side.
  • Backward compatibility: if a user runs a new front-end against an older Python server without S1, the front-end must detect the 404 and fall back to F2a (raw state endpoint). If an old front-end runs against a new server, S1 is simply unused — no harm.

10. Open Questions / Future Work

  • Exact terminal-event names for the SDK version in use (end/error vs. done/result). Confirm against @langchain/langgraph-sdk during implementation; the detector is parameterized on this.
  • True resume (β): track last-event-id (or the V2 thread-stream seq) and replay buffered events on reconnect — seamless, no visible "recovered" flash. Larger front-end effort; defer until α is validated.
  • TUI pivot (γ): separable work to make the in-process TUI the default and eliminate the port-conflict / orphan-process issues. Independent spec if pursued.