9.3 KiB
WebUI SSE Truncation Recovery via Checkpoint Fallback (α)
Date: 2026-07-06
Status: Design — pending implementation plan
Scope: @evoscientist/webui front-end (separate npm repo) + one optional convenience route in this Python repo
1. Background & Root Cause
A WebUI user saw a 127 s research run render truncated output, ending mid-token at
O-8282d7e. Investigation established:
- The server-side run completed cleanly; the thread checkpoint in
~/.evoscientist/sessions.dbholds the full 6818-char final answer (status: idle,error: None). - The disconnect is browser-side:
langgraph devsends an SSE heartbeat every 5 s,send_timeoutisNone, and the server supports resume vialast-event-id(langgraph_api/api/runs.py:623). But/runs/streamis POST, so the browser reads it viafetch(); nativeEventSourceauto-reconnect does not apply, and@evoscientist/webuidoes not re-POST to resume. When the connection drops mid-run, the front-end is left showing the last partial token and never fetches the completed answer.
Core insight: the server already has the complete answer in the checkpoint. The fix is to make the front-end fall back to that checkpoint when the stream ends abnormally. This is small, surgical, and directly addresses the symptom.
2. Goal
The user always sees the complete final answer in the WebUI, even if the SSE stream drops mid-run.
3. Non-Goals
- Seamless stream resume (tracking
last-event-id/seqand replaying buffered events). That is option β — better UX, ~3–5× the code, deferred. - The TUI pivot / launch-mode decoupling (option γ). Separable; only relevant if the port-conflict / orphan-process issues need addressing.
- Patching
langgraph_apiupstream.
4. Detection Logic — What Counts as "Truncated"
The langgraph SSE protocol emits terminal events when a run finishes
(success → end, failure → error; exact names to be confirmed against the
@langchain/langgraph-sdk version during implementation). The front-end's
streaming reader loop currently processes events until the fetch() ReadableStream
closes.
Truncation = the stream closed (reader returned) without a terminal event having been received. Causes: browser tab throttled, network blip, proxy idle timeout. On truncation, the in-progress assistant bubble is left showing a prefix of the real answer.
5. Components
F1 — Truncation detector (front-end)
In the streaming reader loop, track a terminalSeen flag. Set it when a
terminal event arrives. When the reader returns, if !terminalSeen, mark the
run truncated and trigger F2.
File: @evoscientist/webui — the run-stream consumer (the module that wraps
the langgraph-ts SDK's runs.stream and renders into the message list).
F2 — Checkpoint fallback fetch (front-end)
On truncated, fetch the complete final assistant message and replace the
in-progress bubble's content with it (not append — the partial tokens are a
prefix of the same message).
Two implementation choices, pick one:
- F2a (zero server change): call the standard langgraph endpoint
GET /threads/{thread_id}/state, walkvalues.messagesbackwards to the lastAIMessage, join its content blocks. No new route, but the front-end parses raw state and reimplements "find latest assistant text" logic. - F2b (recommended, uses S1 below): call
GET /api/threads/{id}/final-answer→{content, completed_at, complete}. Server owns the parsing; front-end stays dumb.
Render a subtle affordance (e.g., a dim "⚠ stream dropped — recovered from checkpoint" line above the message) so the user knows a disconnect happened and the displayed text is authoritative, not stale streaming.
F3 — Retry until the run settles (front-end)
The browser may drop while the run is still executing server-side. A single fallback fetch at drop-time could return a not-yet-final message. So F2 retries with backoff until either:
complete == true(run finished — render and stop), or- a max wait elapses (default 240 s — comfortably exceeds the observed 127 s run even if the drop happens at t=0), then render whatever is latest and stop retrying.
Backoff: poll at 1 s, 2 s, 4 s, then every 5 s up to the max. Each poll replaces the bubble with the latest content, so if the run finishes mid-retry the user sees it update to the final form.
S1 — Convenience route GET /api/threads/{id}/final-answer (this repo, optional but recommended)
Enables F2b. Mounted on the existing custom ASGI app already wired via
langgraph.json → EvoScientist.langgraph_dev.http:app (currently only serves
/api/models).
Contract:
GET /api/threads/{thread_id}/final-answer
200 → { "content": str, # last AIMessage text, blocks joined
"completed_at": iso8601, # run end timestamp, or null
"complete": bool } # true iff run reached terminal state
404 → thread not found
complete derivation (run reached a terminal state):
thread.status == "idle" or latest checkpoint's next tuple is empty.
File: EvoScientist/langgraph_dev/http.py — add get_final_answer route
beside the existing get_models. Implementation reads the thread state via the
shared checkpointer (AsyncSqliteSaver at ~/.evoscientist/sessions.db,
already used by sessions.py). If direct checkpointer access is awkward from
inside the mounted app, fall back to an internal langgraph_sdk.get_client()
call to GET /threads/{id}/state on localhost — the contract stays identical.
Why prefer F2b+S1 over F2a: the "find latest AIMessage + join content blocks
- decide complete?" logic is non-trivial (content is a list of typed blocks,
including
reasoningblocks that should be excluded from the rendered answer — see the checkpoint: block[0] wasreasoning, block[1] wastext). Centralizing it server-side keeps the front-end thin and makes the same logic reusable by the TUI's future recovery path if γ is ever pursued.
6. Data Flow
browser fetch(/runs/stream, POST) ──SSE──▶ langgraph dev
│ │
│ (mid-run, connection drops) │ run continues to completion
▼ │
reader returns, no terminal event ──▶ truncated │
│ │
│ GET /api/threads/{id}/final-answer │
▼ ▼
http.py::get_final_answer ──read checkpoint──▶ sessions.db
│
▼
{content, complete} ──▶ if !complete, retry (F3); else render (F2)
7. Error Handling
| Case | Behavior |
|---|---|
Stream ends normally (end/error received) |
F1 sets terminalSeen; no fallback; existing render path unchanged |
| Stream drops, run already finished server-side | First fallback fetch returns complete=true; render immediately |
| Stream drops, run still executing | F3 retries with backoff; each retry renders latest; stops when complete or 120 s max |
| Thread unknown / deleted | 404; front-end shows "stream interrupted and the session could not be recovered" |
| Convenience route unreachable (older server without S1) | Front-end falls back to F2a (raw state endpoint) — degrade gracefully |
8. Testing
Front-end (@evoscientist/webui):
- Unit: feed the reader a synthetic event stream that closes without a terminal
event → assert
truncated == true; feed one withend→false. - Integration: mock
fetchto drop mid-stream; assert the fallback fetch fires and the bubble's content is replaced with the checkpoint value.
Server (this repo, if S1): new tests/test_http_final_answer.py
- Completed thread → 200,
complete=true, content matches the last AIMessage text (not the reasoning block). - Mid-run thread (forced
busy/ non-emptynext) → 200,complete=false. - Unknown thread → 404.
- Content with mixed
reasoning+textblocks → onlytextincontent.
9. Cross-Repo Coordination & Rollout
- Order: land S1 in this Python repo first (route + tests, behind no flag —
it's a pure addition). Then F1/F2/F3 in
@evoscientist/webui. - Versioning: the WebUI launches via
npx @evoscientist/webui@latest, so the front-end fix reaches users on their next launch automatically once published. No coordinated upgrade required on the Python side. - Backward compatibility: if a user runs a new front-end against an older Python server without S1, the front-end must detect the 404 and fall back to F2a (raw state endpoint). If an old front-end runs against a new server, S1 is simply unused — no harm.
10. Open Questions / Future Work
- Exact terminal-event names for the SDK version in use (
end/errorvs.done/result). Confirm against@langchain/langgraph-sdkduring implementation; the detector is parameterized on this. - True resume (β): track
last-event-id(or the V2 thread-streamseq) and replay buffered events on reconnect — seamless, no visible "recovered" flash. Larger front-end effort; defer until α is validated. - TUI pivot (γ): separable work to make the in-process TUI the default and eliminate the port-conflict / orphan-process issues. Independent spec if pursued.