10 Commits

Author SHA1 Message Date
Teknium c8b1c049c7 fix(tui_gateway): size replay frames outside the lock; document the byte ceilings
transport.write serializes outside its lock so one large payload cannot block
other threads' frames; the replay ring is on the same hot path (write_json), so
its sizing json.dumps must follow the same rule instead of holding _replay_lock
for the whole encode. seq is not stamped at that point, which is a few bytes off
a MiB budget and irrelevant.

Also restore the original WHY for the count limit, explain the byte ceilings
(512 x 64 KiB tool results is ~32 MiB per session before the cap; 4 MiB per
session, 64 MiB per process) and make the module docstring state the bound in
bytes and the truncation watermark.
2026-09-09 12:20:38 -07:00
Xipong d31f0a647d fix: preserve streaming and replay safety after parent review 2026-09-09 12:20:38 -07:00
Xipong 721128ac02 fix: bound event replay payload bytes 2026-09-09 12:20:38 -07:00
Teknium 052a5dec66 refactor(tui_gateway): table-drive billing error kinds, share pet selection, compact group H docstrings 2026-09-02 23:25:59 -07:00
Teknium 2b538e535a refactor(tui_gateway): AST-neutral layout compaction of group H modules 2026-09-02 23:04:39 -07:00
Teknium 957147365a refactor(tui_gateway): split _apply_model_switch into phase helpers, compact controller/browser/relay/watcher/billing helpers 2026-09-02 22:47:13 -07:00
kshitijk4poor 9f05b06589 Revert "Merge pull request #94245 from kshitijk4poor/feat/gw-event-replay"
This reverts commit df7d7f6e8d, reversing
changes made to 1a66134404.
2026-08-27 11:26:57 +05:30
kshitijk4poor 874fab0ce0 feat(chat-plane): trace_id + turn telemetry, transient-delta split, seq-namespace epoch
Chat/event-plane quality work for the amended Phase 1 scope of #94484
(maintainer restructure: lean chat/event plane, no control-plane
changes). Three fixes came out of a source-level comparison against
OpenHands, Chainlit, VS Code, Zed, LangGraph, and Goose.

1. Per-turn trace_id + active-turn telemetry: _start_inflight_turn mints
   a 12-hex trace_id; _event_frame stamps it on every event frame in the
   turn, so a client can correlate the full lifecycle (dispatch -> first
   token -> tool calls -> complete) from one identifier — none of the six
   surveyed projects has frame-level turn correlation.
   session.events.stats now reports active_turns (session_id, trace_id,
   elapsed_s, streaming).

2. Transient vs durable events (OpenHands StreamingDeltaEvent pattern):
   message.delta / thinking.delta are stamped with seqs (live ordering
   holds) but never buffered — one streaming turn emitted hundreds of
   delta frames and evicted every durable control event from the
   512-slot ring, defeating replay for the exact reconnect window it
   exists to cover. The ring now evicts manually and records the highest
   DURABLE seq dropped, so truncated means real data loss and delta-only
   gaps no longer false-positive.

3. Seq-namespace epoch (Goose stale-cursor recovery): event_replay.EPOCH
   (8-hex per boot) is announced in gateway.ready and echoed by
   session.events.since; the client drops its seq watermarks when the
   epoch changes, so a stale HIGH watermark from a previous gateway
   process can no longer suppress replay/gap-detection forever. Legacy
   backends without an epoch are unaffected (client keeps watermarks).

Validation: Python 85/85 across replay/entry_ws/keepalive/protocol
(12 replay tests, 3 new); vitest 9/9 shared (2 new epoch tests), 66/66
desktop; full tui_gateway sweep 614/615 (1 known ordering flake, passes
in isolation). Live e2e on this tree: 12-event turn -> seqs contiguous
1..12, 8 durable frames buffered + 4 deltas live-only, single trace_id
on all frames, active_turns elapsed_s matches the real turn duration.

Research provenance: NousResearch/hermes-agent#94484 (comparative-scan
comment); techniques credited to OpenHands (transient split), Goose
(epoch/stale-cursor), per maintainer-restructured plan.
2026-08-26 02:55:04 +05:30
Teknium beb7941236 fix(tui-gateway): make WS reconnect replay actually deliver events (follow-up to #94219)
The #94219 replay was a production no-op: the server returned full
JSON-RPC envelopes from session.events.since while the client's replay
loop dispatches only elements with a top-level 'type' — every replayed
event was silently skipped. Each side's tests validated its own
assumption, so both suites stayed green.

- server: events_since() now returns bare event objects (the frame's
  params), the exact shape the live dispatch path consumes; ring stores
  params directly; cross-language contract test added on both sides.
- client: live frames racing an in-flight replay are parked and flushed
  seq-gated afterward — no double dispatch of deltas, no gap-skip from
  a watermark advanced past the replay window.
- restart poisoning: seq counters are in-process, so a backend restart
  reset them while clients kept high watermarks (replay forever empty,
  truncated=false). New replay_epoch advertised in gateway.ready and
  echoed by session.events.since; the client clears watermarks on epoch
  change.
- methods_session no longer reaches into event_replay privates
  (is_truncated() accessor).

Live repro: pre-fix, 3 stamped frames -> 0 dispatchable by the client
gate; post-fix 3/3. Tests: 16 py (replay+ws), 8 vitest, tsc clean, ruff
clean.
2026-08-24 20:11:30 -07:00
kshitijk4poor 87631bd8ae feat(tui-gateway): seq-stamped event replay for lossless desktop reconnect
Server: per-session monotonic seq on every routed event frame, bounded
512-frame replay ring (64 sessions, FIFO eviction), plus two new RPCs —
session.events.since (replay newer-than-watermark, reports latest_seq +
truncated so clients detect gaps) and session.events.stats (telemetry).

Client: per-session seq watermarks recorded from live frames; after any
successful reconnect a fire-and-forget fetchReplay() drains missed events
through the normal dispatch path (recordSeq ignores non-increasing seqs,
so stale replay can never regress a watermark); focus-triggered reconnect
nudge in use-gateway-boot for the Electron unfocused case where macOS wake
skips visibilitychange.

Replay failures are swallowed by design: lossless resume is an upgrade
over the previous lossy reconnect, never a new failure mode.
2026-08-25 02:31:19 +05:30