Commit Graph

83 Commits

Author SHA1 Message Date
Teknium 42e34fa73d refactor(tui_gateway): methods_session — collapse _find_live_unpersisted/_try_get_session, squeeze intra-body blanks, rewrap comments to 100 cols 2026-09-02 23:51:54 -07:00
Teknium 955d362176 refactor(tui_gateway): methods_session — fold steer/redirect into _apply_correction, _save_via_compute_host, _int_param/_b64 helpers, compact payload literals and docstrings 2026-09-02 23:40:05 -07:00
Teknium 8c62bc058f refactor(tui_gateway): methods_session — _listing_rows helper, _Resume.claim/child_history, compact create/resume payloads 2026-09-02 22:53:19 -07:00
Teknium 99433742dc refactor(tui): extract the prompt turn into prompt_turn.py; compact methods_*/session_*/agent_callbacks; split _start_agent_build into scope/wiring helpers
- server.py 7663 -> 5319: _run_prompt_submit and its goal/loop/voice/scope
  phases move to tui_gateway/prompt_turn.py (bound via method_ctx.bind_module);
  _start_agent_build split into _bind/_release_build_profile_scopes,
  _deferred_build_agent_kwargs, _wire_session_agent, _start_session_services;
  _load_enabled_toolsets split (_enabled_mcp_server_names, _resolve_explicit_toolsets).
- methods_slash: _LIVE_SLASH_OUTPUT dispatch table; methods_tools: _SLASH_BUILTINS,
  _guarded; tool_progress: _PROGRESS_HANDLERS; methods_config_set:
  _REASONING_DISPLAY_WORDS; methods_voice: _VOICE_TOGGLE_ACTIONS.
- Unified: _denied_source (methods_session) replaces _WORKER_SOURCES in
  methods_profiles; _compress_live_with_feedback / _compute_host_slash shared by
  the slash mirror and /compress; _end_voice_chat shared by stop-phrase paths;
  _watcher_mtime_ns; _reaper_session_is_detached_idle; _notif_* helpers.
- Dead: _profile_dir_or_err, _resume_info, _slash_builtin_table (refs.py: 0 hits).
- acp_adapter/server.py: docstring compaction only (AST-identical).
- Every file semantically reviewed hunk-by-hunk for wire/log/lock/order parity.
2026-09-02 14:09:13 -07:00
Teknium 28bc8d05b9 refactor(tui): restore compact WHY notes lost in earlier docstring compaction 2026-09-02 14:09:08 -07:00
Teknium a0a41abf72 refactor(tui): move slash.exec mirror + completion helpers out of server.py 2026-09-02 14:09:06 -07:00
Teknium beacf4e942 refactor(tui): resume — verified partial work (acp_adapter + methods_* compaction, bind_module via globals()) 2026-09-02 14:09:01 -07:00
aeonsong 8076c78c87 fix(desktop): scope handoff config to session profile 2026-09-02 06:17:47 -07:00
Teknium 54b2c0ea38 fix(tui_gateway): scope live-session reuse to the requesting profile
`_find_live_session_by_key` matched live runtimes by bare stored session id.
Stored ids are timestamp-based and can exist in more than one profile's
store, so `session.resume` for profile B (fast path, post-build re-check, and
`_claim_or_reuse_live`) could hand back profile A's live runtime — the turn
then ran with A's persona/tools and wrote A's memory (#100029).

Give the lookup an optional `profile_home` (default: any profile, unchanged
for callers that have no profile to scope by) using the same string compare
`_find_live_unpersisted` already uses, and pass the resolved home at every
resume/claim site. `_claim_parked_runtimes` gets the same scope so a resume
under B never finalizes A's parked runtime of the same id.

Reimplements the profile-scope half of #100213 by @Finn763; the Group-title
capability-sync change from that PR is intentionally not carried.

Co-authored-by: Finn763 <165816600+finn763@users.noreply.github.com>
2026-09-02 05:40:49 -07:00
Teknium aa80626764 fix(tui-gateway): adopt late compute-host compress acks instead of a false 120s timeout (#97948)
Manual /compress on a compute-host (turn_isolation) session blocked its RPC
waiter for a hard-coded 120s, answered error 5019, and then DROPPED the
host's late `control.ack`: HostSupervisor.control() popped the pending
queue in `finally`, so `_handle_host_frame` had nothing to deliver to. The
host kept compressing, succeeded minutes later, rotated the session — and
the gateway session never mirrored the new session_key/history_version and
the desktop never refreshed its transcript.

- host_supervisor: `control(..., on_late_ack=)` leaves a one-shot handler
  registered when the waiter times out; control.ack/control.error/error
  frames for that request_id fire it (bounded: 30min TTL, cap 64). A host
  crash fails outstanding handlers with a synthetic control.error.
- server: `_compute_host_compress_wait_seconds()` derives the wait from
  `compression.context_total_ceiling_seconds` (+30s slack, floor 120s,
  cap 630s) instead of the literal 120. `_adopt_late_compute_host_compress_ack`
  applies the metadata mirror and emits the same `session.info` a normal
  compress does plus the existing `status.update kind=compacted` edge; a
  late error goes out through the existing `error` event.
- session.compress / slash.compress (methods_tools + _mirror_slash_side_effects):
  on waiter timeout answer `status: pending` (not 5019) and register the
  late-ack handler.
- desktop: SESSION_COMPRESS_TIMEOUT_MS 120s -> 660s (above the gateway cap);
  `status: 'pending'` renders as an info notice, not `error:`; the
  `compacted` status edge rehydrates an idle active session's transcript
  (mid-turn compaction still defers to the turn settle path).

Minimal extraction of the design in #99630 by @vsd2807 (design trace by
@andrexibiza and @JoaoMarcos44 in the #97948 thread); no new DB tables,
modules, or polling protocol.

Refs #97948

Co-authored-by: VVV <vaibhavdahiya28@gmail.com>
2026-09-02 01:35:59 -07:00
Teknium 52f359a011 fix(sessions): Desktop resume of a heavily-compacted chat no longer fails with 4130
Desktop's cold resume (defer_history + omit_messages, transcript paged over
REST) only ever holds the live tip segment in memory, but session.resume
bounded it against the FULL compression lineage (sessions.max_resume_messages,
default 20000). A Bot Chat with 85 compaction segments / ~29k lineage rows
behind a ~700-row tip was refused at 20001, sent zero model prompts, and sat on
"Waking up default…" forever — the healthiest possible session shape, rejected
by a guard sized for in-memory materialization.

- hermes_state: one `_resume_lineage_ids` definition shared by the resume
  readers (get_resume_conversations, get_ancestor_display_prefix) and the
  guard (assert_resume_safe / get_resume_message_count). Guard grows
  `tip_only=` and names the scope it counted; the branch-aware lineage the
  readers already used is now what the guard counts too (a /branch copy was
  being counted against its parent's rows).
- tui_gateway session.resume: deferred, omit_messages and lazy resumes are
  bounded by the tip; only the full in-memory lineage resume keeps the
  lineage-wide bound. Deferred hydration falls back to tip-only history when
  the lineage exceeds the limit instead of loading the rows the guard refused.
- CLI mid-setup tip-only path routes through the same guard instead of
  borrowing assert_export_safe.
- docs: sessions.max_resume_messages / max_export_messages documented with the
  per-surface scope.

Live repro (real SessionDB fixture, 85 segments / 29,226 lineage rows / 666 tip
rows, real tui_gateway.server.handle_request): before — deferred resume ->
4130; after — ok, hydrated history=666 prefix=0; the non-deferred full resume
still returns 4130 on the same fixture.
2026-09-01 22:14:33 -07:00
ClintonEmok 78eb4ebbfc fix(gateway): compensate half-written seeded branches + observable failures
Review follow-up on #93959:

1. Partial-failure window: if the row commits but the transcript copy or
   title write fails, the durable-but-empty child defeated the lazy
   first-prompt fallback (_ensure_session_db_row is INSERT OR IGNORE), so
   the renderer fail-latched on a transcript-less session again. The seed
   block now compensates: delete just this child so the lazy path can
   retry cleanly. Disk-full is exempt — deleting data on a full disk makes
   things worse.

2. Silent degradation: the best-effort catch now logs at WARNING with
   exc_info instead of DEBUG, so a regression in this user-facing path is
   observable without enabling debug logs.

Tests: compensation deletes the half-written row and preserves
pending_title; disk-full keeps the row and surfaces the WARNING.
2026-09-01 22:58:48 -05:00
ClintonEmok 6556439765 fix(gateway): persist seeded branch children at session.create (#93959)
Desktop branch creation hung on an infinite spinner and lost the branch
on restart. Root cause: the renderer branches via session.create with
parent_session_id + a seeded transcript, but session.create defers the
DB row to the first prompt (the draft-hygiene contract). The renderer's
post-create resume then re-fetches the fresh child through REST and
defer_history hydration — both read the DB. An unpersisted child 404s
and hydrates empty, the client fail-latch (sessionShouldHaveTranscript +
empty messages) refuses to bind a "transcript-less" session, and the
user sees a spinner forever; on restart the rowless child vanishes and
the optimistic "Draft: Branch N" entry disappears with it.

A seeded branch is explicit user intent, not an abandoned draft.
session.create now persists the child immediately when both
parent_session_id AND seeded history are present:

- Row created in the PARENT's profile-scoped state.db, stamped with
  _branched_from + parent_session_id (same shape as TUI /branch).
- Seeded transcript copied via append_messages_batch so REST prefetch
  and defer_history hydration find it on the first read.
- Title assigned from get_next_title_in_lineage(parent) and cleared
  from pending_title — the branch lands in the parent's lineage instead
  of falling back to a message-preview name.

Persistence is best-effort: a broken DB logs and lets create succeed,
leaving the lazy first-prompt path as fallback. Plain drafts keep the
lazy-row contract unchanged.

Fixes #93959
2026-09-01 22:58:48 -05:00
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
Brooklyn Nicholson 5c3bfb6da8 fix(tui_gateway): retire recovery marker on local stop
A confirmed Desktop Stop left the crash-recovery marker on disk until the
run thread finished. If the backend exited in that window, resume treated
the leftover as a crash and auto-continued the turn the user had stopped.

Co-authored-by: Jaime Chieng <164842890+buddhaholic420@users.noreply.github.com>
2026-08-31 20:50:39 -05:00
Brooklyn Nicholson adb23c13cb fix(bot-mode): keep Bot Chat resume on a proven compression tip
Desktop opens the registry id, then session.resume walked the legacy
unmarked-child fallback, so Open Chat still landed in a side chat after
the title lookup was already strict. Recoverable-archive resurrection
uses the same helper.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
2026-08-31 20:47:53 -05:00
Gille bb28056efd fix(bot-mode): keep canonical lookup on compression lineage 2026-08-31 20:47:53 -05:00
fangliquanflq da090aa4ba fix(prompt): preserve resumed workspace provenance 2026-08-31 10:10:25 -07:00
fangliquanflq c6ee4e0809 fix(prompt): skip bundled AGENTS.md for desktop launch cwd 2026-08-31 10:10:25 -07:00
Teknium 6874b99d49 fix(sessions): stamp launch-profile name on new session rows instead of NULL
Sessions created on the launch/default profile were persisted with
profile_name = NULL by all three writers (run_agent._ensure_db_session
None'd out 'default'; the desktop backend's _ensure_session_db_row and
session.branch passed None when no profile_home override was set).

NULL used to mean 'launch profile' by convention, but the desktop now
keys sessions by (profile, id), filters the sidebar by profile scope,
and resolves @session:<profile>/<id> deep links by profile match — a
NULL row matches nothing, so sessions created around a profile switch
vanished from the sidebar and their deep links could not be opened
(#99222). The #94724 one-shot legacy-owner backfill stamps literal
'default' onto old NULL rows, so writers minting NEW NULL rows after
that backfill ran recreated the exact state it exists to repair.

Stamp the real profile name at creation time in all three writers.
E2E-verified against a temp HERMES_HOME: both the desktop create path
and the agent path now persist profile_name='default'.

Fixes #99222
2026-08-31 05:56:18 -07:00
David Dudok de Wit 93c7089f70 feat(bot-mode): run same-gateway Group Chats without Desktop 2026-08-30 22:19:06 -07:00
Teknium c0875ba503 fix(tui): derive resume todo snapshots from already-loaded history
The eager _read_persisted_todo_state(db, target) added a second
get_messages_as_conversation call on every resume, breaking the
one-lineage-SELECT contract pinned by
test_session_resume_uses_parent_lineage_for_display. Derive the
snapshot from the history each resume path already loaded instead;
deferred (defer_history) resumes cache it in the hydration worker once
the transcript arrives.
2026-08-29 18:40:51 -07:00
itsflownium 393af4a310 fix(todo): live task state via revisioned snapshots and a dedicated todo.updated event
Salvaged from PR #97815 by @itsflownium, slimmed to the schema-free core:
- TodoStore gains a monotonic in-memory revision; the todo tool result
  returns it so clients can reject stale updates
- tui_gateway emits a dedicated todo.updated full-snapshot event that
  bypasses optional tool-progress display settings
- session resume/activate responses attach the authoritative todo
  snapshot; renderer restores it with revision arbitration
- desktop store tracks per-session revisions and rejects regressions

The session_todo_state DB table from the original PR is intentionally
dropped: canonical todo tool results already persist in conversation
history, so resume paths derive the snapshot from the stored transcript
instead of a parallel store.
2026-08-29 18:40:51 -07:00
Teknium e05c91ac71 fix(cli): slow /handoff transfers no longer misreported as "gateway not running"
Live-reproduced on main: /handoff poll-waited a flat 60s for a TERMINAL
state, but the gateway's dispatch is a full synthetic agent turn (whole
transcript replay + delivery) that routinely exceeds 60s on long sessions.
The CLI then printed "Timed out waiting for the gateway. Is `hermes
gateway` running?" (false diagnosis), called fail_handoff() on the RUNNING
row (stomping the gateway's claim), and promised "Your CLI session is
intact" after switch_session had already re-pointed the session. The
watcher later overwrote failed -> completed: split-brain.

- hermes_state.fail_handoff gains only_states CAS; waiters can only fail
  rows still pending. Owner (gateway watcher) keeps the unconditional form.
- CLI wait loop is two-phase: 60s for the CLAIM (pending) — a timeout
  there really does mean no gateway — then up to 15 min for the claimed
  dispatch with 30s heartbeats; a running row is never failed by the CLI.
- Desktop handoff.fail RPC now CAS-fails pending rows only; a running row
  returns {failed: false, state: running} instead of stomping the claim.

Repro (real _handoff_watcher, real state.db, CLI as separate process,
75s dispatch): before — CLI timeout @60s + false message + row stomped;
after — pending->running@5s->completed@80s, clean CLI exit.
2026-08-29 18:08:55 -07:00
David Tyler 84e17db0bd fix(bot-mode): canonical bot DMs always follow the profile's current config
Bot-Mode canonical chats (the ONE forever DM per bot) and room plumbing
sessions are plugin-owned scratch conversations. They are now created with
an explicit follow_profile_config contract, persisted in the session row's
model_config, so session.resume rebuilds from the member profile's CURRENT
config instead of restoring the stored model/provider pin from an old row.

That stale pin is what left bot DMs stuck on a dead provider (e.g. 'out of
Nous credits' after the profile was switched to ollama-cloud) while the
same bot worked fine in rooms — the mirror image of the room-plumbing bug
(#89497 class). Normal 1:1 user chats keep the stored-runtime restore:
opening an older chat must show the model it actually used.

- tui_gateway/methods_session.py: accept follow_profile_config on session.create
- tui_gateway/server.py: persist the marker in the row; skip stored-runtime
  overrides on resume when present
- apps/desktop/src/plugins/hermes-bots/plugin.js: send the contract from
  createCanonicalChat and ensureGroupChatSession
- tests: backend override + row-persist coverage; desktop source-contract
  coverage for both session kinds
2026-08-28 02:59:17 -07:00
David Tyler 316e51ae72 fix(bot-mode): room plumbing sessions always follow the profile's current config
Room member sessions in Bot Mode are per-member scratch conversations
inside a group chat. session.resume restored their stored model/provider
pin from the row's model_config, so a room bot stayed stuck on whatever
provider was pinned when the row was first written — even after the
profile was switched. Every room message then failed on the stale
provider (e.g. 'out of Nous credits' after switching a profile from
Nous to ollama-cloud) while the same bot worked fine in DMs.

Add an explicit room_plumbing contract:
- session.create accepts room_plumbing: true, persisted in model_config
- _stored_session_runtime_overrides() returns {} for marked rows, so a
  room session always rebuilds from the member profile's CURRENT config
- hidden + 'Group:' title shape is kept as a legacy fallback for rows
  created by older desktop builds that never sent the marker; hidden
  non-room chats keep the stored-runtime restore
- Desktop Bot Mode sends room_plumbing: true when creating the hidden
  per-member room sessions

Fixes #89497
2026-08-28 02:59:17 -07:00
kshitijk4poor 9f05b06589 Revert "Merge pull request #94245 from kshitijk4poor/feat/gw-event-replay"
This reverts commit df7d7f6e8d, reversing
changes made to 1a66134404.
2026-08-27 11:26:57 +05:30
kshitijk4poor 874fab0ce0 feat(chat-plane): trace_id + turn telemetry, transient-delta split, seq-namespace epoch
Chat/event-plane quality work for the amended Phase 1 scope of #94484
(maintainer restructure: lean chat/event plane, no control-plane
changes). Three fixes came out of a source-level comparison against
OpenHands, Chainlit, VS Code, Zed, LangGraph, and Goose.

1. Per-turn trace_id + active-turn telemetry: _start_inflight_turn mints
   a 12-hex trace_id; _event_frame stamps it on every event frame in the
   turn, so a client can correlate the full lifecycle (dispatch -> first
   token -> tool calls -> complete) from one identifier — none of the six
   surveyed projects has frame-level turn correlation.
   session.events.stats now reports active_turns (session_id, trace_id,
   elapsed_s, streaming).

2. Transient vs durable events (OpenHands StreamingDeltaEvent pattern):
   message.delta / thinking.delta are stamped with seqs (live ordering
   holds) but never buffered — one streaming turn emitted hundreds of
   delta frames and evicted every durable control event from the
   512-slot ring, defeating replay for the exact reconnect window it
   exists to cover. The ring now evicts manually and records the highest
   DURABLE seq dropped, so truncated means real data loss and delta-only
   gaps no longer false-positive.

3. Seq-namespace epoch (Goose stale-cursor recovery): event_replay.EPOCH
   (8-hex per boot) is announced in gateway.ready and echoed by
   session.events.since; the client drops its seq watermarks when the
   epoch changes, so a stale HIGH watermark from a previous gateway
   process can no longer suppress replay/gap-detection forever. Legacy
   backends without an epoch are unaffected (client keeps watermarks).

Validation: Python 85/85 across replay/entry_ws/keepalive/protocol
(12 replay tests, 3 new); vitest 9/9 shared (2 new epoch tests), 66/66
desktop; full tui_gateway sweep 614/615 (1 known ordering flake, passes
in isolation). Live e2e on this tree: 12-event turn -> seqs contiguous
1..12, 8 durable frames buffered + 4 deltas live-only, single trace_id
on all frames, active_turns elapsed_s matches the real turn duration.

Research provenance: NousResearch/hermes-agent#94484 (comparative-scan
comment); techniques credited to OpenHands (transient split), Goose
(epoch/stale-cursor), per maintainer-restructured plan.
2026-08-26 02:55:04 +05:30
Teknium beb7941236 fix(tui-gateway): make WS reconnect replay actually deliver events (follow-up to #94219)
The #94219 replay was a production no-op: the server returned full
JSON-RPC envelopes from session.events.since while the client's replay
loop dispatches only elements with a top-level 'type' — every replayed
event was silently skipped. Each side's tests validated its own
assumption, so both suites stayed green.

- server: events_since() now returns bare event objects (the frame's
  params), the exact shape the live dispatch path consumes; ring stores
  params directly; cross-language contract test added on both sides.
- client: live frames racing an in-flight replay are parked and flushed
  seq-gated afterward — no double dispatch of deltas, no gap-skip from
  a watermark advanced past the replay window.
- restart poisoning: seq counters are in-process, so a backend restart
  reset them while clients kept high watermarks (replay forever empty,
  truncated=false). New replay_epoch advertised in gateway.ready and
  echoed by session.events.since; the client clears watermarks on epoch
  change.
- methods_session no longer reaches into event_replay privates
  (is_truncated() accessor).

Live repro: pre-fix, 3 stamped frames -> 0 dispatchable by the client
gate; post-fix 3/3. Tests: 16 py (replay+ws), 8 vitest, tsc clean, ruff
clean.
2026-08-24 20:11:30 -07:00
kshitijk4poor 87631bd8ae feat(tui-gateway): seq-stamped event replay for lossless desktop reconnect
Server: per-session monotonic seq on every routed event frame, bounded
512-frame replay ring (64 sessions, FIFO eviction), plus two new RPCs —
session.events.since (replay newer-than-watermark, reports latest_seq +
truncated so clients detect gaps) and session.events.stats (telemetry).

Client: per-session seq watermarks recorded from live frames; after any
successful reconnect a fire-and-forget fetchReplay() drains missed events
through the normal dispatch path (recordSeq ignores non-increasing seqs,
so stale replay can never regress a watermark); focus-triggered reconnect
nudge in use-gateway-boot for the Electron unfocused case where macOS wake
skips visibilitychange.

Replay failures are swallowed by design: lossless resume is an upgrade
over the previous lossy reconnect, never a new failure mode.
2026-08-25 02:31:19 +05:30
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
kshitijk4poor 80b202f53a harden(adoption): review findings — exact-id donors only, divergence guard, honest donor_retired
Review batch (3 reviewers) on the final diff surfaced:
- H1: title-based donor matching could adopt AND non-recoverably retire
  an UNRELATED default-store conversation (bot titles collide by design;
  get_session_by_title has no archived filter/ordering). Donor probe is
  now exact-id only — the stranded repro always has the id.
- H2: re-adoption after a partial run could retire a donor that had
  accumulated NEWER messages than the profile copy (skip-based
  idempotency never merges). New divergence guard compares message
  counts and refuses retirement when the donor is ahead (still adopts).
- M1: donor_retired reported True even when every retirement step
  failed under suppress. Now per-segment tracked + warn-logged;
  True only when all applied.
- M3: adopted=False (e.g. import validation limits) was silent — now
  warn-logged with import errors.
- M4: archived donors are never re-adopted (no cross-profile cloning).
- Dead 'from pathlib import Path' dropped; contextlib no longer needed.

5 new red-first-verified regressions (title-collision immunity,
archived-donor immunity, non-vacuous owns_db gating with a real donor
seeded, divergent-donor retirement refusal, donor_retired truthfulness).
tests/tui_gateway: 578 passed. ruff clean.
2026-08-23 18:58:40 -07:00
kshitijk4poor 26a4f89ada fix(gateway): adopt stranded bot sessions from the default store on profile resume
Pre-#93296, the desktop routed session RPCs by the focused tile, so a
profile bot's turns executed on the default backend and its canonical
session accumulated in the DEFAULT profile's state.db. Post-fix, the
profile backend correctly receives the resume — but its store has never
seen the session, so the same chat 4001s forever (unreachable instead
of misrouted). Live repro: Teknium's Developer bot, session c93770.

- hermes_state_portability: SessionDB.adopt_session_lineage_from() —
  composes the existing export_session_lineage()/import_sessions()
  primitives; donor rows are archived (never deleted) with
  end_reason=adopted_by_profile, which is deliberately NOT in
  RECOVERABLE_END_REASONS so canonical-lookup resurrection cannot undo
  an adoption. Idempotent (already-present ids skip).
- tui_gateway/methods_session: profile-scoped session.resume falls back
  to adoption from the default store right before the 4007; ids unknown
  to BOTH stores still 4007 exactly as before, and launch-profile
  resumes never consult the fallback.
- tests: 10 new (7 unit on the primitive incl. compression-lineage
  unit adoption + non-resurrectable archive; 3 handler-level through
  server.handle_request incl. the live repro shape); db-ownership
  leak test taught that the shared launch handle probe is by design.

Follow-up to #93296/#93311; part of #93091.
2026-08-23 18:58:40 -07:00
Teknium b4d4167d42 fix(gateway): lazy/unpersisted resume also rebinds transport and cancels the pending reap
Live WS E2E after the #93361 merge (real web_server + tui_gateway, isolated
HERMES_HOME, 2s grace): drop socket -> re-resume stored id on a new socket
still produced a ws_orphan_reap reclaim. The lazy/unpersisted resume branch
(no state.db row yet -- every fresh Bot Chat) returned the sentinel-parked
live record without rebinding its transport or cancelling the armed reap
Timer, so the storm survived for exactly the Bot Mode sessions the cluster
targeted. The unit-covered paths (_live_session_payload, _reuse_live_response,
_claim_or_reuse_live) were all correct; this branch bypassed them.

Regression test drives the real session.resume RPC against a sentinel-parked
unpersisted record (sabotage-verified: fails without the fix). After the fix
the full live E2E passes 10/10 scenarios including a 4-cycle drop/resume storm
loop with zero reclaim broadcasts.
2026-08-23 18:24:26 -07:00
Teknium 4aa162b30d fix(gateway): cancel pending WS-orphan reaps on resume and supersede stale runtimes quietly
Storm killer for the reap->broadcast->auto-re-resume feedback loop:

- New _pending_ws_reaps registry (sid -> Timer): _schedule_ws_orphan_reap
  registers, _reap pops, and _cancel_ws_orphan_reap(sid) is called from
  every resume/reuse/rebind path — the session.resume fast-path reuse
  (methods_session.py), _claim_or_reuse_live winners, and the
  _live_session_payload live-transport rebind.
- When a resume mints a fresh runtime for stored session id S, any prior
  runtime for S still parked on the detached-WS sentinel is claimed under
  the resume lock, its reap Timer cancelled, and the record finalized
  quietly with end_reason superseded_by_resume — NOT in
  _RECLAIM_END_REASONS, so no session.reclaimed broadcast fires and the
  client's auto-re-resume can't storm.
- superseded_by_resume added to _RECOVERABLE_END_REASONS in
  hermes_state_common.py so canonical Bot Chat resurrection still applies.

Unit tests: resume cancels the reap timer, superseded runtimes finalize
without a reclaimed broadcast, and the normal orphan reap still fires
when nobody re-resumes.
2026-08-23 17:43:39 -07:00
Kyzcreig 14b50f5edd fix(tui-gateway): interrupt turns after websocket disconnect
After the existing reconnect grace, route a still-detached running session through the same interrupt mechanism as session.interrupt. Preserve delegation deferral, sidecar teardown, partial history, and single-owner reap semantics.

Verified on upstream main: RED 4 failed/2 passed without production changes; GREEN 603 related gateway/compute-host tests. Ruff and py_compile passed. Momus pass 2: APPROVE.
2026-08-23 17:43:39 -07:00
kshitij 4865194772 fix(bot-mode): review follow-ups for recoverable-archive resurrection
- Clear the accidental end stamp on resurrection (at the lineage tip):
  a surviving ws_orphan_reap/agent_close reason made a LATER deliberate
  archive auto-resurrect on the next lookup — the user could never retire
  the canonical chat. Test pins the resurrect -> deliberate-archive ->
  stays-archived cycle.
- Judge recoverability at the compression TIP: the registry row of a
  compressed lineage carries end_reason='compression', so tip-stamped
  accidents were unrecoverable through the registry row. Lineage test.
- Heal the third lookup: the api_server exact-title listing (hermes peer
  dm resolution) filtered archived rows out via list_sessions_rich and
  still failed for reap-archived canonical chats.
- Single source of truth for the recoverable set: tuple moved to
  hermes_state_common (mirroring _RESET_END_REASONS_SQL) and interpolated
  into all three recovery SQL sites — literals cannot drift.
- methods_session gate uses BOT_CHAT_TITLE (not a literal) and re-fetches
  by id after resurrection (title has no DB-level UNIQUE).
- Idempotence pinned: two consecutive profiles.list calls both resolve.
2026-08-24 03:27:11 +05:30
kshitijk4poor bef31fb06b fix(bot-mode): resurrect canonical Bot Chat archived by recoverable reasons on reopen (#92687) 2026-08-24 03:06:11 +05:30
Teknium fe2e6b76c4 fix(tui-gateway): messaging a never-used bot no longer fails with 'session not found'
session.create intentionally persists no state.db row until the first
prompt, but session.resume only looked in the database — so resuming a
live lazy session by its stored key or pending title hard-404'd. Bot
Mode hits this on every fresh non-default bot: the canonical Bot Chat
is created lazily on the profile, the open/send resumes it, and the
user gets 'session not found' on their first message to that bot.

session.resume now falls back to the in-memory session registry,
matching by stored key or pending title scoped to the SAME profile
home. Cross-profile lookups still fail closed; unknown ids still 404.
2026-08-23 04:57:50 -07:00
poisdahl fd41164861 fix(history): keep carrier rewinds race-safe after refresh 2026-08-22 17:30:35 +02:00
poisdahl a5b326a471 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821
# Conflicts:
#	tests/agent/test_reference_handoff_active_turn.py
2026-08-22 16:47:39 +02:00
liuhao1024 bd2afde48f fix(tui): never transfer the shared launch SessionDB to one agent
The eager session.resume path called _transfer_db_to_agent(agent, db)
unconditionally. With no non-launch profile selected, db resolves to the
SHARED launch handle (_get_db()), so the transfer succeeded on identity
alone — the agent IS holding that handle — and session.close() then
closed the process-wide database under every unrelated session:
subsequent writes failed with "'NoneType' object has no attribute
'execute'" and the Desktop could not open chats until restart (#91610).
This directly violated _transfer_db_to_agent's own contract ("Never
called for the shared launch handle", introduced with the ownership
lifecycle in #81071).

Gate the transfer on owns_db (dedicated handles only), and add defense
in depth: _transfer_db_to_agent now refuses db is _get_db() even when a
caller invokes it incorrectly.
2026-08-22 15:25:35 +05:30
Teknium a9860d413d fix(bot-mode): the canonical Bot Chat is found by NAME — session-id pins removed
A bot's forever-chat now has exactly one identity: the session titled
"Bot Chat" on that bot's profile. Core UNIQUE(title) makes (profile,
'Bot Chat') an exact registry, and every open consults it directly via
session.list {title, include_hidden}. The stored-id pin
(ui_meta['hermes-bots'].chat) and its entire verification apparatus —
preferred_session_ids resolution, drifted-pin keep branches, last_session
grandfathering, dead-pin recovery re-anchoring, newerVisibleBotChat — are
removed, not deprecated. Legacy ui_meta.chat keys are ignored and dropped
from merges on sight.

Every lost-canonical-chat incident (#88146, #88200, #90524, #90705, and
five hardening waves) traced to that pointer dangling or being stolen,
then later guards welding the wrong session in. A name cannot dangle:
corrupt pins self-heal on first click because the pointer is simply never
read.

Gateway: profiles.list now reports canonical_session per profile row
(registry row resolved server-side by title — hidden rows resolve,
deny-listed sources and archived rows do not, compression lineages
resolve to the live tip), replacing the preferred_session_ids request
contract. The roster preview, activity signals, and the /new→/compact
guard all read canonical_session, so preview identity and click identity
are the same row by construction.

No migration shims: this IS the system.
2026-08-22 01:23:39 -07:00
poisdahl abf87e7248 Merge current main into composite-carrier fix 2026-08-21 15:56:45 +02:00
Teknium 21e9d4532f fix(bot-mode): canonical-chat adoption survives busy profiles via exact-title lookup
The #90732 adoption scan used session.list's 200-row recency window. A busy
bot profile (group-chat traffic, routines, or accumulated fork spam) pushes
an older forever-chat past row 200, the scan misses it, and the mint path
re-enters the unique-title-conflict fork loop — same pathology, higher
trigger threshold.

Profile → Named Session is an exact registry (UNIQUE title index), so
consult it exactly:

- session.list gains a `title` param: indexed WHERE title = ? lookup,
  window-free, hidden rows resolve, archived/deny-listed do not,
  compression lineages resolve to the live tip (resolved_id), mirroring
  profiles.list's preferred_session resolver.
- findExistingCanonicalChat sends title: 'Bot Chat'. Older gateways ignore
  the unknown param and return the windowed listing — the local scan stays
  as the compatibility rung.
- Adoption opens the lineage tip (resolved_id) while pinning the durable id,
  same split as the preferred_session path.
2026-08-20 03:49:15 -07:00
Teknium 66221397a1 fix(bot-mode): always hide Bot Mode sessions from the global Sessions sidebar
Bot Mode's group chats spawned one per-member session per room, and those
"Group: ..." rows (plus canonical Bot Chats when the old eye-toggle pref was
off) flooded the global Sessions sidebar — a 6-bot room dumped six identical
rows into recents (reported with screenshot, Aug 17).

Plugin (apps/desktop/src/plugins/hermes-bots/plugin.js):
- session.create now passes hidden:true UNCONDITIONALLY for both canonical
  Bot Chats and group-room member sessions; the $hideBotChats pref, its eye
  toggle, and its storage hydrate are removed (Bot Mode sessions are plumbing
  or plugin-owned forever-chats, never scratch conversations).
- hideOwnedBotSessions(): idempotent reconciliation sweep over every owned
  session id (bot meta canonical chats + each room's member sessions) via
  session.set_hidden, run on plugin load and on each gateway reconnect, so
  rows born visible under the old pref get cleaned up.
- The Bots session browser and canonical-chat recovery scan pass
  include_hidden:true so they still see the rows they own.

Gateway (tui_gateway/methods_session.py):
- session.list honors an include_hidden param (default off — the resume
  picker and all global callers keep dropping hidden rows).
- session.set_hidden gains a durable fallback: when no LIVE runtime session
  matches, resolve the stored session id in the target profile's state.db
  (via resolve_session_id) and flip the flag there. The sweep holds stored
  ids for chats that aren't live; the old live-only lookup 4001'd them.

Validated E2E with real imports against a temp HERMES_HOME: born-hidden row
(hidden=1), profile-scoped session.list default vs include_hidden (0 vs 1),
and stored-id sweep on a non-live legacy row (hidden=1). Plugin suite
167/167; new RPC regression tests in tests/tui_gateway/test_session_hidden_rpc.py.
2026-08-17 16:02:42 -07:00
poisdahl 7ca1987459 Merge upstream main into PR 81234 2026-08-16 12:20:50 +02:00
Ayush Nangia be8b58dcbc fix(gateway): let an explicit workspace move win for a running session
session.workspace.move refused a running live session with 4009
(session busy), but the desktop's Move-to-project flow calls exactly
this RPC — so the UI updated its local grouping while state.db kept the
old cwd and the agent's tools kept running in the old workspace. Two
sources of truth disagreed (#86626).

An explicit move now wins: the stored row and the live session re-anchor
together. In-flight tool calls keep the cwd they were launched with; the
next tool call uses the new workspace.
2026-08-16 01:59:34 -07:00
Teknium 69f7c655b4 docs(tui): document defer_history vs omit_messages precedence
Follow-up to the #62799 salvage: Desktop sends both defer_history and
omit_messages on a cold resume. Make explicit in the deferred branch that
defer_history supersedes omit_messages — the single history read happens in
the background hydration worker and the synchronous omit_messages read on
the cold-resume default path is skipped entirely, so the transcript is
never loaded twice for one resume.
2026-08-16 01:28:26 -07:00
embwl0x 60be8ef26d perf(desktop): make session resume incremental 2026-08-16 01:28:26 -07:00