Commit Graph

107 Commits

Author SHA1 Message Date
teknium1 f537c0b4c3 refactor(status): CLI, gateway and TUI /status render the same field set from hermes_cli/status_report.py
The three /status renderers (hermes_cli/cli_session_mixin.py::_show_session_status,
gateway/slash_commands_status.py::_handle_status_command, tui_gateway/methods_session.py
session.status) each hand-built Session ID / Path / Title / Model (provider) / Created /
Last Activity / Tokens / Agent Running with their own getattr(agent, "model") fallback
chain, their own updated_at/last_updated_at/last_activity_at scan and their own timestamp
format. A fix to one (a new last-activity column, a placeholder change) silently missed the
other two.

hermes_cli/status_report.py::build_status_fields now derives the common facts once and
returns them as structured, display-ready data; status_lines() renders the English
"Label: value" form for the CLI and TUI. The gateway keeps translating through its
existing t("gateway.status.*") catalog keys (no locale change); the CLI keeps reasoning /
approvals / context, the gateway keeps free-tier / context / queue depth / Matrix scope,
the TUI keeps its Project line. tui_gateway/methods_session.py::_status_dt and the
CLI's inline updated_at loop are gone; cli_session_mixin._timestamp_or stays for its
remaining history-timestamp caller.

Behavior change: none intended for populated sessions. Unified edge cases: a
SessionDB row with an unparseable started_at now falls back to now() on the TUI as it
already did on the CLI, and the TUI's fallback on a bad updated_at is the created stamp on
both surfaces.

Test: tests/hermes_cli/test_status_report_contract.py drives the three real renderers with
one session (distinctive model, provider, title, stamps, token count) and asserts each
output carries every common value. Sabotage-verified: builder dropping tokens -> red;
TUI hand-formatting the model line -> red; restored -> green.
2026-09-13 05:21:02 -07:00
Siddharth Balyan 3b01b4ce0f feat(desktop): Nous free tier on Hermes Desktop (#105260)
* feat(desktop): free-tier state over RPC, status routes that name it, and a sign-in that keeps connectors

The desktop learns about the Nous free tier by reading local auth state (pull): free_tier.status
answers has_guest / enabled / carries_inference / notice_pending with zero network, and
free_tier.ack_notice persists the one-time notice flag on the identity itself. setup.runtime_check
reports free_tier for the selected route; /api/portal, the Nous card in /api/providers/oauth and
billing.state carry free_tier (billing answers the free tier locally instead of a portal call that
can only fail). The free-tier picker row carries an explicit free_tier_row flag and is never priced or
locked. POST /api/providers/oauth/nous/start over a free-tier identity registers the connector
transfer and returns its code and consent URL; the poller waits for the transfer before the token
grant, persists the account, runs settle_after_upgrade, and the poll response gains reason,
account_email and model.

* feat(desktop): free tier on Hermes Desktop: ready screen, notice strip, status chip, Billing view, one sign-in dialog

The renderer reads the free tier from free_tier.status (pull) into one store; the first-launch
intro is the same state rendered two ways, keyed on the backend's one-time flag: the onboarding
overlay opens on a ready screen when the free tier carries inference, else a one-time strip above
the composer. Settings > Billing gains a free_tier view (notice with one Sign in, Plan / Model /
Connectors summary, plan card, footnote; no payment or usage rows). A status-bar chip names the
tier and model while it carries inference. Every entry point opens one claimed sign-in dialog that
drives the extended oauth/nous route and maps the poll's status and reason to the ruled screens;
Done settles billing, model options, providers and re-homes a session still on nous/welcome. The
picker badge also fires on free_tier_row. Docs: Desktop section in the free-tier guide, AGENTS notes.

* fix(desktop): free_tier.status starts the free tier's background setup when no identity exists

A served backend has no session-setup moment like the CLI's, so beside an explicit provider the free
tier was never set up on the desktop: no connectors, no notice strip. The first status read now
starts the same one-attempt background setup; the call itself never waits.

* fix(desktop): one Sign in on the Billing page; Settings > Providers names the free tier, never Connected

The free-tier plan card is the what-you-get text alone (the notice carries the page's one Sign in).
The Nous provider row reads Nous · free tier with a Free tier tag while the identity is the free
tier, instead of Nous Portal · Connected.

* fix(desktop): Settings > Providers never files the free tier under Connected

* fix(desktop): the intro's shape is keyed on the route, not on the identity

free_tier.status reports available (an identity exists and the tier is on); whether inference
runs on the free tier is setup.runtime_check.free_tier, keyed on the resolved endpoint. The ready
screen shows when that route is the free tier; the composer strip when the user's own provider
carries inference. An own-key install used to get the ready screen.

* docs(desktop): say what the free-tier chip is keyed on

* fix(desktop): the featured Nous row's pitch on the free tier says what signing in adds

* fix(desktop): a cancelled or superseded sign-in attempt can no longer change the identity or hide the intro

Four lifecycle holes from review. The Nous poller checks the session's cancelled flag after the
transfer wait, after the token grant, and once more under the session lock together with the
save, so a sign-in the user abandoned never persists. The renderer's sign-in store carries an
attempt generation that every continuation checks after each await, so a poll from a closed
attempt cannot publish over the one on screen (and its backend session is cancelled). The ready
screen comes down only after the backend recorded the acknowledgement. A composer still mounted
takes over the notice claim when its owner unmounts. One thin test per hole.
2026-09-11 03:45:32 +05:30
Siddharth Balyan d5aaaa4a1b fix(tui-gateway): a hidden seed row stays out of search, a partial seed copy is rolled back, live resume counts the wire (#107562)
Two independent reviews of the seeded-create change found three more
places where the newly durable hidden row, or the new create-time copy,
was not handled by the same rule as the rest of the path:

- Message search (dashboard search and the session_search tool) had no
  display_kind filter, so a hidden opening row matched a query the
  person never saw. The shared search predicate now skips hidden rows.
- _seed_row left the fresh session row behind when the transcript copy
  failed after the row was committed. The first prompt's retry copies
  the whole seed, so a kept partial copy would be duplicated. The row
  is now deleted when the copy did not complete, the compensation
  _persist_branch applies to branch children; the first prompt then
  starts clean.
- _live_session_payload (a resume that reuses a live session) reported
  message_count as the raw history length while its messages array was
  filtered. It now follows _resume_response: the stored size when
  messages are omitted, else the wire count.

Tests: the two seeded-create tests now drive the first-submit path
through _persist_session_row_for_submit, the function prompt.submit
calls, and assert search and the reuse-live count; a third test pins
the rollback (no row after a failed copy, one copy after the retry).
2026-09-10 18:05:33 +00:00
Siddharth Balyan c22a8d8e3f Seeded sessions survive a gateway restart and store their seed once (tui_gateway) (#107549)
* fix(tui-gateway): a seeded session is durable at create, and its seed is written once

session.create accepts opening messages. Three defects sat in that path:

- A seeded session without a parent was never persisted at create, so a
  restart before the first prompt lost it and session.resume answered
  4007. Only branch children (#93959) were persisted up front. The
  same rationale applies to any seeded create: seeded content is
  intent, not an abandoned draft. Parentless seeds now persist their
  row, transcript and client title at create; empty drafts stay lazy.
- _coerce_seed_history dropped display_kind, so a seeded row tagged
  "hidden" (model-facing scaffolding) rendered as a user bubble. The
  coercion keeps "hidden" and only "hidden"; every other kind is
  stamped by the gateway at turn time and is not accepted from the wire.
- A branch child's seed was written twice: _seed_branch_row copied it at
  create but never marked it persisted, so the first prompt's
  _persist_branch_seed appended the copy again. The create path now
  sets _branch_seed_persisted, and the gate is a create-time `seeded`
  stamp instead of parent_session_id, so a resumed session (whose
  history comes from the DB) can never re-append its transcript.

Two invariant tests, both red on main: a parentless seed survives a
gateway restart with the hidden row kept out of the wire transcript and
not re-written by the first-submit path; a branch child's seed is stored
exactly once. The reasoning-fields fixture stamps `seeded`, the flag
session.create sets.

* fix(tui-gateway): a hidden seed row stays out of the list preview and the create count

Live-testing the seeded create on every surface showed two places where
the newly durable hidden row (display_kind="hidden") still surfaced:

- session.list built a session's preview from its first user row with no
  display_kind filter, so a hidden opening row (model-facing scaffolding
  the gateway never paints) became the sidebar preview. The preview
  predicate now skips hidden rows, in every listing query that shares it.
- session.create reported message_count as the raw seed length while its
  messages array already filtered the hidden row (2 vs 1). It now counts
  what is on the wire, the same rule session.resume applies.

Both are covered by the existing seeded-create test: the create count
equals the wire transcript, and the preview of a session whose first
user row is hidden is its first visible user row.

* fix(tui-gateway): a live unpersisted resume counts the wire transcript

session.resume on a live session that has no row yet reported message_count as
the raw history length while its messages array was already filtered, the same
mismatch the previous commit fixed on session.create. Count the wire, as the
cold, deferred and reuse-live resume paths already do.

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-10 17:55:24 +00:00
joaomarcos e1b0ee6f03 fix(tui): fail closed on unavailable profile targets matching custom root basenames 2026-09-09 12:39:39 +05:30
joaomarcos 298893df66 fix(tui): resolve default profile session names 2026-09-09 12:39:39 +05:30
Teknium 924c5ded2e fix: scope subagent stops and publish authoritative live progress 2026-09-08 03:06:30 -07:00
Ryan Tucker de25545dce fix(tui_gateway): fan session events out instead of rebinding the transport slot
A session held exactly one transport, and prompt.submit, session.resume, session.activate, and the queued-prompt drain all rebound that slot. A second client therefore took the stream away from the first: the earlier client stopped receiving the turn it was already rendering, and either client disconnecting parked the whole session on the drop sentinel.

FanoutTransport goes in the same slot and satisfies the same Transport protocol, so write_json and every other reader of the slot are unchanged. It delivers each frame to a snapshot of its peers, concurrently when more than one peer is attached and the caller is not on an event loop, and prunes any peer that returns False or raises. A dead client is dropped; a slow one costs the emitter at most one write timeout per frame rather than one per peer. Request/response RPCs are unaffected: they still answer on the request's context-bound transport, so a client only ever sees replies to its own calls.

The rebind sites become attach sites through _attach_session_transport, whose ladder keeps the single-client shape identical. The same object already in the slot is a no-op; an empty, stdio, or parked slot is taken outright; only the arrival of a second live client wraps both. The queued-prompt drain is included because it pinned the drained turn to the queuer and silenced everyone else. A non-peer newcomer such as stdio or the drop sentinel never displaces a live client, so an activate dispatched without a bound websocket cannot silence the socket that owns the session.

Disconnect detaches first. A session that retains another client keeps streaming and is neither parked nor reaped, and only the clientless ones follow the existing close_on_disconnect and park-sentinel path, so a single-client disconnect, the orphan reaper, and its grace window behave as before. _ws_session_is_orphaned is unchanged: it still asks whether the drop sentinel is in the slot, and a fan-out is never the sentinel, so a session that still has a peer is never reported as orphaned.

Attach performs no entitlement check: any authenticated peer may mirror any session.

The fan-out architecture follows the approach in #40822 by @OmarB97.

What the slot's later history forces. _close_sessions_for_transport drops the #83716 rebind-to-the-most-recent-surviving-viewer, which fan-out membership subsumes — a pop-out window is a peer, so a session that still shows in one is never returned as clientless — and keeps the #77129 revalidation before parking, now expressed as a liveness check under _session_transport_lock so it is race-free against attach and detach. _transport_is_live_peer defers its last answer to _transport_is_dead: a socket that already latched _closed is a departed client, and admitting it would keep a session out of both the park and the reap. _transport_is_dead also learns the fan-out: a FanoutTransport with no live peer is dead, so a session whose peers were all pruned by failed writes cannot outlive the TTL and LRU reapers. The upstream test that pinned the #83716 rebind, test_close_transport_rebinds_session_to_remaining_viewer, is re-expressed in fan-out terms: both windows attached, the pop-out closes, the session stays with the main window unparked and still receiving frames.
2026-09-07 22:25:12 -07:00
Teknium 04767e7aaa fix(display): distinguish estimated context from provider usage 2026-09-07 08:13:01 -07:00
kshitijk4poor 311b980bb6 fix(prefix-cache): drop the workspace pin at session boundaries; bind session cwd for /context
The pin from the previous commit lives in _SESSION_STATE but nothing cleared it, so a CLI
/new, /resume or /branch (same AIAgent, reset_session_state + _invalidate_system_prompt)
replayed the previous session's git snapshot into the new session's prompt. Clear it in
reset_session_state next to the other session anchors; one invariant test (red without it).

The TUI/Desktop session.context_breakdown RPC ran the prompt builder on the RPC thread
with no session cwd bound, so it re-probed against the backend's cwd and overwrote the
session's pin — one /context between compactions restored the divergence this fix removes.
Bind the session context around the build like the live rebuild in server.py does.

Also: trim _coding_parts' docstring to the WHY, drop the isinstance/len guard on a value only
this function writes, and remove the tests' assertions on the private pin shape.
2026-09-06 22:45:31 +05:30
kshitijk4poor 212ed99d96 refactor(tui-gateway): lazy resume reattaches through _rebind_live_transport
_resume_live_unpersisted hand-rolled the same transport + viewers + reap-cancel
sequence _rebind_live_transport now owns. Route it through the helper; the
stdio case (no current transport) keeps cancelling the reap as before.
2026-09-06 14:14:29 +05:30
kshitijk4poor c6cb7111a2 perf(tui-gateway): session.activate builds its payload outside _session_resume_lock
The salvaged fix correctly put the activate guard + transport rebind under the
process-wide resume lock, but it dragged _live_session_payload in with them. The
Desktop passes omit_messages=true (cheap), the Ink TUI does not: every TUI session
switch then read the full persisted history from the profile DB while holding the
lock that serializes every resume, disconnect and reap Timer.

Extract _rebind_live_transport from _live_session_payload; activate does guard +
rebind under the lock (the part that must be atomic with grace expiry) and builds
the payload after releasing it.
2026-09-06 14:14:29 +05:30
kshitijk4poor f2077a0209 refactor(tui-gateway): one _reattach_refusal helper for the live-identity/interrupt-settling guard
The 4007/4009 guard the salvaged fix added to prompt.submit, session.activate,
_resume_live_unpersisted and _resume_reuse_live_locked was the same six lines
four times. One helper in session_lifecycle.py (next to _cancel_ws_orphan_reap,
whose contract it mirrors) so the next reattach path cannot drift from the others.
Behaviour unchanged; the 19 race tests cover all four call sites.
2026-09-06 14:14:29 +05:30
BearHuddleston 8513c5984c [verified] fix(gateway): fence orphan timers and reconnect interrupt claims
Re-scope #98106 onto the activity-based orphan policy from #100504. Keep timer ownership across callbacks and continuations, serialize reconnect paths against interrupt claims, and avoid recursive eager-resume locking. Leave cleanup polling armed when concurrent cold reuse is rejected.
2026-09-06 14:14:29 +05:30
Teknium 53db597201 simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
2026-09-03 13:46:50 -07:00
Teknium eeb7671e69 simplify(compat): hermes_cli small facades — drop 7 re-exports/aliases (+relay_runtime alias module), repoint 12 callers/tests 2026-09-03 13:05:57 -07:00
Teknium e7a0f4695b review-fix(suppress-audit): skill_usage hub-lock walk, watchdog tick finally, session.create cwd — restore BASE exception semantics 2026-09-03 10:00:07 -07:00
Teknium ad62a15e94 review-fix(suppress-audit): tui_gateway/methods_session.py — restore BASE exception semantics 2026-09-03 09:43:43 -07:00
Teknium 5209d1d487 refactor(tui_gateway): methods_session — steer/redirect factory, one-line docstrings, final compaction (2134 -> 2099 LOC) 2026-09-03 03:06:33 -07:00
Teknium b1d541d111 refactor(tui_gateway): methods_session — fold oneshot/handoff/pet/spawn_tree ladders, resume mint prologue (2157 -> 2134 LOC) 2026-09-03 02:59:45 -07:00
Teknium 6375896d4d refactor(tui_gateway): methods_session — table-drive billing write routes, fold resume/branch/title ladders (2214 -> 2157 LOC) 2026-09-03 02:55:18 -07:00
Teknium c94d8e3689 refactor(tui_gateway): methods_session — compact resume phases, handoff, spawn_tree and title-read ladders (2235 -> 2214 LOC) 2026-09-03 02:42:37 -07:00
Teknium 104bbe564e refactor(tui_gateway): methods_session — unify db decorators, billing view factory, branch persistence, session_method (2289 -> 2235 LOC) 2026-09-03 02:39:15 -07:00
Teknium 4876472bbd refactor(tui_gateway): fold methods_session resume/pet/branch/compress plumbing into shared helpers (2471 -> 2289 LOC) 2026-09-03 02:27:04 -07:00
Teknium 42e34fa73d refactor(tui_gateway): methods_session — collapse _find_live_unpersisted/_try_get_session, squeeze intra-body blanks, rewrap comments to 100 cols 2026-09-02 23:51:54 -07:00
Teknium 955d362176 refactor(tui_gateway): methods_session — fold steer/redirect into _apply_correction, _save_via_compute_host, _int_param/_b64 helpers, compact payload literals and docstrings 2026-09-02 23:40:05 -07:00
Teknium 8c62bc058f refactor(tui_gateway): methods_session — _listing_rows helper, _Resume.claim/child_history, compact create/resume payloads 2026-09-02 22:53:19 -07:00
Teknium 99433742dc refactor(tui): extract the prompt turn into prompt_turn.py; compact methods_*/session_*/agent_callbacks; split _start_agent_build into scope/wiring helpers
- server.py 7663 -> 5319: _run_prompt_submit and its goal/loop/voice/scope
  phases move to tui_gateway/prompt_turn.py (bound via method_ctx.bind_module);
  _start_agent_build split into _bind/_release_build_profile_scopes,
  _deferred_build_agent_kwargs, _wire_session_agent, _start_session_services;
  _load_enabled_toolsets split (_enabled_mcp_server_names, _resolve_explicit_toolsets).
- methods_slash: _LIVE_SLASH_OUTPUT dispatch table; methods_tools: _SLASH_BUILTINS,
  _guarded; tool_progress: _PROGRESS_HANDLERS; methods_config_set:
  _REASONING_DISPLAY_WORDS; methods_voice: _VOICE_TOGGLE_ACTIONS.
- Unified: _denied_source (methods_session) replaces _WORKER_SOURCES in
  methods_profiles; _compress_live_with_feedback / _compute_host_slash shared by
  the slash mirror and /compress; _end_voice_chat shared by stop-phrase paths;
  _watcher_mtime_ns; _reaper_session_is_detached_idle; _notif_* helpers.
- Dead: _profile_dir_or_err, _resume_info, _slash_builtin_table (refs.py: 0 hits).
- acp_adapter/server.py: docstring compaction only (AST-identical).
- Every file semantically reviewed hunk-by-hunk for wire/log/lock/order parity.
2026-09-02 14:09:13 -07:00
Teknium 28bc8d05b9 refactor(tui): restore compact WHY notes lost in earlier docstring compaction 2026-09-02 14:09:08 -07:00
Teknium a0a41abf72 refactor(tui): move slash.exec mirror + completion helpers out of server.py 2026-09-02 14:09:06 -07:00
Teknium beacf4e942 refactor(tui): resume — verified partial work (acp_adapter + methods_* compaction, bind_module via globals()) 2026-09-02 14:09:01 -07:00
aeonsong 8076c78c87 fix(desktop): scope handoff config to session profile 2026-09-02 06:17:47 -07:00
Teknium 54b2c0ea38 fix(tui_gateway): scope live-session reuse to the requesting profile
`_find_live_session_by_key` matched live runtimes by bare stored session id.
Stored ids are timestamp-based and can exist in more than one profile's
store, so `session.resume` for profile B (fast path, post-build re-check, and
`_claim_or_reuse_live`) could hand back profile A's live runtime — the turn
then ran with A's persona/tools and wrote A's memory (#100029).

Give the lookup an optional `profile_home` (default: any profile, unchanged
for callers that have no profile to scope by) using the same string compare
`_find_live_unpersisted` already uses, and pass the resolved home at every
resume/claim site. `_claim_parked_runtimes` gets the same scope so a resume
under B never finalizes A's parked runtime of the same id.

Reimplements the profile-scope half of #100213 by @Finn763; the Group-title
capability-sync change from that PR is intentionally not carried.

Co-authored-by: Finn763 <165816600+finn763@users.noreply.github.com>
2026-09-02 05:40:49 -07:00
Teknium aa80626764 fix(tui-gateway): adopt late compute-host compress acks instead of a false 120s timeout (#97948)
Manual /compress on a compute-host (turn_isolation) session blocked its RPC
waiter for a hard-coded 120s, answered error 5019, and then DROPPED the
host's late `control.ack`: HostSupervisor.control() popped the pending
queue in `finally`, so `_handle_host_frame` had nothing to deliver to. The
host kept compressing, succeeded minutes later, rotated the session — and
the gateway session never mirrored the new session_key/history_version and
the desktop never refreshed its transcript.

- host_supervisor: `control(..., on_late_ack=)` leaves a one-shot handler
  registered when the waiter times out; control.ack/control.error/error
  frames for that request_id fire it (bounded: 30min TTL, cap 64). A host
  crash fails outstanding handlers with a synthetic control.error.
- server: `_compute_host_compress_wait_seconds()` derives the wait from
  `compression.context_total_ceiling_seconds` (+30s slack, floor 120s,
  cap 630s) instead of the literal 120. `_adopt_late_compute_host_compress_ack`
  applies the metadata mirror and emits the same `session.info` a normal
  compress does plus the existing `status.update kind=compacted` edge; a
  late error goes out through the existing `error` event.
- session.compress / slash.compress (methods_tools + _mirror_slash_side_effects):
  on waiter timeout answer `status: pending` (not 5019) and register the
  late-ack handler.
- desktop: SESSION_COMPRESS_TIMEOUT_MS 120s -> 660s (above the gateway cap);
  `status: 'pending'` renders as an info notice, not `error:`; the
  `compacted` status edge rehydrates an idle active session's transcript
  (mid-turn compaction still defers to the turn settle path).

Minimal extraction of the design in #99630 by @vsd2807 (design trace by
@andrexibiza and @JoaoMarcos44 in the #97948 thread); no new DB tables,
modules, or polling protocol.

Refs #97948

Co-authored-by: VVV <vaibhavdahiya28@gmail.com>
2026-09-02 01:35:59 -07:00
Teknium 52f359a011 fix(sessions): Desktop resume of a heavily-compacted chat no longer fails with 4130
Desktop's cold resume (defer_history + omit_messages, transcript paged over
REST) only ever holds the live tip segment in memory, but session.resume
bounded it against the FULL compression lineage (sessions.max_resume_messages,
default 20000). A Bot Chat with 85 compaction segments / ~29k lineage rows
behind a ~700-row tip was refused at 20001, sent zero model prompts, and sat on
"Waking up default…" forever — the healthiest possible session shape, rejected
by a guard sized for in-memory materialization.

- hermes_state: one `_resume_lineage_ids` definition shared by the resume
  readers (get_resume_conversations, get_ancestor_display_prefix) and the
  guard (assert_resume_safe / get_resume_message_count). Guard grows
  `tip_only=` and names the scope it counted; the branch-aware lineage the
  readers already used is now what the guard counts too (a /branch copy was
  being counted against its parent's rows).
- tui_gateway session.resume: deferred, omit_messages and lazy resumes are
  bounded by the tip; only the full in-memory lineage resume keeps the
  lineage-wide bound. Deferred hydration falls back to tip-only history when
  the lineage exceeds the limit instead of loading the rows the guard refused.
- CLI mid-setup tip-only path routes through the same guard instead of
  borrowing assert_export_safe.
- docs: sessions.max_resume_messages / max_export_messages documented with the
  per-surface scope.

Live repro (real SessionDB fixture, 85 segments / 29,226 lineage rows / 666 tip
rows, real tui_gateway.server.handle_request): before — deferred resume ->
4130; after — ok, hydrated history=666 prefix=0; the non-deferred full resume
still returns 4130 on the same fixture.
2026-09-01 22:14:33 -07:00
ClintonEmok 78eb4ebbfc fix(gateway): compensate half-written seeded branches + observable failures
Review follow-up on #93959:

1. Partial-failure window: if the row commits but the transcript copy or
   title write fails, the durable-but-empty child defeated the lazy
   first-prompt fallback (_ensure_session_db_row is INSERT OR IGNORE), so
   the renderer fail-latched on a transcript-less session again. The seed
   block now compensates: delete just this child so the lazy path can
   retry cleanly. Disk-full is exempt — deleting data on a full disk makes
   things worse.

2. Silent degradation: the best-effort catch now logs at WARNING with
   exc_info instead of DEBUG, so a regression in this user-facing path is
   observable without enabling debug logs.

Tests: compensation deletes the half-written row and preserves
pending_title; disk-full keeps the row and surfaces the WARNING.
2026-09-01 22:58:48 -05:00
ClintonEmok 6556439765 fix(gateway): persist seeded branch children at session.create (#93959)
Desktop branch creation hung on an infinite spinner and lost the branch
on restart. Root cause: the renderer branches via session.create with
parent_session_id + a seeded transcript, but session.create defers the
DB row to the first prompt (the draft-hygiene contract). The renderer's
post-create resume then re-fetches the fresh child through REST and
defer_history hydration — both read the DB. An unpersisted child 404s
and hydrates empty, the client fail-latch (sessionShouldHaveTranscript +
empty messages) refuses to bind a "transcript-less" session, and the
user sees a spinner forever; on restart the rowless child vanishes and
the optimistic "Draft: Branch N" entry disappears with it.

A seeded branch is explicit user intent, not an abandoned draft.
session.create now persists the child immediately when both
parent_session_id AND seeded history are present:

- Row created in the PARENT's profile-scoped state.db, stamped with
  _branched_from + parent_session_id (same shape as TUI /branch).
- Seeded transcript copied via append_messages_batch so REST prefetch
  and defer_history hydration find it on the first read.
- Title assigned from get_next_title_in_lineage(parent) and cleared
  from pending_title — the branch lands in the parent's lineage instead
  of falling back to a message-preview name.

Persistence is best-effort: a broken DB logs and lets create succeed,
leaving the lazy first-prompt path as fallback. Plain drafts keep the
lazy-row contract unchanged.

Fixes #93959
2026-09-01 22:58:48 -05:00
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
Brooklyn Nicholson 5c3bfb6da8 fix(tui_gateway): retire recovery marker on local stop
A confirmed Desktop Stop left the crash-recovery marker on disk until the
run thread finished. If the backend exited in that window, resume treated
the leftover as a crash and auto-continued the turn the user had stopped.

Co-authored-by: Jaime Chieng <164842890+buddhaholic420@users.noreply.github.com>
2026-08-31 20:50:39 -05:00
Brooklyn Nicholson adb23c13cb fix(bot-mode): keep Bot Chat resume on a proven compression tip
Desktop opens the registry id, then session.resume walked the legacy
unmarked-child fallback, so Open Chat still landed in a side chat after
the title lookup was already strict. Recoverable-archive resurrection
uses the same helper.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
2026-08-31 20:47:53 -05:00
Gille bb28056efd fix(bot-mode): keep canonical lookup on compression lineage 2026-08-31 20:47:53 -05:00
fangliquanflq da090aa4ba fix(prompt): preserve resumed workspace provenance 2026-08-31 10:10:25 -07:00
fangliquanflq c6ee4e0809 fix(prompt): skip bundled AGENTS.md for desktop launch cwd 2026-08-31 10:10:25 -07:00
Teknium 6874b99d49 fix(sessions): stamp launch-profile name on new session rows instead of NULL
Sessions created on the launch/default profile were persisted with
profile_name = NULL by all three writers (run_agent._ensure_db_session
None'd out 'default'; the desktop backend's _ensure_session_db_row and
session.branch passed None when no profile_home override was set).

NULL used to mean 'launch profile' by convention, but the desktop now
keys sessions by (profile, id), filters the sidebar by profile scope,
and resolves @session:<profile>/<id> deep links by profile match — a
NULL row matches nothing, so sessions created around a profile switch
vanished from the sidebar and their deep links could not be opened
(#99222). The #94724 one-shot legacy-owner backfill stamps literal
'default' onto old NULL rows, so writers minting NEW NULL rows after
that backfill ran recreated the exact state it exists to repair.

Stamp the real profile name at creation time in all three writers.
E2E-verified against a temp HERMES_HOME: both the desktop create path
and the agent path now persist profile_name='default'.

Fixes #99222
2026-08-31 05:56:18 -07:00
David Dudok de Wit 93c7089f70 feat(bot-mode): run same-gateway Group Chats without Desktop 2026-08-30 22:19:06 -07:00
Teknium c0875ba503 fix(tui): derive resume todo snapshots from already-loaded history
The eager _read_persisted_todo_state(db, target) added a second
get_messages_as_conversation call on every resume, breaking the
one-lineage-SELECT contract pinned by
test_session_resume_uses_parent_lineage_for_display. Derive the
snapshot from the history each resume path already loaded instead;
deferred (defer_history) resumes cache it in the hydration worker once
the transcript arrives.
2026-08-29 18:40:51 -07:00
itsflownium 393af4a310 fix(todo): live task state via revisioned snapshots and a dedicated todo.updated event
Salvaged from PR #97815 by @itsflownium, slimmed to the schema-free core:
- TodoStore gains a monotonic in-memory revision; the todo tool result
  returns it so clients can reject stale updates
- tui_gateway emits a dedicated todo.updated full-snapshot event that
  bypasses optional tool-progress display settings
- session resume/activate responses attach the authoritative todo
  snapshot; renderer restores it with revision arbitration
- desktop store tracks per-session revisions and rejects regressions

The session_todo_state DB table from the original PR is intentionally
dropped: canonical todo tool results already persist in conversation
history, so resume paths derive the snapshot from the stored transcript
instead of a parallel store.
2026-08-29 18:40:51 -07:00
Teknium e05c91ac71 fix(cli): slow /handoff transfers no longer misreported as "gateway not running"
Live-reproduced on main: /handoff poll-waited a flat 60s for a TERMINAL
state, but the gateway's dispatch is a full synthetic agent turn (whole
transcript replay + delivery) that routinely exceeds 60s on long sessions.
The CLI then printed "Timed out waiting for the gateway. Is `hermes
gateway` running?" (false diagnosis), called fail_handoff() on the RUNNING
row (stomping the gateway's claim), and promised "Your CLI session is
intact" after switch_session had already re-pointed the session. The
watcher later overwrote failed -> completed: split-brain.

- hermes_state.fail_handoff gains only_states CAS; waiters can only fail
  rows still pending. Owner (gateway watcher) keeps the unconditional form.
- CLI wait loop is two-phase: 60s for the CLAIM (pending) — a timeout
  there really does mean no gateway — then up to 15 min for the claimed
  dispatch with 30s heartbeats; a running row is never failed by the CLI.
- Desktop handoff.fail RPC now CAS-fails pending rows only; a running row
  returns {failed: false, state: running} instead of stomping the claim.

Repro (real _handoff_watcher, real state.db, CLI as separate process,
75s dispatch): before — CLI timeout @60s + false message + row stomped;
after — pending->running@5s->completed@80s, clean CLI exit.
2026-08-29 18:08:55 -07:00
David Tyler 84e17db0bd fix(bot-mode): canonical bot DMs always follow the profile's current config
Bot-Mode canonical chats (the ONE forever DM per bot) and room plumbing
sessions are plugin-owned scratch conversations. They are now created with
an explicit follow_profile_config contract, persisted in the session row's
model_config, so session.resume rebuilds from the member profile's CURRENT
config instead of restoring the stored model/provider pin from an old row.

That stale pin is what left bot DMs stuck on a dead provider (e.g. 'out of
Nous credits' after the profile was switched to ollama-cloud) while the
same bot worked fine in rooms — the mirror image of the room-plumbing bug
(#89497 class). Normal 1:1 user chats keep the stored-runtime restore:
opening an older chat must show the model it actually used.

- tui_gateway/methods_session.py: accept follow_profile_config on session.create
- tui_gateway/server.py: persist the marker in the row; skip stored-runtime
  overrides on resume when present
- apps/desktop/src/plugins/hermes-bots/plugin.js: send the contract from
  createCanonicalChat and ensureGroupChatSession
- tests: backend override + row-persist coverage; desktop source-contract
  coverage for both session kinds
2026-08-28 02:59:17 -07:00
David Tyler 316e51ae72 fix(bot-mode): room plumbing sessions always follow the profile's current config
Room member sessions in Bot Mode are per-member scratch conversations
inside a group chat. session.resume restored their stored model/provider
pin from the row's model_config, so a room bot stayed stuck on whatever
provider was pinned when the row was first written — even after the
profile was switched. Every room message then failed on the stale
provider (e.g. 'out of Nous credits' after switching a profile from
Nous to ollama-cloud) while the same bot worked fine in DMs.

Add an explicit room_plumbing contract:
- session.create accepts room_plumbing: true, persisted in model_config
- _stored_session_runtime_overrides() returns {} for marked rows, so a
  room session always rebuilds from the member profile's CURRENT config
- hidden + 'Group:' title shape is kept as a legacy fallback for rows
  created by older desktop builds that never sent the marker; hidden
  non-room chats keep the stored-runtime restore
- Desktop Bot Mode sends room_plumbing: true when creating the hidden
  per-member room sessions

Fixes #89497
2026-08-28 02:59:17 -07:00