* feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable)
Adds a source-orthogonal, archive-orthogonal 'hidden' session flag meaning
'don't show in the global Sessions sidebar, but stay fully resumable by the
surface that owns it'. Mirrors the existing archived/pinned capability end to
end, so it's a generic widening (any plugin that owns its own session lifecycle
- kanban, Bot Mode, future plugins - can keep its sessions out of the shared
recents list) rather than a per-plugin special-case.
- Schema: hidden INTEGER NOT NULL DEFAULT 0 on sessions (additive; lands on
existing DBs via the declarative _reconcile_columns ADD COLUMN path, same as
archived/pinned - no version-gated migration).
- DB: SessionDB.set_session_hidden(session_id, hidden) (clones set_session_pinned
incl. the compression-lineage recursive CTE); list_sessions_rich gains
include_hidden=False, appending 's.hidden = 0' by default so hidden rows drop
from every listing path (and the REST sidebar endpoints inherit it with no
change).
- Gateway: session.set_hidden RPC (mirrors session.title); session.create accepts
hidden=true, deferred via pending_hidden and applied in _ensure_session_db_row
when the row is lazily created (mirrors pending_title).
- REST parity: PATCH /api/sessions/{id} accepts+bool-validates 'hidden' ->
set_session_hidden; _session_response exposes it.
Enables Hermes-Bot-Mode to hide canonical 'Bot Chat' sessions from the sidebar
(NousResearch/Hermes-Bot-Mode#46) WITHOUT retagging source (which would mis-set
the agent platform). Bot Chats keep source=desktop. Gateway RPC needs a
SERVE-backend restart to take effect live. 1 focused test (default-exclude /
include_hidden / unhide round-trip).
* fix: teach lost-and-found recovery about the 55-column sessions layout
Adding the 'hidden' column makes the current sessions table 55 columns. The
SQLite lost-and-found recovery classifier keys off the physical field count
(SESSIONS_LAYOUT_NFIELDS) to identify a salvaged sessions row, so a recovered
current-layout row (nfield=55) would otherwise be unrecognized and dropped.
Add 55 to the frozenset (54/52 stay as historical prefixes) and update the
column-count assertions + synthetic current-layout insert in the recovery test.
---------
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
The original __del__ only closed _conn (the writer connection),
skipping the read-only connection pool, token writer thread, and
atexit unregister. Delegates to self.close() instead so all
cleanup paths run. Uses __dict__.get('_conn') guard to stay
safe on partially-constructed instances and during interpreter
shutdown.
Two call sites create SessionDB instances without closing them on error:
1. gateway/slash_commands.py: /insights command — db.close() was on the
success path but not in a finally block, so exceptions between
SessionDB() and db.close() leak the connection.
2. hermes_cli/sessions_cmd.py: sessions repair — SessionDB() created
inline with no .close() at all, leaking the FD on every call.
Additionally, add a __del__ safety net to SessionDB itself so that
instances orphaned by callers who forget .close() are cleaned up when
garbage collected, rather than pinning FDs alive until process exit
via the atexit hook.
Fixes#83226
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).
- Close partially initialized SessionDB connections on every constructor
exception path via a finally block guarded by an initialization-complete
flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
thread, close on worker exit, reject new enqueues after shutdown starts,
and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
API disconnect failures, shutdown recovery, RetainDB late enqueue, and
foreign-loop async clients.
Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A turn writing against a session already closed by compression died with
session_persistence_failed and a misleading "this is often a full disk"
dialog, even though the store was healthy and a live continuation existed
(#82001). Depth-1 recovery (find_live_compression_child) could not resolve
lineages with >=2 compression hops (root -> mid -> tip), reproduced
independently on two- and three-hop chains.
- run_agent.py flush chokepoint: on CompressionSessionClosedError, resolve
tip = db.get_compression_tip(old_id) (canonical bounded transitive walk),
adopt only when tip != old_id AND the tip row is live, retry the flush
exactly once (adoption budget); otherwise fail closed.
- gateway/session.py append_to_transcript: replace the depth-1 live-child
lookup with the same tip + liveness contract, so gateway transcript
reroutes follow full chains.
- agent/conversation_compression.py _adopt_live_compression_child: turn-start
recovery preflight now resolves via get_compression_tip with the same
liveness check, closing the last depth-1 consumer in this family.
- classify_persistence_error: new "compression_closed" bucket; the turn-end
explanation names compression rotation and tells the client to refresh the
session id instead of blaming a full disk.
Tests: depth-1 adoption, multi-hop chain adoption (agent + gateway), fail
closed with no continuation / stale-closed (ws_orphan_reap) tip, exactly-once
adoption budget, and error-wording guards (compression-closed never mentions
disk; real disk failures keep disk guidance).
Closes#82001
Co-authored-by: Al3xand3r1987 <125030427+Al3xand3r1987@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
Dedupe key now includes tool_call_id/tool_calls/tool_name: compaction
copies carry those fields verbatim, so identical tool messages across
generations still collapse, while distinct tool calls sharing
role/content/timestamp are never merged. Add endpoint-level coverage
for the desktop's real read path (limit + order=latest +
include_compacted=true).
Refresh-loss interrupt is cooperative, so a stalled writer could still flush after another process reclaimed the conversation. Carry the holder into append_message / append_messages_batch and reject the write in the same SQLite transaction when the lease row is missing, expired, or owned by someone else.
Presence-only _delegate_from/_branched_from checks stopped the lease walk on
continuations that copied a delegate's model_config, so the first refresh
after rotation missed the parent-key lease and hard-interrupted. A failed
get_session probe also skipped acquire entirely. Walk the lineage inside
the write transaction and treat a probe error as contended, not a fresh
session.
Honor interrupts while waiting for admission, stop the turn when refresh
loses the lease, poll once per second under contention, and test dead-PID
reclaim.
Ports Claude Code's /loop (and its /proactive alias) across every Hermes
surface. /loop [interval] <prompt> re-runs a prompt or slash command on a
recurring cadence inside the live session; omitting the interval enables
self-paced mode (starts at the floor, backs off exponentially while the
agent's replies stop changing, snaps back on change — local digest
comparison, zero extra LLM cost).
Stop conditions: agent-emitted LOOP_COMPLETE marker, --times N,
--until <condition> (judged by the existing goal_judge aux task,
fail-open), /loop stop, and a loops.max_ticks backstop budget.
Core: hermes_cli/loops.py (LoopState + LoopManager + shared
dispatch_loop_command), persisted per session in SessionDB state_meta
(loop:<sid>) so /resume picks it up; migrates across compression
boundaries like /goal. New SessionDB.list_meta_prefix() powers the
gateway's cross-session scan.
Surfaces:
- CLI: /loop handler + idle-fire and post-turn-complete hooks in
process_loop (mirrors the /goal hook shape; Ctrl+C pauses the loop)
- Gateway: /loop handler with route capture, mid-run control-verb guard,
post-turn tick completion, and a supervised loop_wakeup_watcher that
injects due wakeups into idle chats via the synthetic-message path
- TUI/dashboard/desktop: command.dispatch handler + per-session
notification-poller wakeup driver + post-turn completion in the turn
dispatcher; /loop added to the desktop slash palette
- /goal mixing: an active non-parked goal owns the idle boundary — loop
ticks defer until it finishes, pauses, or parks; real user input always
wins over both
Config: loops.{min_interval_seconds,max_ticks,self_paced_floor_seconds,
self_paced_ceiling_seconds}. Docs page + sidebar entry. 77 new tests.
Slack's 50-slash cap: /version moves to /hermes version to free the
native slot for /loop.
Salvage of #79604 (webtecnica) + #85721 (pierrenode), combined and
rebased onto current main with simplify-code findings folded in.
#79604: update_session_model() wrote the model name to sessions.model
but never persisted the provider into model_config. On resume, the
runtime recombined the persisted model with the config.yaml primary
provider (which may not serve that model), producing auth errors.
Fix: add optional provider parameter to update_session_model, merged
into model_config via the shared _merge_model_config_json helper (not
hand-rolled SQL). Wire both gateway /model call sites to pass
result.target_provider.
#85721: session_gateway_runtime() had no billing_provider fallback.
A CLI session that never ran /model has no gateway_runtime or
top-level provider in model_config — billing_provider (written on
every session's first accounted API call) is the only durable record.
Fix: add billing_provider as the last-resort fallback in
session_gateway_runtime(), filtering bare billing buckets (auto/custom)
that are not routable identities.
Simplify-code findings addressed:
- Use _merge_model_config_json instead of 40 lines of branched SQL
- Share _BARE_BILLING_PROVIDERS from hermes_state.py (was duplicated
as a set in tui_gateway/server.py)
- Merge None-filtering from #85920 with the billing_provider fallback
into one coherent return path
Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
The route dict in _persist_model_switch_to_session used filtering
(omits falsy values) while the top-level keys used (writes
explicit None to trigger deletion in _merge_model_config_json). This
asymmetry meant stale keys from a previous /model switch survived in
the nested gateway_runtime dict even after the fix in #85261 that
properly deleted them from the top-level keys.
Fix: build the route dict with and derive the top-level keys
from **route so both shapes always use identical deletion semantics.
Also filter None values in session_gateway_runtime's reader since
gateway_runtime is replaced as a whole dict (not deep-merged), so None
values written by the persist path survive in the nested dict.
Found by /simplify-code 3-reviewer review on #85261 (all 3 reviewers
converged on the route dict asymmetry as the verdict-relevant finding).
- Extract the two duplicated /model session-persist blocks into
_persist_model_switch_to_session; persist the route BOTH nested
(gateway_runtime, CLI reader) and top-level (TUI gateway's
_stored_session_runtime_overrides reader) so a CLI switch also
survives a desktop/TUI session.resume.
- Add SessionDB.session_gateway_runtime as the canonical tolerant
row-level route reader (session_yolo_enabled precedent); use it in
_restore_session_model instead of hand-rolled JSON parsing.
- Clear stale launch-time _explicit_api_key/_explicit_base_url when
resume restores a different provider (same leak guard
_apply_model_switch_result already has).
- 12 new tests incl. a real-SessionDB round trip; mutation-checked.
Follow-ups to the salvaged #84009 commits:
- Add 'session_switch' to _RESET_END_REASONS: a reset continuation's
parent can be promoted to session_switch (resume the reset parent,
then switch away), which permanently hid pre-marker legacy children —
reopen-time stamping cannot rescue them because the parent is being
ended, not reopened. Probe-verified before/after.
- Share the legacy reset-child heuristic via _legacy_reset_child_sql()
so _RESET_CHILD_SQL and reopen_session()'s stamping UPDATE cannot
drift, and derive find_latest_gateway_session_for_peer's two recovery
fence literals from _RESET_END_REASONS_SQL (was a third hand-written
copy of the same set).
- Exclude reset children (marker + legacy shape) from the
resolve_resume_session_id forward walker: resuming a reset parent
could redirect into the post-reset conversation — the exact context
the user reset away. Regression tests cover both shapes plus the
walker's original compression-tip behavior; mutation-checked.
Address rewinds/edits via SQLite messages.id (truncate_before_row_id)
instead of shifting user ordinals. Resolve against in-memory stamps,
then durable session history when live turns drop _row_id; refuse
unknown durable targets with 4018 (no ordinal fallback) and 4030 on
ordinal/row_id mismatch. Stamp _row_id on insert, load row ids on
resume paths, send rowId from Desktop, filter renderer-synthetic ids,
and stop silently resending failed targeted edits without truncation.
Add production-shaped SessionDB tests for resolve and fail-closed paths.
Fixes#82959
Salvaged remainder of PR #82280 (state.db hardening rollup):
- Runtime connection corruption: a sibling process replacing/truncating
the backing file breaks the live write connection — every subsequent
write raises 'file is not a database' and the gateway wedges
permanently (messages pile up in memory). Add a bounded one-shot
reconnect on the write path: close the broken connection, reopen the
DB file (re-running WAL activation + schema reconciliation), retry
the failed write once.
- _on_disk_journal_mode: retry transient 'disk i/o error' (virtualized
block devices) a few times before returning None, so a one-shot EIO
doesn't push callers onto the fail-closed unknown-mode branch.
The rollup's write-lock machinery, checkpoint-strategy changes, and
repair serialization are intentionally NOT included — superseded by
PRs #84277 and #69609, or wrong-direction per the POSIX
lock-cancellation findings (#71724 lineage).
`sessions optimize` could consume several GB of disk instead of freeing
any, filling the host to 100% on exactly the large databases it exists to
shrink.
Two causes, both in the WAL lifecycle:
1. No `journal_size_limit`. SQLite defaults to -1 (unlimited), so after a
checkpoint the WAL is reused in place and never truncated —
`state.db-wal` permanently keeps the high-water mark of the largest
transaction ever run. `hermes_cli/kanban_db.py` already bounds its WAL
with `wal_autocheckpoint=100`; the session store, by far the larger
database, had no equivalent.
2. `vacuum()` checkpoints BEFORE `VACUUM` but not after. VACUUM rewrites
every page through the WAL, so the pre-checkpoint does nothing about
the slack VACUUM itself creates.
Measured on a 3.0 GB state.db: `hermes sessions optimize` reported
"3143.9 MB -> 3155.1 MB (reclaimed -11.2 MB)" while leaving a 3.07 GB
state.db-wal behind. Free space fell from 6.9 GB to 772 MB (100% full)
and stayed there. A manual `PRAGMA wal_checkpoint(TRUNCATE)` recovered
the full 3.07 GB, confirming it was slack, not data.
Fix: set `journal_size_limit` (64 MiB) when enabling WAL, and truncate
the WAL again after VACUUM. Both are best-effort and never raise — a
failure costs disk slack and must not stop the DB from opening.
Tests assert the contract (limit is a finite positive bound; VACUUM does
not leave an oversized WAL) rather than pinning the byte count, which is
a tunable. They skip where WAL is unavailable — including hosts where
Hermes falls back to journal_mode=DELETE due to the SQLite 3.50.4
WAL-reset bug.
Verified: 462 passed / 3 skipped in tests/test_hermes_state.py, and
_apply_wal_size_limit flips a real WAL database from -1 to 67108864.
Tested on Linux (aarch64, Python 3.11).
The Aug 2026 incident in #69603 documented a fail-open: when the
pre-repair backup was refused (another same-process handle open),
repair_state_db_schema() recorded backup_path=None and proceeded —
leaving the writable_schema surgery, FTS-schema deletion, REINDEX and
VACUUM strategies reachable against the only remaining copy of the
damaged DB.
_backup_db_file() now returns (path, reason) and the repair path treats
any refused/failed backup as an unconditional hard stop: abort before
the first mutating strategy and surface the reason in report['error'].
Explicit backup=False (CLI --no-backup) is unchanged — that is the
operator opting out, not a silent failure.
Three new tests: refusal hard-stops with source bytes untouched,
OS-level copy failure hard-stops with the reason surfaced, and
backup=False still repairs.
`repair_state_db_schema()` performs `PRAGMA writable_schema=ON` +
`sqlite_master` surgery + `VACUUM` on a private connection. The only guard
around it is `_repair_attempt_lock`, a `threading.Lock`, whose docstring
claims it "serialises concurrent web_server / gateway opens" — but a
threading lock covers threads inside one interpreter, not processes.
A normal host runs four independent processes against the same state.db:
the gateway service, the Desktop app's own `hermes serve` backend (it
spawns one per launch, not a thin client), interactive CLI sessions, and
the TUI slash worker. When two of them hit a malformed DB, both entered
the critical section and each ran the full surgery while the other was
mid-rewrite. Observed as a repair/re-corrupt cascade: the DB is repaired,
then re-corrupts minutes later, repeatedly.
Two fixes:
1. Wrap the surgery in a bounded `flock` on `<db>.repair.lock`. `flock` is
the right primitive — the kernel drops it when the holder dies, so a
crashed repairer cannot wedge future repairs the way a pidfile would.
The acquire is bounded (#36644's failure shape) and, unlike the kanban
init lock, a caller that times out must NOT proceed: here "proceed
anyway" is exactly the unsafe interleaving. It re-probes instead, and
reports success if the holder already healed the file.
Under the lock, the existing `_db_opens_cleanly()` check becomes a
double-check: a queued process finds the DB healthy and returns
`already_healthy` rather than re-running surgery on a repaired DB.
2. Bump the schema cookie after direct `sqlite_master` edits. Ordinary DDL
bumps it for free and every other connection compares it before running
a prepared statement — that is how they learn to drop a cached schema.
Editing `sqlite_master` under `writable_schema=ON` does not, so live
connections in other processes kept writing `messages` rows through
triggers into `messages_fts*` shadow tables the surgery had just
deleted. SQLite's writable_schema docs call out incrementing
`schema_version` as the required companion to such an edit.
Tests: four new cases in tests/test_state_db_malformed_repair.py, all
using real child processes and a real flock. All four fail on main and
pass with this change; the concurrency case asserts exactly one
`malformed-backup-*` file is produced by two simultaneous repairers
(two on main). Full state suite: 558 passed.
Complements #43742, which makes the *in-process* claim loser retry rather
than raise; it explicitly leaves `repair_state_db_schema()` unchanged and
does nothing cross-process. The two are independent and compose.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
SessionDB.close() ran `PRAGMA wal_checkpoint(TRUNCATE)`. Every cron
run_agent opens and closes its own transient SessionDB, so on a busy
fleet this fired a full WAL reset many times an hour, racing the
gateway's long-lived writer on a large WAL database and tearing hot
B-tree pages -- structurally the same corruption this module's own
periodic checkpoint was already switched to PASSIVE to avoid (#45383).
Only close() and two manual-maintenance paths still used TRUNCATE.
Route every checkpoint on the shared state.db through PASSIVE:
- close() (hermes_state.py)
- pre-VACUUM in vacuum() (hermes_state.py)
- post-optimize-storage (hermes_state_search.py)
PASSIVE never resets/truncates the WAL and never takes the exclusive
checkpoint lock, so it cannot lose a transient closer's race with the
live writer. The WAL is instead bounded by `journal_size_limit` and the
writer's natural post-checkpoint reset. TRUNCATE belongs only on a
sole-opener/quiescent connection (e.g. offline maintenance); this change
does not try to detect that -- PASSIVE is the safe default.
Diagnosed as the root cause of three state.db B-tree corruptions in
2026-08: damage localized to the hottest-written pages (gateway_routing
and the sessions indexes), with whole zero-filled pages still live and
off the freelist -- the checkpoint/reset-race signature, not disk or
application SQL.
Tests: tests/test_wal_checkpoint_strategy.py now asserts PASSIVE at
close(), before vacuum(), and after optimize_fts_storage() VACUUM;
tests/test_hermes_state.py asserts close() likewise. Focused run:
226 passed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review of this PR was right that maxsize=8 bounds the wrong thing. The
LifoQueue caps how many connections are RETURNED; _checkout_read_conn opened
unconditionally on a miss, so N readers arriving on a cold pool all missed, all
opened, and peaked at N. The surplus was closed on release, so nothing
accumulated forever -- but EMFILE is a peak-instant condition and the burst
that empties the pool is exactly the burst that exhausts the fd table, so the
original wedge was still reachable. Measured on the previous commit: 64
concurrent readers held 64 live connections at once.
A connection now holds a permit for its whole lifetime -- acquired in
_get_read_conn() before the open, released in _close_read_conn() after the
close -- so open+checked-out is bounded together. A pool hit costs no permit
because the connection it hands back already holds one, which leaves
_get_read_conn() as the only place that can open. The acquire is non-blocking:
past the ceiling readers fall back to the locked writer connection rather than
queueing, since blocking would convert descriptor exhaustion into a stall,
which is the same outage with a different stack trace. Same burst now peaks at
8. BoundedSemaphore rather than Semaphore so an unpaired release raises instead
of silently widening the ceiling.
Two latent leaks in the same function, found while doing this:
- a CJK extension load that failed after a successful open returned None
without closing the connection, leaking a descriptor the tracking registry
still counted -- the same leak shape one level down;
- any non-sqlite3.Error between open and return stranded a permit
permanently, which would ratchet the ceiling down to zero and silently
demote every later read to the writer lock.
On the test: the existing one joins every worker before counting, so it
measures the pool at rest and structurally cannot observe peak -- which is why
this got through. The new one uses a barrier so all 64 workers hold their
connections until every worker has checked out, making the count taken at that
moment the actual simultaneous peak. Verified it fails against the previous
commit (64 checked out, 65 live) and passes at 8/9. Also covers the
writer-connection fallback, permit recovery after a failed open, and that
close() releases exactly the permits it drained.
`projects.tree` answers for the backend's own profile, so the grouped
sidebar had nothing to draw once the user asked to see every profile.
Run the same authoritative builder once per profile against that
profile's state.db and merge the results by folder, so one checkout is
one group no matter how many profiles work in it, and the owning profile
rides on each session row where the badge and filter can read it.
Group totals are summed in SQL rather than over the loaded page — a
number that shrank as you scrolled would be worse than no number.
Scope the batched sidebar slices while we're here: cron and messaging
came back cross-profile unconditionally, which is why a concrete profile
showed another profile's Telegram threads and cronjobs.
Closes#65710Closes#42651Closes#70629
Review follow-ups on the composite salvage (whole-bug-class sweep):
- session.branch and _persist_branch_seed copied parent history without
display_kind/display_metadata, so a tagged timeline marker (personality
pivot, model switch, auto-continue) re-entered the branched session as a
bare role=user row after a restart — re-planting the phantom-ordinal
class this PR fixes. Both projection dicts now carry the tags; regression
asserts added to both branch tests (mutation-checked: fail without the
fix).
- ui-tui renderer learns display_kind=personality_switch (was falling
through to an opaque user bubble; desktop got the case in commit 1).
- programmatic-integration docs: document the two new 4004 refusals
(boolean ordinal, bare confirm_truncate).
- hermes_state comment: archived rows are searchable only with
include_inactive=True, not by default search — align comment with the
actual FTS filter.
- strip stray trailing blank line in test_tui_gateway_server.py
Guarding the *aim* of a rewind still leaves every other way of aiming it
wrong terminal. All three reported incidents (#70516, #80763, #82756) ended
at the same write — `replace_messages()` in the `prompt.submit` truncation
path — and all three were unrecoverable for the same reason: the rows are
DELETEd, which also evicts them from the FTS index, so there is no `active=0`
archive and nothing to restore from.
The codebase already draws this distinction and already has the safe half of
it. `archive_and_compact` is documented as "the durability-preserving
alternative to replace_messages"; `rewind_to_message` — the `/undo` path —
soft-deletes to `active=0, compacted=0` and keeps the rows "on disk for audit
/ forensic inspection". The desktop rewind is the same user-facing operation
as `/undo` and was the one taking the destructive branch.
`replace_messages(..., archive_dropped=True)` flips the DELETE to a
content-preserving `UPDATE messages SET active = 0`, reusing the existing
transaction and the existing `active=0, compacted=0` marking so the dropped
turns stay readable via `get_messages(..., include_inactive=True)` and stay
out of session search (`compacted=0` = "the user took it back", vs
compaction's `compacted=1` = "summarized away, still discoverable").
The live transcript is byte-identical either way — only the durability of the
dropped turns changes. The parameter defaults to False, so the fork handler,
the ACP adapter and `gateway/session.py` keep their current semantics
untouched; a test pins that.
`active_only=True` stays on the call: #80216 still applies, and archiving must
not disturb rows an earlier compaction deliberately archived.
Test doubles for `replace_messages` in the gateway suite are widened to the
real signature — they are stand-ins for SessionDB, and a double that does not
accept what production passes silently converts this write into a 5008.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
find_latest_gateway_session_for_peer filtered non-recoverable rows out of
candidacy BEFORE ordering, so recovery could search behind a /new reset
boundary and resurrect an older still-open row for the same peer —
silently restoring the exact context the user reset.
Rebuilt against the #82633 finder (has-messages ranking +
COALESCE(last_activity_at, started_at) recency): the fence is expressed
as a NOT EXISTS guard inside both the exact-key and peer-fallback
queries — a candidate is rejected when an intentional boundary row
(session_reset / session_switch / idle / daily / suspended /
resume_pending_expired) for the same peer ended after the candidate's
last activity. If the conversation's most recent event is an intentional
reset, recovery returns nothing rather than reaching behind it.
Cherry-picked from #68617 and adapted to the rewritten finder.
(cherry picked from commit bb2c562a165d91e00f64d42cf7495e6c8a5da9d7)
When state.db's write path fails (corrupt FTS, or a crash landing between
routing publication and row creation), the live gateway conversation can end
up in a session row that never received its identity columns: session_key,
chat_id, chat_type and origin_json are all NULL. In-memory routing hides the
damage for as long as the gateway stays up. After a restart the chat is
resolved from the DB, and find_latest_gateway_session_for_peer cannot see
that row — both of its queries match on the very columns it lacks — so the
chat resumes the last keyed sibling instead, days older. The messages were
never lost, only unreachable.
Hardening the write side cannot reach a row that is already damaged, so add
the offline repair path the tracking issue asks for:
- SessionDB.find_orphaned_gateway_sessions() reports message-bearing rows
with no session_key, and names the predecessor each one continues only
when the evidence is unambiguous — a recorded parent_session_id
("lineage"), or exactly one keyed row of the same source and compatible
user_id that fell quiet within 15 minutes of the orphan's start
("contiguity"). Contested pairs are reported with a reason and left alone:
a wrong adoption would splice one person's conversation into another
person's chat. Branch, delegate and tool rows are excluded — they are
unkeyed by design, not by damage.
- SessionDB.adopt_orphaned_gateway_session() stamps the orphan from the
predecessor (never overwriting a column that already has a value), records
the lineage, and retires the predecessor under end_reason
'superseded_by_repair' — a reason recovery does not treat as resumable, so
the repaired row wins the chat from then on. The pair is re-verified inside
the write transaction, making a concurrent heal a no-op rather than a
conflicting write.
- `hermes sessions repair-routing` drives both. It reports without touching
the database; --apply confirms first and warns that a running gateway
still holds the old mapping in memory.
Refs #82616.
Root cause of #82616: gateway session identity (session_key/chat_id/
origin_json) was written best-effort in a separate UPDATE after row
creation, both reset-path DB writes swallowed failures silently
(logger.debug / bare print), transcript reads ignored the reroute map
that writes follow, and restart recovery ranked candidate rows by
started_at while hard-rejecting empty rows. A single failed write could
therefore strand the live conversation in an unroutable orphan row while
a days-old zombie kept the routing key — after any gateway restart the
chat silently resumed the zombie (user-visible context loss, 5 confirmed
incidents on one install since June).
Four class fixes:
1. Identity lands atomically in the session INSERT: origin_json and
display_name join _insert_session_row's column list + COALESCE
backfill; both gateway creation paths (get_or_create + reset) pass
full identity including parent_session_id lineage (fixes#12857).
2. record_gateway_session_peer self-heals: when the target row is
missing (failed/deferred create, crash window) it INSERTs the row
with full identity instead of silently no-opping — every per-turn
peer refresh is now a repair opportunity, and an identity-less lazy
writer (update_token_counts/record_auxiliary_usage) can never leave
a gateway session permanently unroutable.
3. load_transcript follows the write-side reroute chain and the durable
compression tip before querying, so reads can no longer return 0
rows for a session whose messages live under its compression child;
read exceptions are WARNING, distinguishable from an empty result.
4. find_latest_gateway_session_for_peer ranks by
COALESCE(last_activity_at, started_at) (message-bearing rows first)
and returns an empty-but-keyed row instead of None — a zombie
predecessor can no longer beat the live conversation, and recovery
never mints a fresh id when a keyed row exists.
Reset-path DB write failures now log at WARNING with the routing
consequence spelled out.
Tests: tests/gateway/test_session_continuity_82616.py (11 tests) —
sabotage-verified: 6/11 fail without the fixes. E2E incident replay
(real SessionDB, temp HERMES_HOME) confirms the production shape now
resolves to the live session.
Fixes#82616. Related: #12857, #78182 (read-path half), #79576.
A session row can say whether its work is open, merged or closed, and link
to it. The join is the session's own repo + branch, asked of GitHub in one
batched GraphQL request per repo (branch aliases, not a `gh pr list` page
that a busy repo crowds ours out of), through the remote-aware git facade so
a desktop on a remote gateway asks the backend's `gh`.
Two ways a session's branch can't answer, both covered:
- It ran on trunk. Fork PRs share our branch namespace, so asking about
`main` badges a stranger's PR onto it — trunk is never asked about, and
cross-repository PRs are dropped server-side either way.
- It worked in a worktree, so the branch it recorded at start isn't where
the PR came from. Creating a PR from the review pane binds the session to
the branch it actually used, and for sessions that predate that, the PR is
recovered from the transcript: `gh pr create` prints a bare PR url and
nothing else, so a tool result whose whole output is one is a claim rather
than a mention. Scanned read-only across profiles, once per session ever.
get_messages() only deserializes content and tool_calls; the structured
reasoning columns (reasoning_details, codex_reasoning_items,
codex_message_items) come back as the raw TEXT they were stored as.
Feeding those rows straight back into a write, which is exactly what
the POST /api/sessions/{id}/fork handler does by piping get_messages()
into replace_messages(), hit an unguarded json.dumps() and stored the
already-serialized string encoded a second time. On replay of the fork,
json.loads() then yields the inner string instead of a list, and every
consumer's isinstance(..., list) gate silently drops it: preserved
Anthropic thinking blocks, Codex encrypted-reasoning/message-item
replay, and OpenRouter multi-turn reasoning context are all lost after
a fork, with one more encoding layer added per fork.
The /branch copy loop had the same defect from the other side: it
forwarded reasoning but none of the structured columns, and both TUI
branch writers persisted role/content alone, dropping reasoning and
reasoning_content along with them.
Route the six dumps sites in append_message and _insert_message_rows
through a shared guard that keeps already-serialized strings as-is;
structured values from the live runtime are dumped exactly as before.
Forward the reasoning fields in all three branch writers, matching the
set gateway/slash_commands.py already forwards on its own /branch path.
A session title had no notion of who set it, so two bugs followed. An
auto-generated title could clobber a name the user typed, and every
compression rotation renumbered the conversation it forked - one piece of
work reaching 'Smallville Map Architecture Plan #10' in the sidebar.
Titles now carry a source (derived < llm < user) enforced by one
compare-and-swap, so an automatic write can only ever replace a title of
strictly lower authority. Compression carries the name across unchanged.
Legacy NULL rows rank as user, so auto-titling only fills genuinely
empty titles on existing data.
- Move classify_persistence_error into hermes_state beside is_disk_full_error
and delegate the disk bucket to it (fixes 'ENOSPC writing state.db' and
'not enough space' classifying as unknown). run_agent keeps a thin lazy
delegating wrapper so the documented import path and fast import survive.
- Classify CompressionSessionBusyError (and its RPC-wrapped message forms)
as 'locked': the motivating #81227 failure mode stringifies to 'is being
compressed by another writer', which the substring heuristic missed.
- Export PERSISTENCE_ERROR_CAUSES and iterate it in the cron explainer
suppression instead of a hardcoded tuple, so a future cause bucket cannot
silently desynchronize cron delivery.
- Hedge the gateway locked/unknown recovery wording ('should already be
saved' instead of 'was recorded') to match the explainer - the early
turn-start persist may also have failed.
- Drop STATE_DB_WAL_WARN_BYTES (speculative dead constant with no consumer;
the pre-existing 50 MB doctor WAL check covers the warning).
- Tests: compression-busy classification, is_disk_full_error delegation,
causes-tuple coverage; mutation-checked red-green.
Review follow-ups from the pre-push falsification pass:
- collect_state_db_stats now routes through _connect_tracked_db so the
module's byte-probe guard sees the read-only connection (consistency
with the module's own ro-connection precedent; prevents a raw header
probe from cancelling this reader's locks in multi-threaded callers).
- Drop the new >256 MiB WAL warning from the stats renderer: doctor's
pre-existing 50 MB WAL check (with --fix checkpoint) already covers
WAL runaway, and two warnings for one condition is noise. The test now
locks in the dedup decision.
- Locked-cause explainer says the message 'should already be saved'
rather than overclaiming when the early turn-start persist also failed.
Operators had no Hermes surface showing state.db size, WAL health, index
family shape, or how many processes hold the database — all of which were
needed to diagnose a lock-contention incident on a 4.5 GB multi-writer
install.
Adds collect_state_db_stats() (strictly read-only URI connection, no
SessionDB instantiation, per-field best-effort) and a /proc-based
count_db_holders() to hermes_state, and wires a stats block into hermes
doctor's state.db section: logical size, pages/freelist, WAL size,
message/session counts, journal mode, holder count, FTS table presence
and deferred-rebuild status. Advisory warnings at >1 GiB (suggest
sessions.auto_prune and, when the v23 rebuild is pending or the legacy
trigram shape is detected, an offline 'hermes sessions optimize-storage')
and >256 MiB WAL (checkpoint health). Any stats failure degrades to a
single info line.
SessionResumeTooLargeError said 'across its lineage' even when the CLI
mid-setup path counted only the tip segment; the exception now takes a
scope phrase.
With sessions.max_*_messages: 0 the guards previously still ran an
unbounded COUNT (full lineage for resume) — the exact pathological work
disabling them is meant to avoid. Live callers use the raise side
effect only, so return 0 without touching the messages table.
OFFSET paging made the streaming export O(n^2) on huge transcripts;
after_id keyset paging keeps each page seek O(1). Adds after_id to
SessionDB.get_messages (ascending-only, guarded against latest/offset
combos).
sessions.max_resume_messages / sessions.max_export_messages (default
20000, 0 disables) replace the hardcoded hard-rejects, and the CLI
'sessions export' guard becomes per-session instead of cumulative so
full-DB backups of many small sessions keep working. Error guidance now
points at the config override instead of the (corruption-only) repair
command.
#80216 fixed /retry (and a follow-up fixed yuanbao recall) destroying
soft-archived active=0/compacted=1 in-place-compaction rows via the
destructive replace_messages default. Two sibling sites still carried the
same class:
- acp_adapter/session.py _persist (non-owned-agent branch): probed
has_archived_messages and FAILED OPEN into the destructive full replace
on any probe error; the probe can also race a concurrent
archive_and_compact. Now passes active_only=True unconditionally — on a
fresh create/fork every row is active=1 so behavior is identical, and
the probe (its only production caller) is deleted.
- tui_gateway/methods_prompt.py edit/regenerate truncation: bare
replace_messages() deleted the archived transcript of a compacted
session on every edit/regenerate. Now active_only=True.
hermes_state.has_archived_messages docstring updated (probe is now
test/diagnostic-only). Test stubs in test_tui_gateway_server.py accept the
new kwarg. New regression tests: real-SQLite archive-survival for both
write shapes, fresh-session equivalence (the claim the unconditional
switch rests on), and source-level guards pinning that neither site
re-grows the fail-open probe (both mutation-checked: revert either fix and
its guard fails).
Adversarial review of the salvaged recovery found a reachable fail-open:
compression continuations inherit the rotated agent's model_config
verbatim (publish_compression_child callers pass
agent._session_init_model_config), so a delegate subagent's continuation
carries _delegate_from=<the delegate's own parent>. The marker-PRESENCE
filters in reopen_orphaned_compression_session and
find_live_compression_child misclassified such a REAL continuation as a
delegate child:
- reopen: parent 'orphaned' -> reopened while a live continuation exists
-> two live heads in one lineage (verified with a live repro)
- find_live: adoption misses the continuation (fail-closed, masked the
fork pre-PR; the PR made it active)
Fix: markers only disqualify a child when they point at the queried
parent (shared _NON_CONTINUATION_CHILD_FILTER_SQL fragment, also
resolving the duplicated-SQL drift risk flagged by the reuse reviewer).
Both directions regression-tested: reopen fails closed on an
inherited-marker continuation; find_live adopts it.
Also from review: reopen-failure log raised debug->warning (the failure
hard-fails the turn moments later), commit-semantics hardening comment
on the lease DELETE path, blank-line nit.
The three read-only projection walks (get_compression_tip,
list_sessions_rich chain, resume walk) share the marker-presence shape
but fail closed (skip a continuation -> resume shows the parent), and
the fixed adoption path self-heals that case at turn start; left as-is.