Commit Graph

362 Commits

Author SHA1 Message Date
DannyFengTianYu b066f2b373 fix(history): isolate branch transcripts from parent updates
(cherry picked from commit 57c51cc401e867e0b315e051ba138d0b8f7a5f27)
2026-08-15 00:35:40 -07:00
embwl0x e89532d97e fix(desktop): order async session git metadata 2026-08-15 00:33:11 -07:00
Teknium fbaea9bddc feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable) (#86797)
* feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable)

Adds a source-orthogonal, archive-orthogonal 'hidden' session flag meaning
'don't show in the global Sessions sidebar, but stay fully resumable by the
surface that owns it'. Mirrors the existing archived/pinned capability end to
end, so it's a generic widening (any plugin that owns its own session lifecycle
- kanban, Bot Mode, future plugins - can keep its sessions out of the shared
recents list) rather than a per-plugin special-case.

- Schema: hidden INTEGER NOT NULL DEFAULT 0 on sessions (additive; lands on
  existing DBs via the declarative _reconcile_columns ADD COLUMN path, same as
  archived/pinned - no version-gated migration).
- DB: SessionDB.set_session_hidden(session_id, hidden) (clones set_session_pinned
  incl. the compression-lineage recursive CTE); list_sessions_rich gains
  include_hidden=False, appending 's.hidden = 0' by default so hidden rows drop
  from every listing path (and the REST sidebar endpoints inherit it with no
  change).
- Gateway: session.set_hidden RPC (mirrors session.title); session.create accepts
  hidden=true, deferred via pending_hidden and applied in _ensure_session_db_row
  when the row is lazily created (mirrors pending_title).
- REST parity: PATCH /api/sessions/{id} accepts+bool-validates 'hidden' ->
  set_session_hidden; _session_response exposes it.

Enables Hermes-Bot-Mode to hide canonical 'Bot Chat' sessions from the sidebar
(NousResearch/Hermes-Bot-Mode#46) WITHOUT retagging source (which would mis-set
the agent platform). Bot Chats keep source=desktop. Gateway RPC needs a
SERVE-backend restart to take effect live. 1 focused test (default-exclude /
include_hidden / unhide round-trip).

* fix: teach lost-and-found recovery about the 55-column sessions layout

Adding the 'hidden' column makes the current sessions table 55 columns. The
SQLite lost-and-found recovery classifier keys off the physical field count
(SESSIONS_LAYOUT_NFIELDS) to identify a salvaged sessions row, so a recovered
current-layout row (nfield=55) would otherwise be unrecognized and dropped.
Add 55 to the frozenset (54/52 stay as historical prefixes) and update the
column-count assertions + synthetic current-layout insert in the recovery test.

---------

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-15 00:31:37 -07:00
kshitij 7a390f7c43 fix: __del__ delegates to close() for full cleanup
The original __del__ only closed _conn (the writer connection),
skipping the read-only connection pool, token writer thread, and
atexit unregister. Delegates to self.close() instead so all
cleanup paths run. Uses __dict__.get('_conn') guard to stay
safe on partially-constructed instances and during interpreter
shutdown.
2026-08-15 10:26:24 +05:30
RelaxJonh 1db4801063 fix: close leaked SessionDB connections on exception paths (#83226)
Two call sites create SessionDB instances without closing them on error:

1. gateway/slash_commands.py: /insights command — db.close() was on the
   success path but not in a finally block, so exceptions between
   SessionDB() and db.close() leak the connection.

2. hermes_cli/sessions_cmd.py: sessions repair — SessionDB() created
   inline with no .close() at all, leaking the FD on every call.

Additionally, add a __del__ safety net to SessionDB itself so that
instances orphaned by callers who forget .close() are cleaned up when
garbage collected, rather than pinning FDs alive until process exit
via the atexit hook.

Fixes #83226
2026-08-15 10:26:24 +05:30
joaomarcos 39e480c051 fix(state): close leaked SessionDB connections on exception paths (#83226)
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).

- Close partially initialized SessionDB connections on every constructor
  exception path via a finally block guarded by an initialization-complete
  flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
  CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
  thread, close on worker exit, reject new enqueues after shutdown starts,
  and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
  API disconnect failures, shutdown recovery, RetainDB late enqueue, and
  foreign-loop async clients.

Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 21:41:26 -07:00
Teknium 8dc5608a78 fix(compression): adopt live continuation tip at flush across multi-hop chains
A turn writing against a session already closed by compression died with
session_persistence_failed and a misleading "this is often a full disk"
dialog, even though the store was healthy and a live continuation existed
(#82001). Depth-1 recovery (find_live_compression_child) could not resolve
lineages with >=2 compression hops (root -> mid -> tip), reproduced
independently on two- and three-hop chains.

- run_agent.py flush chokepoint: on CompressionSessionClosedError, resolve
  tip = db.get_compression_tip(old_id) (canonical bounded transitive walk),
  adopt only when tip != old_id AND the tip row is live, retry the flush
  exactly once (adoption budget); otherwise fail closed.
- gateway/session.py append_to_transcript: replace the depth-1 live-child
  lookup with the same tip + liveness contract, so gateway transcript
  reroutes follow full chains.
- agent/conversation_compression.py _adopt_live_compression_child: turn-start
  recovery preflight now resolves via get_compression_tip with the same
  liveness check, closing the last depth-1 consumer in this family.
- classify_persistence_error: new "compression_closed" bucket; the turn-end
  explanation names compression rotation and tells the client to refresh the
  session id instead of blaming a full disk.

Tests: depth-1 adoption, multi-hop chain adoption (agent + gateway), fail
closed with no continuation / stale-closed (ws_orphan_reap) tip, exactly-once
adoption budget, and error-wording guards (compression-closed never mentions
disk; real disk failures keep disk guidance).

Closes #82001

Co-authored-by: Al3xand3r1987 <125030427+Al3xand3r1987@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
2026-08-14 21:39:44 -07:00
zuowen7 bbb6cc7e99 test: cover tool-message dedupe and latest paging for include_compacted (#80680)
Dedupe key now includes tool_call_id/tool_calls/tool_name: compaction
copies carry those fields verbatim, so identical tool messages across
generations still collapse, while distinct tool calls sharing
role/content/timestamp are never merged. Add endpoint-level coverage
for the desktop's real read path (limit + order=latest +
include_compacted=true).
2026-08-14 21:08:14 -07:00
zuowen7 f71f91a39b fix(desktop): surface compaction-archived messages in transcript reads (#80680) 2026-08-14 21:08:14 -07:00
fangliquanflq 19b1204392 fix(sessions): revive uncontested expired turn leases 2026-08-15 03:22:16 +05:30
fangliquan f1025b2c00 fix(sessions): fence transcript writes with the turn-lease holder
Refresh-loss interrupt is cooperative, so a stalled writer could still flush after another process reclaimed the conversation. Carry the holder into append_message / append_messages_batch and reject the write in the same SQLite transaction when the lease row is missing, expired, or owned by someone else.
2026-08-15 03:22:16 +05:30
fangliquan c21efeeb52 fix(sessions): keep turn lease across inherited-marker compressions
Presence-only _delegate_from/_branched_from checks stopped the lease walk on
continuations that copied a delegate's model_config, so the first refresh
after rotation missed the parent-key lease and hard-interrupted. A failed
get_session probe also skipped acquire entirely. Walk the lineage inside
the write transaction and treat a probe error as contended, not a fresh
session.
2026-08-15 03:22:16 +05:30
fangliquan 5e2be43fd4 fix(sessions): harden cross-process turn lease wait and refresh
Honor interrupts while waiting for admission, stop the turn when refresh
loses the lease, poll once per second under contention, and test dead-PID
reclaim.
2026-08-15 03:22:16 +05:30
fangliquan 3b09456019 fix(sessions): surface wait status for cross-process turn leases
Emit lifecycle notices while waiting on another process, and return a
resend-friendly timeout result instead of a bare TimeoutError.
2026-08-15 03:22:16 +05:30
fangliquan 6e929a9694 fix(sessions): serialize turns across processes 2026-08-15 03:22:16 +05:30
teknium1 f79440e0f4 feat: /loop — recurring in-session wakeups (Claude Code parity)
Ports Claude Code's /loop (and its /proactive alias) across every Hermes
surface. /loop [interval] <prompt> re-runs a prompt or slash command on a
recurring cadence inside the live session; omitting the interval enables
self-paced mode (starts at the floor, backs off exponentially while the
agent's replies stop changing, snaps back on change — local digest
comparison, zero extra LLM cost).

Stop conditions: agent-emitted LOOP_COMPLETE marker, --times N,
--until <condition> (judged by the existing goal_judge aux task,
fail-open), /loop stop, and a loops.max_ticks backstop budget.

Core: hermes_cli/loops.py (LoopState + LoopManager + shared
dispatch_loop_command), persisted per session in SessionDB state_meta
(loop:<sid>) so /resume picks it up; migrates across compression
boundaries like /goal. New SessionDB.list_meta_prefix() powers the
gateway's cross-session scan.

Surfaces:
- CLI: /loop handler + idle-fire and post-turn-complete hooks in
  process_loop (mirrors the /goal hook shape; Ctrl+C pauses the loop)
- Gateway: /loop handler with route capture, mid-run control-verb guard,
  post-turn tick completion, and a supervised loop_wakeup_watcher that
  injects due wakeups into idle chats via the synthetic-message path
- TUI/dashboard/desktop: command.dispatch handler + per-session
  notification-poller wakeup driver + post-turn completion in the turn
  dispatcher; /loop added to the desktop slash palette
- /goal mixing: an active non-parked goal owns the idle boundary — loop
  ticks defer until it finishes, pauses, or parks; real user input always
  wins over both

Config: loops.{min_interval_seconds,max_ticks,self_paced_floor_seconds,
self_paced_ceiling_seconds}. Docs page + sidebar entry. 77 new tests.
Slack's 50-slash cap: /version moves to /hermes version to free the
native slot for /loop.
2026-08-14 13:40:19 -07:00
webtecnica 56a41715dc fix: persist provider on model switch and add billing_provider fallback
Salvage of #79604 (webtecnica) + #85721 (pierrenode), combined and
rebased onto current main with simplify-code findings folded in.

#79604: update_session_model() wrote the model name to sessions.model
but never persisted the provider into model_config. On resume, the
runtime recombined the persisted model with the config.yaml primary
provider (which may not serve that model), producing auth errors.
Fix: add optional provider parameter to update_session_model, merged
into model_config via the shared _merge_model_config_json helper (not
hand-rolled SQL). Wire both gateway /model call sites to pass
result.target_provider.

#85721: session_gateway_runtime() had no billing_provider fallback.
A CLI session that never ran /model has no gateway_runtime or
top-level provider in model_config — billing_provider (written on
every session's first accounted API call) is the only durable record.
Fix: add billing_provider as the last-resort fallback in
session_gateway_runtime(), filtering bare billing buckets (auto/custom)
that are not routable identities.

Simplify-code findings addressed:
- Use _merge_model_config_json instead of 40 lines of branched SQL
- Share _BARE_BILLING_PROVIDERS from hermes_state.py (was duplicated
  as a set in tui_gateway/server.py)
- Merge None-filtering from #85920 with the billing_provider fallback
  into one coherent return path

Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
2026-08-14 14:39:08 +05:30
kshitij 9cb456a9b9 fix: unify route dict or-None discipline in /model persist
The route dict in _persist_model_switch_to_session used  filtering
(omits falsy values) while the top-level keys used  (writes
explicit None to trigger deletion in _merge_model_config_json). This
asymmetry meant stale keys from a previous /model switch survived in
the nested gateway_runtime dict even after the fix in #85261 that
properly deleted them from the top-level keys.

Fix: build the route dict with  and derive the top-level keys
from **route so both shapes always use identical deletion semantics.
Also filter None values in session_gateway_runtime's reader since
gateway_runtime is replaced as a whole dict (not deep-merged), so None
values written by the persist path survive in the nested dict.

Found by /simplify-code 3-reviewer review on #85261 (all 3 reviewers
converged on the route dict asymmetry as the verdict-relevant finding).
2026-08-14 13:18:53 +05:30
kshitij b8f85f18e8 refactor: shared /model persist helper, canonical gateway_runtime reader, cross-surface route persistence, tests
- Extract the two duplicated /model session-persist blocks into
  _persist_model_switch_to_session; persist the route BOTH nested
  (gateway_runtime, CLI reader) and top-level (TUI gateway's
  _stored_session_runtime_overrides reader) so a CLI switch also
  survives a desktop/TUI session.resume.
- Add SessionDB.session_gateway_runtime as the canonical tolerant
  row-level route reader (session_yolo_enabled precedent); use it in
  _restore_session_model instead of hand-rolled JSON parsing.
- Clear stale launch-time _explicit_api_key/_explicit_base_url when
  resume restores a different provider (same leak guard
  _apply_model_switch_result already has).
- 12 new tests incl. a real-SessionDB round trip; mutation-checked.
2026-08-14 02:15:22 +05:30
kshitij 91c7a67f44 fix(sessions): close the session_switch legacy gap and fence the resume walker at reset boundaries
Follow-ups to the salvaged #84009 commits:

- Add 'session_switch' to _RESET_END_REASONS: a reset continuation's
  parent can be promoted to session_switch (resume the reset parent,
  then switch away), which permanently hid pre-marker legacy children —
  reopen-time stamping cannot rescue them because the parent is being
  ended, not reopened. Probe-verified before/after.
- Share the legacy reset-child heuristic via _legacy_reset_child_sql()
  so _RESET_CHILD_SQL and reopen_session()'s stamping UPDATE cannot
  drift, and derive find_latest_gateway_session_for_peer's two recovery
  fence literals from _RESET_END_REASONS_SQL (was a third hand-written
  copy of the same set).
- Exclude reset children (marker + legacy shape) from the
  resolve_resume_session_id forward walker: resuming a reset parent
  could redirect into the post-reset conversation — the exact context
  the user reset away. Regression tests cover both shapes plus the
  walker's original compression-tip behavior; mutation-checked.
2026-08-13 23:45:21 +05:30
embwl0x 5a10537b24 fix(sessions): stabilize legacy reset lineage on resume 2026-08-13 23:45:21 +05:30
embwl0x ce89afa59c fix(sessions): keep reset conversations listable 2026-08-13 23:45:21 +05:30
StanleyStetson 23da6d6fe2 fix(gateway/desktop): durable row-id addressing for rewind truncation
Address rewinds/edits via SQLite messages.id (truncate_before_row_id)
instead of shifting user ordinals. Resolve against in-memory stamps,
then durable session history when live turns drop _row_id; refuse
unknown durable targets with 4018 (no ordinal fallback) and 4030 on
ordinal/row_id mismatch. Stamp _row_id on insert, load row ids on
resume paths, send rowId from Desktop, filter renderer-synthetic ids,
and stop silently resending failed targeted edits without truncation.
Add production-shaped SessionDB tests for resolve and fail-closed paths.

Fixes #82959
2026-08-13 13:35:55 +05:30
zhouou6 66d7a39ea6 fix(state): self-heal 'file is not a database' write connections + retry transient EIO on journal-mode probe
Salvaged remainder of PR #82280 (state.db hardening rollup):

- Runtime connection corruption: a sibling process replacing/truncating
  the backing file breaks the live write connection — every subsequent
  write raises 'file is not a database' and the gateway wedges
  permanently (messages pile up in memory). Add a bounded one-shot
  reconnect on the write path: close the broken connection, reopen the
  DB file (re-running WAL activation + schema reconciliation), retry
  the failed write once.
- _on_disk_journal_mode: retry transient 'disk i/o error' (virtualized
  block devices) a few times before returning None, so a one-shot EIO
  doesn't push callers onto the fail-closed unknown-mode branch.

The rollup's write-lock machinery, checkpoint-strategy changes, and
repair serialization are intentionally NOT included — superseded by
PRs #84277 and #69609, or wrong-direction per the POSIX
lock-cancellation findings (#71724 lineage).
2026-08-12 19:43:41 -07:00
Aldo 9cf4fe5513 fix(state): bound WAL growth and checkpoint after VACUUM
`sessions optimize` could consume several GB of disk instead of freeing
any, filling the host to 100% on exactly the large databases it exists to
shrink.

Two causes, both in the WAL lifecycle:

1. No `journal_size_limit`. SQLite defaults to -1 (unlimited), so after a
   checkpoint the WAL is reused in place and never truncated —
   `state.db-wal` permanently keeps the high-water mark of the largest
   transaction ever run. `hermes_cli/kanban_db.py` already bounds its WAL
   with `wal_autocheckpoint=100`; the session store, by far the larger
   database, had no equivalent.

2. `vacuum()` checkpoints BEFORE `VACUUM` but not after. VACUUM rewrites
   every page through the WAL, so the pre-checkpoint does nothing about
   the slack VACUUM itself creates.

Measured on a 3.0 GB state.db: `hermes sessions optimize` reported
"3143.9 MB -> 3155.1 MB (reclaimed -11.2 MB)" while leaving a 3.07 GB
state.db-wal behind. Free space fell from 6.9 GB to 772 MB (100% full)
and stayed there. A manual `PRAGMA wal_checkpoint(TRUNCATE)` recovered
the full 3.07 GB, confirming it was slack, not data.

Fix: set `journal_size_limit` (64 MiB) when enabling WAL, and truncate
the WAL again after VACUUM. Both are best-effort and never raise — a
failure costs disk slack and must not stop the DB from opening.

Tests assert the contract (limit is a finite positive bound; VACUUM does
not leave an oversized WAL) rather than pinning the byte count, which is
a tunable. They skip where WAL is unavailable — including hosts where
Hermes falls back to journal_mode=DELETE due to the SQLite 3.50.4
WAL-reset bug.

Verified: 462 passed / 3 skipped in tests/test_hermes_state.py, and
_apply_wal_size_limit flips a real WAL database from -1 to 67108864.
Tested on Linux (aarch64, Python 3.11).
2026-08-12 19:43:35 -07:00
Teknium d724bd0376 fix(state): make a refused pre-repair backup a hard stop (#69603)
The Aug 2026 incident in #69603 documented a fail-open: when the
pre-repair backup was refused (another same-process handle open),
repair_state_db_schema() recorded backup_path=None and proceeded —
leaving the writable_schema surgery, FTS-schema deletion, REINDEX and
VACUUM strategies reachable against the only remaining copy of the
damaged DB.

_backup_db_file() now returns (path, reason) and the repair path treats
any refused/failed backup as an unconditional hard stop: abort before
the first mutating strategy and surface the reason in report['error'].
Explicit backup=False (CLI --no-backup) is unchanged — that is the
operator opting out, not a silent failure.

Three new tests: refusal hard-stops with source bytes untouched,
OS-level copy failure hard-stops with the reason surfaced, and
backup=False still repairs.
2026-08-12 19:43:29 -07:00
Ernst Bablick 923d86e099 fix(state): serialize state.db schema surgery across processes
`repair_state_db_schema()` performs `PRAGMA writable_schema=ON` +
`sqlite_master` surgery + `VACUUM` on a private connection. The only guard
around it is `_repair_attempt_lock`, a `threading.Lock`, whose docstring
claims it "serialises concurrent web_server / gateway opens" — but a
threading lock covers threads inside one interpreter, not processes.

A normal host runs four independent processes against the same state.db:
the gateway service, the Desktop app's own `hermes serve` backend (it
spawns one per launch, not a thin client), interactive CLI sessions, and
the TUI slash worker. When two of them hit a malformed DB, both entered
the critical section and each ran the full surgery while the other was
mid-rewrite. Observed as a repair/re-corrupt cascade: the DB is repaired,
then re-corrupts minutes later, repeatedly.

Two fixes:

1. Wrap the surgery in a bounded `flock` on `<db>.repair.lock`. `flock` is
   the right primitive — the kernel drops it when the holder dies, so a
   crashed repairer cannot wedge future repairs the way a pidfile would.
   The acquire is bounded (#36644's failure shape) and, unlike the kanban
   init lock, a caller that times out must NOT proceed: here "proceed
   anyway" is exactly the unsafe interleaving. It re-probes instead, and
   reports success if the holder already healed the file.

   Under the lock, the existing `_db_opens_cleanly()` check becomes a
   double-check: a queued process finds the DB healthy and returns
   `already_healthy` rather than re-running surgery on a repaired DB.

2. Bump the schema cookie after direct `sqlite_master` edits. Ordinary DDL
   bumps it for free and every other connection compares it before running
   a prepared statement — that is how they learn to drop a cached schema.
   Editing `sqlite_master` under `writable_schema=ON` does not, so live
   connections in other processes kept writing `messages` rows through
   triggers into `messages_fts*` shadow tables the surgery had just
   deleted. SQLite's writable_schema docs call out incrementing
   `schema_version` as the required companion to such an edit.

Tests: four new cases in tests/test_state_db_malformed_repair.py, all
using real child processes and a real flock. All four fail on main and
pass with this change; the concurrency case asserts exactly one
`malformed-backup-*` file is produced by two simultaneous repairers
(two on main). Full state suite: 558 passed.

Complements #43742, which makes the *in-process* claim loser retry rather
than raise; it explicitly leaves `repair_state_db_schema()` unchanged and
does nothing cross-process. The two are independent and compose.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-12 19:43:29 -07:00
lkz-de ba80f3b86d fix(state): PASSIVE not TRUNCATE for all state.db checkpoints (#45383)
SessionDB.close() ran `PRAGMA wal_checkpoint(TRUNCATE)`. Every cron
run_agent opens and closes its own transient SessionDB, so on a busy
fleet this fired a full WAL reset many times an hour, racing the
gateway's long-lived writer on a large WAL database and tearing hot
B-tree pages -- structurally the same corruption this module's own
periodic checkpoint was already switched to PASSIVE to avoid (#45383).
Only close() and two manual-maintenance paths still used TRUNCATE.

Route every checkpoint on the shared state.db through PASSIVE:
  - close()                    (hermes_state.py)
  - pre-VACUUM in vacuum()     (hermes_state.py)
  - post-optimize-storage      (hermes_state_search.py)

PASSIVE never resets/truncates the WAL and never takes the exclusive
checkpoint lock, so it cannot lose a transient closer's race with the
live writer. The WAL is instead bounded by `journal_size_limit` and the
writer's natural post-checkpoint reset. TRUNCATE belongs only on a
sole-opener/quiescent connection (e.g. offline maintenance); this change
does not try to detect that -- PASSIVE is the safe default.

Diagnosed as the root cause of three state.db B-tree corruptions in
2026-08: damage localized to the hottest-written pages (gateway_routing
and the sessions indexes), with whole zero-filled pages still live and
off the freelist -- the checkpoint/reset-race signature, not disk or
application SQL.

Tests: tests/test_wal_checkpoint_strategy.py now asserts PASSIVE at
close(), before vacuum(), and after optimize_fts_storage() VACUUM;
tests/test_hermes_state.py asserts close() likewise. Focused run:
226 passed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-12 19:43:22 -07:00
Yishova 0472c31aa1 state: bound PEAK read connections with a permit, not just pooled returns
Review of this PR was right that maxsize=8 bounds the wrong thing. The
LifoQueue caps how many connections are RETURNED; _checkout_read_conn opened
unconditionally on a miss, so N readers arriving on a cold pool all missed, all
opened, and peaked at N. The surplus was closed on release, so nothing
accumulated forever -- but EMFILE is a peak-instant condition and the burst
that empties the pool is exactly the burst that exhausts the fd table, so the
original wedge was still reachable. Measured on the previous commit: 64
concurrent readers held 64 live connections at once.

A connection now holds a permit for its whole lifetime -- acquired in
_get_read_conn() before the open, released in _close_read_conn() after the
close -- so open+checked-out is bounded together. A pool hit costs no permit
because the connection it hands back already holds one, which leaves
_get_read_conn() as the only place that can open. The acquire is non-blocking:
past the ceiling readers fall back to the locked writer connection rather than
queueing, since blocking would convert descriptor exhaustion into a stall,
which is the same outage with a different stack trace. Same burst now peaks at
8. BoundedSemaphore rather than Semaphore so an unpaired release raises instead
of silently widening the ceiling.

Two latent leaks in the same function, found while doing this:

  - a CJK extension load that failed after a successful open returned None
    without closing the connection, leaking a descriptor the tracking registry
    still counted -- the same leak shape one level down;
  - any non-sqlite3.Error between open and return stranded a permit
    permanently, which would ratchet the ceiling down to zero and silently
    demote every later read to the writer lock.

On the test: the existing one joins every worker before counting, so it
measures the pool at rest and structurally cannot observe peak -- which is why
this got through. The new one uses a barrier so all 64 workers hold their
connections until every worker has checked out, making the count taken at that
moment the actual simultaneous peak. Verified it fails against the previous
commit (64 checked out, 65 live) and passes at 8/9. Also covers the
writer-connection fallback, permit recovery after a failed open, and that
close() releases exactly the permits it drained.
2026-08-10 17:02:56 -07:00
Yishova 87aedbe7b6 state: pool SessionDB read connections instead of leaking one per (SessionDB x thread) 2026-08-10 17:02:56 -07:00
Brooklyn Nicholson 5b68d2271b feat(profiles): serve a cross-profile project tree and per-profile usage totals
`projects.tree` answers for the backend's own profile, so the grouped
sidebar had nothing to draw once the user asked to see every profile.
Run the same authoritative builder once per profile against that
profile's state.db and merge the results by folder, so one checkout is
one group no matter how many profiles work in it, and the owning profile
rides on each session row where the badge and filter can read it.

Group totals are summed in SQL rather than over the loaded page — a
number that shrank as you scrolled would be worse than no number.

Scope the batched sidebar slices while we're here: cron and messaging
came back cross-profile unconditionally, which is why a concrete profile
showed another profile's Telegram threads and cronjobs.

Closes #65710
Closes #42651
Closes #70629
2026-08-10 03:13:08 -05:00
kshitij 327f7efab8 fix: close sibling display_kind drops and ui-tui parity for #82756
Review follow-ups on the composite salvage (whole-bug-class sweep):

- session.branch and _persist_branch_seed copied parent history without
  display_kind/display_metadata, so a tagged timeline marker (personality
  pivot, model switch, auto-continue) re-entered the branched session as a
  bare role=user row after a restart — re-planting the phantom-ordinal
  class this PR fixes. Both projection dicts now carry the tags; regression
  asserts added to both branch tests (mutation-checked: fail without the
  fix).
- ui-tui renderer learns display_kind=personality_switch (was falling
  through to an opaque user bubble; desktop got the case in commit 1).
- programmatic-integration docs: document the two new 4004 refusals
  (boolean ordinal, bare confirm_truncate).
- hermes_state comment: archived rows are searchable only with
  include_inactive=True, not by default search — align comment with the
  actual FTS filter.
- strip stray trailing blank line in test_tui_gateway_server.py
2026-08-10 11:01:15 +05:30
joaomarcos 60645f8a53 fix(state): make a rewind truncation recoverable instead of a hard DELETE (#82756)
Guarding the *aim* of a rewind still leaves every other way of aiming it
wrong terminal. All three reported incidents (#70516, #80763, #82756) ended
at the same write — `replace_messages()` in the `prompt.submit` truncation
path — and all three were unrecoverable for the same reason: the rows are
DELETEd, which also evicts them from the FTS index, so there is no `active=0`
archive and nothing to restore from.

The codebase already draws this distinction and already has the safe half of
it. `archive_and_compact` is documented as "the durability-preserving
alternative to replace_messages"; `rewind_to_message` — the `/undo` path —
soft-deletes to `active=0, compacted=0` and keeps the rows "on disk for audit
/ forensic inspection". The desktop rewind is the same user-facing operation
as `/undo` and was the one taking the destructive branch.

`replace_messages(..., archive_dropped=True)` flips the DELETE to a
content-preserving `UPDATE messages SET active = 0`, reusing the existing
transaction and the existing `active=0, compacted=0` marking so the dropped
turns stay readable via `get_messages(..., include_inactive=True)` and stay
out of session search (`compacted=0` = "the user took it back", vs
compaction's `compacted=1` = "summarized away, still discoverable").

The live transcript is byte-identical either way — only the durability of the
dropped turns changes. The parameter defaults to False, so the fork handler,
the ACP adapter and `gateway/session.py` keep their current semantics
untouched; a test pins that.

`active_only=True` stays on the call: #80216 still applies, and archiving must
not disturb rows an earlier compaction deliberately archived.

Test doubles for `replace_messages` in the gateway suite are widened to the
real signature — they are stand-ins for SessionDB, and a double that does not
accept what production passes silently converts this write into a 5008.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 11:01:15 +05:30
Tranquil-Flow 6e99531c8e fix(gateway): respect reset boundaries during recovery (#68539)
find_latest_gateway_session_for_peer filtered non-recoverable rows out of
candidacy BEFORE ordering, so recovery could search behind a /new reset
boundary and resurrect an older still-open row for the same peer —
silently restoring the exact context the user reset.

Rebuilt against the #82633 finder (has-messages ranking +
COALESCE(last_activity_at, started_at) recency): the fence is expressed
as a NOT EXISTS guard inside both the exact-key and peer-fallback
queries — a candidate is rejected when an intentional boundary row
(session_reset / session_switch / idle / daily / suspended /
resume_pending_expired) for the same peer ended after the candidate's
last activity. If the conversation's most recent event is an intentional
reset, recovery returns nothing rather than reaching behind it.

Cherry-picked from #68617 and adapted to the rewritten finder.
(cherry picked from commit bb2c562a165d91e00f64d42cf7495e6c8a5da9d7)
2026-08-09 15:02:27 -07:00
izumi0uu 1527a81b5e fix(state): keep canonical writes available when FTS is corrupt 2026-08-09 14:10:21 -07:00
joaomarcos c790ed2a5d fix(state): recover gateway sessions stranded without a routing identity
When state.db's write path fails (corrupt FTS, or a crash landing between
routing publication and row creation), the live gateway conversation can end
up in a session row that never received its identity columns: session_key,
chat_id, chat_type and origin_json are all NULL. In-memory routing hides the
damage for as long as the gateway stays up. After a restart the chat is
resolved from the DB, and find_latest_gateway_session_for_peer cannot see
that row — both of its queries match on the very columns it lacks — so the
chat resumes the last keyed sibling instead, days older. The messages were
never lost, only unreachable.

Hardening the write side cannot reach a row that is already damaged, so add
the offline repair path the tracking issue asks for:

- SessionDB.find_orphaned_gateway_sessions() reports message-bearing rows
  with no session_key, and names the predecessor each one continues only
  when the evidence is unambiguous — a recorded parent_session_id
  ("lineage"), or exactly one keyed row of the same source and compatible
  user_id that fell quiet within 15 minutes of the orphan's start
  ("contiguity"). Contested pairs are reported with a reason and left alone:
  a wrong adoption would splice one person's conversation into another
  person's chat. Branch, delegate and tool rows are excluded — they are
  unkeyed by design, not by damage.
- SessionDB.adopt_orphaned_gateway_session() stamps the orphan from the
  predecessor (never overwriting a column that already has a value), records
  the lineage, and retires the predecessor under end_reason
  'superseded_by_repair' — a reason recovery does not treat as resumable, so
  the repaired row wins the chat from then on. The pair is re-verified inside
  the write transaction, making a concurrent heal a no-op rather than a
  conflicting write.
- `hermes sessions repair-routing` drives both. It reports without touching
  the database; --apply confirms first and warns that a running gateway
  still holds the old mapping in memory.

Refs #82616.
2026-08-09 14:06:06 -07:00
Teknium d2a4d373eb fix(gateway): make session identity durable so chat continuity survives crashes and restarts
Root cause of #82616: gateway session identity (session_key/chat_id/
origin_json) was written best-effort in a separate UPDATE after row
creation, both reset-path DB writes swallowed failures silently
(logger.debug / bare print), transcript reads ignored the reroute map
that writes follow, and restart recovery ranked candidate rows by
started_at while hard-rejecting empty rows. A single failed write could
therefore strand the live conversation in an unroutable orphan row while
a days-old zombie kept the routing key — after any gateway restart the
chat silently resumed the zombie (user-visible context loss, 5 confirmed
incidents on one install since June).

Four class fixes:

1. Identity lands atomically in the session INSERT: origin_json and
   display_name join _insert_session_row's column list + COALESCE
   backfill; both gateway creation paths (get_or_create + reset) pass
   full identity including parent_session_id lineage (fixes #12857).

2. record_gateway_session_peer self-heals: when the target row is
   missing (failed/deferred create, crash window) it INSERTs the row
   with full identity instead of silently no-opping — every per-turn
   peer refresh is now a repair opportunity, and an identity-less lazy
   writer (update_token_counts/record_auxiliary_usage) can never leave
   a gateway session permanently unroutable.

3. load_transcript follows the write-side reroute chain and the durable
   compression tip before querying, so reads can no longer return 0
   rows for a session whose messages live under its compression child;
   read exceptions are WARNING, distinguishable from an empty result.

4. find_latest_gateway_session_for_peer ranks by
   COALESCE(last_activity_at, started_at) (message-bearing rows first)
   and returns an empty-but-keyed row instead of None — a zombie
   predecessor can no longer beat the live conversation, and recovery
   never mints a fresh id when a keyed row exists.

Reset-path DB write failures now log at WARNING with the routing
consequence spelled out.

Tests: tests/gateway/test_session_continuity_82616.py (11 tests) —
sabotage-verified: 6/11 fail without the fixes. E2E incident replay
(real SessionDB, temp HERMES_HOME) confirms the production shape now
resolves to the live session.

Fixes #82616. Related: #12857, #78182 (read-path half), #79576.
2026-08-09 12:37:06 -07:00
Brooklyn Nicholson 21aaa8b4f8 feat(desktop): resolve a session's pull request
A session row can say whether its work is open, merged or closed, and link
to it. The join is the session's own repo + branch, asked of GitHub in one
batched GraphQL request per repo (branch aliases, not a `gh pr list` page
that a busy repo crowds ours out of), through the remote-aware git facade so
a desktop on a remote gateway asks the backend's `gh`.

Two ways a session's branch can't answer, both covered:

- It ran on trunk. Fork PRs share our branch namespace, so asking about
  `main` badges a stranger's PR onto it — trunk is never asked about, and
  cross-repository PRs are dropped server-side either way.
- It worked in a worktree, so the branch it recorded at start isn't where
  the PR came from. Creating a PR from the review pane binds the session to
  the branch it actually used, and for sessions that predate that, the PR is
  recovered from the transcript: `gh pr create` prints a bare PR url and
  nothing else, so a tool result whose whole output is one is a claim rather
  than a mention. Scanned read-only across profiles, once per session ever.
2026-08-09 06:21:04 -05:00
cryptoyasenka f2d03c1f2a fix(state,cli,tui-gateway): keep reasoning fields intact across forks and branches
get_messages() only deserializes content and tool_calls; the structured
reasoning columns (reasoning_details, codex_reasoning_items,
codex_message_items) come back as the raw TEXT they were stored as.
Feeding those rows straight back into a write, which is exactly what
the POST /api/sessions/{id}/fork handler does by piping get_messages()
into replace_messages(), hit an unguarded json.dumps() and stored the
already-serialized string encoded a second time. On replay of the fork,
json.loads() then yields the inner string instead of a list, and every
consumer's isinstance(..., list) gate silently drops it: preserved
Anthropic thinking blocks, Codex encrypted-reasoning/message-item
replay, and OpenRouter multi-turn reasoning context are all lost after
a fork, with one more encoding layer added per fork.

The /branch copy loop had the same defect from the other side: it
forwarded reasoning but none of the structured columns, and both TUI
branch writers persisted role/content alone, dropping reasoning and
reasoning_content along with them.

Route the six dumps sites in append_message and _insert_message_rows
through a shared guard that keeps already-serialized strings as-is;
structured values from the live runtime are dumped exactly as before.
Forward the reasoning fields in all three branch writers, matching the
set gateway/slash_commands.py already forwards on its own /branch path.
2026-08-08 17:37:26 -07:00
Brooklyn Nicholson 5566379f57 fix(sessions): give titles provenance so they stop overwriting themselves
A session title had no notion of who set it, so two bugs followed. An
auto-generated title could clobber a name the user typed, and every
compression rotation renumbered the conversation it forked - one piece of
work reaching 'Smallville Map Architecture Plan #10' in the sidebar.

Titles now carry a source (derived < llm < user) enforced by one
compare-and-swap, so an automatic write can only ever replace a title of
strictly lower authority. Compression carries the name across unchanged.
Legacy NULL rows rank as user, so auto-titling only fills genuinely
empty titles on existing data.
2026-08-08 17:07:21 -05:00
kshitij 1005a057f0 review follow-ups: canonical classifier in hermes_state, compression-busy=locked, hedged gateway wording, drop dead constant
- Move classify_persistence_error into hermes_state beside is_disk_full_error
  and delegate the disk bucket to it (fixes 'ENOSPC writing state.db' and
  'not enough space' classifying as unknown). run_agent keeps a thin lazy
  delegating wrapper so the documented import path and fast import survive.
- Classify CompressionSessionBusyError (and its RPC-wrapped message forms)
  as 'locked': the motivating #81227 failure mode stringifies to 'is being
  compressed by another writer', which the substring heuristic missed.
- Export PERSISTENCE_ERROR_CAUSES and iterate it in the cron explainer
  suppression instead of a hardcoded tuple, so a future cause bucket cannot
  silently desynchronize cron delivery.
- Hedge the gateway locked/unknown recovery wording ('should already be
  saved' instead of 'was recorded') to match the explainer - the early
  turn-start persist may also have failed.
- Drop STATE_DB_WAL_WARN_BYTES (speculative dead constant with no consumer;
  the pre-existing 50 MB doctor WAL check covers the warning).
- Tests: compression-busy classification, is_disk_full_error delegation,
  causes-tuple coverage; mutation-checked red-green.
2026-08-08 14:18:26 +05:30
Victor Kyriazakos a24cbaf426 review: tracked ro-connection for stats, single WAL warning, hedged locked-cause wording
Review follow-ups from the pre-push falsification pass:

- collect_state_db_stats now routes through _connect_tracked_db so the
  module's byte-probe guard sees the read-only connection (consistency
  with the module's own ro-connection precedent; prevents a raw header
  probe from cancelling this reader's locks in multi-threaded callers).
- Drop the new >256 MiB WAL warning from the stats renderer: doctor's
  pre-existing 50 MB WAL check (with --fix checkpoint) already covers
  WAL runaway, and two warnings for one condition is noise. The test now
  locks in the dedup decision.
- Locked-cause explainer says the message 'should already be saved'
  rather than overclaiming when the early turn-start persist also failed.
2026-08-08 14:18:26 +05:30
Victor Kyriazakos 64c342c1c9 feat(doctor): state.db health stats — size, WAL, FTS shape, holders, growth warnings
Operators had no Hermes surface showing state.db size, WAL health, index
family shape, or how many processes hold the database — all of which were
needed to diagnose a lock-contention incident on a 4.5 GB multi-writer
install.

Adds collect_state_db_stats() (strictly read-only URI connection, no
SessionDB instantiation, per-field best-effort) and a /proc-based
count_db_holders() to hermes_state, and wires a stats block into hermes
doctor's state.db section: logical size, pages/freelist, WAL size,
message/session counts, journal mode, holder count, FTS table presence
and deferred-rebuild status. Advisory warnings at >1 GiB (suggest
sessions.auto_prune and, when the v23 rebuild is pending or the legacy
trigram shape is detected, an offline 'hermes sessions optimize-storage')
and >256 MiB WAL (checkpoint health). Any stats failure degrades to a
single info line.
2026-08-08 14:18:26 +05:30
kshitij a8ccd52123 refactor(sessions): accurate scope wording for tip-only resume rejections
SessionResumeTooLargeError said 'across its lineage' even when the CLI
mid-setup path counted only the tip segment; the exception now takes a
scope phrase.
2026-08-08 13:36:08 +05:30
kshitij edf2cb4bf8 perf(sessions): skip counting entirely when transcript guards are disabled
With sessions.max_*_messages: 0 the guards previously still ran an
unbounded COUNT (full lineage for resume) — the exact pathological work
disabling them is meant to avoid. Live callers use the raise side
effect only, so return 0 without touching the messages table.
2026-08-08 13:36:08 +05:30
kshitij e8b05dc6c2 perf(dashboard): keyset pagination for streaming session export
OFFSET paging made the streaming export O(n^2) on huge transcripts;
after_id keyset paging keeps each page seek O(1). Adds after_id to
SessionDB.get_messages (ascending-only, guarded against latest/offset
combos).
2026-08-08 13:36:08 +05:30
kshitij f0794640f6 feat(sessions): config-gate transcript safety limits
sessions.max_resume_messages / sessions.max_export_messages (default
20000, 0 disables) replace the hardcoded hard-rejects, and the CLI
'sessions export' guard becomes per-session instead of cumulative so
full-DB backups of many small sessions keep working. Error guidance now
points at the config override instead of the (corruption-only) repair
command.
2026-08-08 13:36:08 +05:30
kinsolee c750d5354a fix(sessions): prevent oversized transcripts from exhausting memory 2026-08-08 13:36:08 +05:30
kshitij ee6d79648a fix(state): finish the #80216 bug class — archive-preserving rewrites at the two remaining sibling sites
#80216 fixed /retry (and a follow-up fixed yuanbao recall) destroying
soft-archived active=0/compacted=1 in-place-compaction rows via the
destructive replace_messages default. Two sibling sites still carried the
same class:

- acp_adapter/session.py _persist (non-owned-agent branch): probed
  has_archived_messages and FAILED OPEN into the destructive full replace
  on any probe error; the probe can also race a concurrent
  archive_and_compact. Now passes active_only=True unconditionally — on a
  fresh create/fork every row is active=1 so behavior is identical, and
  the probe (its only production caller) is deleted.
- tui_gateway/methods_prompt.py edit/regenerate truncation: bare
  replace_messages() deleted the archived transcript of a compacted
  session on every edit/regenerate. Now active_only=True.

hermes_state.has_archived_messages docstring updated (probe is now
test/diagnostic-only). Test stubs in test_tui_gateway_server.py accept the
new kwarg. New regression tests: real-SQLite archive-survival for both
write shapes, fresh-session equivalence (the claim the unconditional
switch rests on), and source-level guards pinning that neither site
re-grows the fail-open probe (both mutation-checked: revert either fix and
its guard fails).
2026-08-07 14:42:32 +05:30
kshitij a0801b878a fix: bind continuation-marker exclusions to the queried parent (fail-open fix)
Adversarial review of the salvaged recovery found a reachable fail-open:
compression continuations inherit the rotated agent's model_config
verbatim (publish_compression_child callers pass
agent._session_init_model_config), so a delegate subagent's continuation
carries _delegate_from=<the delegate's own parent>. The marker-PRESENCE
filters in reopen_orphaned_compression_session and
find_live_compression_child misclassified such a REAL continuation as a
delegate child:

- reopen: parent 'orphaned' -> reopened while a live continuation exists
  -> two live heads in one lineage (verified with a live repro)
- find_live: adoption misses the continuation (fail-closed, masked the
  fork pre-PR; the PR made it active)

Fix: markers only disqualify a child when they point at the queried
parent (shared _NON_CONTINUATION_CHILD_FILTER_SQL fragment, also
resolving the duplicated-SQL drift risk flagged by the reuse reviewer).
Both directions regression-tested: reopen fails closed on an
inherited-marker continuation; find_live adopts it.

Also from review: reopen-failure log raised debug->warning (the failure
hard-fails the turn moments later), commit-semantics hardening comment
on the lease DELETE path, blank-line nit.

The three read-only projection walks (get_compression_tip,
list_sessions_rich chain, resume walk) share the marker-presence shape
but fail closed (skip a continuation -> resume shows the parent), and
the fixed adoption path self-heals that case at turn start; left as-is.
2026-08-07 13:24:56 +05:30