Commit Graph

66 Commits

Author SHA1 Message Date
Brooklyn Nicholson fdd399dfd3 test(sessions): pin display-projection parity across compaction
Assert the invariant — every display read of one session returns the same
transcript — plus the two things that must not grow with it: the model-fed
projection stays compressed, and Undo/Rewind rows stay hidden. Six of these
fail on the unfixed read paths.
2026-09-02 12:11:15 -05:00
kshitijk4poor e9fa7bc05e fix(state): keep the #94736 teardown self-heal alive under the deleted-WAL guard
Follow-ups on the salvaged #101081 guard:

- A clean close() lets SQLite unlink the WAL sidecars legitimately; the
  guard treated that as a lost generation and permanently halted the
  handle, so the #94736 late-write self-heal reopen dropped transcript
  tails (4 existing tests failed). close() now clears the recorded
  sidecar generation, and _wal_generation_was_lost() re-adopts the
  current sidecars after a clean /proc/self probe instead of relying on
  a stale snapshot.
- Healthy writes no longer walk /proc/self/fd: once a sidecar
  generation is recorded, the stat-based inode check alone detects an
  unlink/replace. The fd probe only runs in the empty-identity state
  (fresh DB, post-close reopen).
- DeletedWalGenerationError now subclasses StateDbReplacedError, so the
  gateway retry queue and run_agent flush divert transcripts to the
  JSONL fallback exactly as they do for a replaced store, instead of
  retrying forever against a halted handle.
- __init__ refuses once (under the startup lock) instead of twice per
  open, halving the system-wide /proc scan; dropped the dead
  include_self parameter and the dead _IS_WINDOWS clause.
- Test fixes: rstrip(' (deleted)') char-set bug -> removesuffix; the
  non-linux test now patches sys.platform (the real gate) instead of
  _IS_WINDOWS.
2026-09-02 18:24:54 +05:30
Cursor Agent 7f7df1ce44 fix(state): refuse SessionDB open and writes on a deleted WAL generation
A live writer can keep a deleted state.db-wal inode while a second opener
mints a fresh WAL at the same path. Fail closed on writable open (before
connect) and on the write-path sidecar identity check so the second
generation is never created.

Co-authored-by: Noa <rainbowgore@users.noreply.github.com>
2026-09-02 18:24:54 +05:30
Teknium 89e2e4f572 docs(sessions): explain why the canonical Bot Chat guard is provenance-blind (#99517)
Follow-up to the salvaged #99560 commit: restore the original docstring
wording (user writes still always land everywhere else), name the
llm-outranks-derived hole the guard now closes, and annotate the two new
tests with what each pins.
2026-09-02 05:39:24 -07:00
fangliquanflq fb9a294794 fix(sessions): protect canonical Bot Chat from auto-title 2026-09-02 05:39:24 -07:00
kshitijk4poor e9160625dc fix(state): also disable SQLite's internal close-time checkpoint on quarantine (py3.12+)
Skipping the explicit PRAGMA wal_checkpoint(PASSIVE) in close() left
sqlite3.Connection.close() running SQLite's own last-connection PASSIVE
checkpoint, which still checkpoints the WAL and unlinks -wal/-shm on a
structurally corrupt file (E2E: the -wal vanished on close despite the
quarantine). Python 3.12+ exposes SQLITE_DBCONFIG_NO_CKPT_ON_CLOSE via
Connection.setconfig(); arm it in _halt_db_corrupt so the WAL image
survives close() for forensics/recovery. On 3.11 the switch does not
exist; the docstring and docs now say so instead of claiming sqlite3
cannot reach it at all.

Follow-up to #101095; flagged by JoaoMarcos44 on #101093.
2026-09-02 16:57:21 +05:30
leomcamilo bcc2e65818 fix(state): quarantine SessionDB handle after structural corruption
A bare SQLITE_CORRUPT/NOTADB on a live write (not FTS-scoped, not a
replaced file) now sets a sticky per-instance flag: later writes fail
fast with StateDbCorruptError, the handle never reopens after close(),
and close() skips its explicit PASSIVE WAL checkpoint. Gateway and agent
flush paths divert pending transcripts to JSONL/spool like the replaced
case instead of retrying forever.

Field evidence: a handle that kept writing for ~50 minutes after the
first structural error checkpointed 15 pages under the wrong page
numbers on shutdown (page 1 <- messages_fts_trigram_data leaf), turning
"malformed" into "file is not a database".

Refs #90837, #90950, #97940, #89332, #45383

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CNX8rNYHqA5pT4tAGSzXtb
2026-09-02 16:57:21 +05:30
kshitijk4poor 6f1733ca22 test(state): reset the single-flight _opening map in the registry fixture
The _clean_registry fixture clears _generations and _retired between
tests; the new _opening map needs the same reset so a test that aborts
mid-construction cannot leave a stale opening event that stalls the
next test's cold acquire.
2026-09-02 16:37:21 +05:30
fangliquanflq 61635e1b53 fix(state): single-flight shared database opens 2026-09-02 16:37:21 +05:30
Teknium 9bc249c7e5 fix(state): report closed stale-open count from auto-maintenance, document the sweep (#54189)
Follow-up on top of the salvaged #94095 commit:
- maybe_auto_prune_and_vacuum() now returns 'closed' (stale open state-owned
  sessions marked ended) alongside 'pruned', so entrypoints can report the
  reconciliation without parsing logs.
- Docstring explains the two-window lifecycle (close now, delete after a
  further retention window).
- Regression test: cron/kanban/subagent rows with ended_at NULL are closed on
  pass 1 and deleted on pass 2; a telegram row is never touched.
- website/docs sessions.md documents the automatic stale-open sweep.
2026-09-02 01:20:47 -07:00
Alex 0aa84bb3fd fix(state): reap stale state-owned sessions safely 2026-09-02 01:20:47 -07:00
Teknium 09b88bab88 fix(state): stop the on-write identity probe cancelling our own POSIX locks (#100368) 2026-09-01 10:51:52 -07:00
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
Teknium 0685bf1992 test: align file-identity guard tests with post-18ac3c4 FTS recovery contract
Main removed the live-path runtime FTS rebuild (18ac3c4fb6) and #99652
narrowed the corruption classifier, after the salvaged PR #89364 branched.
Update its tests to assert the current contract: no fail-open detach on a
replaced file (fts_enabled stays True, no stale marker), fail-open (not
rebuild) for genuine same-file FTS corruption, and FTS-provenance errors
for the guard probe since generic malformed no longer reaches fail-open.
2026-08-31 14:02:55 -07:00
rainbowgits 71256dfd01 fix(state): fail loudly when state.db is replaced under a live process
Detect same-inode cp via a generation stamp, halt FTS repair, and divert
unwritten transcripts to sessions/<id>.jsonl plus the gateway pending spool.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-31 14:02:55 -07:00
Teknium f57802c7ad fix(state): bound the rewind-tail walk at the watermark and rewind concurrent-tail originals too 2026-08-31 10:02:22 -07:00
BrunoBza 9d9d9194d4 fix(state): archive carried-forward compaction tail as rewind rows (#86366)
archive_and_compact() soft-archives every active row with compacted=1 and
then re-inserts compacted_messages as fresh live rows. When the
compressor's protected tail rides inside that list verbatim - which is
the normal batch-compaction shape ([summary] + tail) - the tail's
ORIGINALS end up stored twice per compaction: (active=0, compacted=1)
next to their live clones. search_messages() recalls both flags without
DISTINCT, so every carried-forward message came back once per compaction
(measured up to 4 identical hits) and was mislabeled to users and the
agent as archived "summarized away" content.

Add an optional tail_count parameter: the last tail_count archived rows
are superseded byte-identical duplicates, stamped rewind-style
(active=0, compacted=0, hidden from recall) instead of compacted=1.

Callers:
- batch in-place compaction counts the compressor-tagged tail dicts
  (_COMPACTION_TAIL_MARKER set by compress() on every carried-forward
  message);
- micro-compaction splices [prefix, marker, suffix] - everything except
  the single marker row is carried forward, so tail_count=len-1;
- proactive tool-result pruning rewrites content in place (not verbatim),
  keeping the historical archive-everything behavior.

Fixes #86366
2026-08-31 10:02:22 -07:00
Teknium c0667439ec fix(compression): rotation heals stale automatic ended_at stamps instead of wedging (#88197)
TUI server shutdown stamps ended_at/end_reason='tui_shutdown' on sessions
whose agent keeps running; every rotation then aborts at
publish_compression_child's liveness check forever (the #88197 wedge; the
amplification half was fixed by #88411).

Class fix: is_automatic_end_reason() in hermes_state_common owns the
"accidental infrastructure cleanup vs deliberate boundary" taxonomy.
publish_compression_child clears automatic stamps in its own transaction
and proceeds (parent re-closes with its TRUE boundary,
end_reason='compression'); the #88411 pre-flush guard no longer aborts on
stamps the publish can heal. Deliberate boundaries (compression,
session_reset, explicit close) still fail closed at both sites.

TEST REPIN (deliberate contract change):
test_ended_parent_aborts_before_the_prepublish_flush pinned
"tui_shutdown stamp => rotation aborts and parent must not grow" — the
abort it required IS the #88197 wedge. Repinned as two tests:
- test_automatic_stamp_no_longer_wedges_rotation: automatic stamp =>
  rotation COMMITS (no abort loop, so no growth-by-abort is possible);
- test_deliberately_ended_parent_aborts_before_the_prepublish_flush:
  session_reset (deliberate boundary) => still aborts BEFORE the #47202
  flush, preserving #88411's no-growth contract where an abort remains
  correct.
The class invariant "no aborted rotation grows the parent" holds
everywhere: automatic stamps no longer produce aborts, deliberate
boundaries still abort pre-flush.
2026-08-31 09:58:11 -07:00
Finn763 cae58be1f5 fix(state.db): cross-backend heartbeat gates orphan sweep
A live sibling serve sharing state.db is no longer treated as a dead process by the startup orphan sweep.

Covers the sweep half of #94895. The launchd Errno 48 KeepAlive loop is not addressed here.

Credit: @Finn763
2026-08-27 11:31:59 -05:00
joaomarcos 59993be6e9 docs(sessions): name the reachable ordinal-0 rewind path in the #95868 guards
Precision pass on the comments added by the previous commit: prompt.submit
reaches replace_messages(archive_dropped=True) with an empty prefix on a
confirmed ordinal-0 rewind, which is the production shape that lands a
populated session on message_count = 0. archive_and_compact normally
publishes at least a summary row, so it is pinned as defense in depth
rather than claimed as an equally reachable trigger.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011zTDHnWcBvQsJSX1fBLhuv
2026-08-27 19:54:32 +05:30
joaomarcos 0bfe715e4d fix(sessions): don't let the empty-session sweep delete an archived transcript
`count_empty_sessions` / `delete_empty_sessions` — the dashboard's
"Delete empty (N)" affordance — defined "empty" as `sessions.message_count
= 0`. That column is a denormalized counter over the LIVE (`active = 1`)
rows only, and two production transcript-rewrite paths reset it on purpose
while keeping every dropped turn on disk as `active = 0`:

  * `replace_messages(..., archive_dropped=True)` — the rewind / edit /
    regenerate mode added in #82756 so a taken-back turn stays recoverable.
  * `archive_and_compact` — in-place compaction, which archives the
    pre-compaction transcript under the same session id (#38763).

A chat rewound to its first turn, or compacted with an empty live set,
therefore reports `message_count = 0` while still holding its entire
history — and those soft-archived rows are the only copy. A gateway reload
is what makes the row eligible: it stamps `ended_at` on every detached
session (`end_reason='ws_orphan_reap'`), satisfying the sweep's
`ended_at IS NOT NULL` gate. The next sweep then hard-deleted the session
row AND `DELETE FROM messages`, destroying the transcript silently.

Every other emptiness test in `hermes_state` already defends the counter
with a real `EXISTS (SELECT 1 FROM messages ...)` probe
(`delete_session_if_empty`, `prune_empty_ghost_sessions`,
`list_never_active_keyed_sessions`, `find_recoverable_session`). This
sweep was the only destructive path that trusted the counter alone. It now
uses the same probe, via one `_EMPTY_SESSION_WHERE` selector shared by the
count and the delete so the button's N and the sweep it triggers can never
disagree again. The counter stays as a cheap prefilter; `EXISTS` is the
authority.

Genuinely message-less rows are still swept — the feature is unchanged for
the case it was built for.

Fixes #95868

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011zTDHnWcBvQsJSX1fBLhuv
2026-08-27 19:54:32 +05:30
Teknium 4896cab0eb fix(tui): _init_session cwd hydration must not fall back to the launch DB
Sibling site of the same class fixed in the previous commit: a failed
profile-store open during _init_session fell back to _get_db(), hydrating
and persisting a named-profile session's cwd row against the launch
state.db. Fail closed instead — skip the hydration (log a warning) so
nothing ever reads or writes the wrong profile's store. The other
profile-store open sites (_db_for_profile, _ensure_session_db_row,
_session_db) already degrade to None/skip and were left as-is.
2026-08-26 21:37:32 -07:00
EndeavorYen 0a4d3abaf3 fix(tui): fail closed when a named profile's state.db won't open
A deferred agent build for a named-profile session swallowed a failed
profile-store open (except Exception -> session_db = None), so _make_agent
silently bound the launch _get_db() handle and every turn bled into the
wrong profile's state.db exactly when the profile store was briefly
unopenable. Opening the named profile then looked blank.

Route the open through _open_profile_session_db, which raises a clear
'profile session store unavailable' error instead; the deferred build's
existing except path turns that into agent_error + an error event, so the
user gets a clear failure and no agent turn against the wrong store.

Salvaged from #90219 (hardening half), adapted to main's current
deferred-build/_transfer_db_to_agent structure.

Related to #87723 and #89789. #88532 covered SessionStore only.
2026-08-26 21:37:32 -07:00
Teknium ad7b7255ab fix(state): renaming a bot's canonical Bot Chat is refused — the title IS the identity (#92473)
Bot Mode resolves the forever-chat by exact-title lookup on
(profile, 'Bot Chat'); no session-id pointer exists. A user rename
therefore orphaned the whole conversation: resolution missed, the next
click minted an empty replacement, and UNIQUE(title) then blocked ever
renaming back. Refuse the rename at SessionDB._set_session_title — the
single write path every surface funnels through (gateway session.title,
/title, CLI rename, REST). Hidden discriminates the registry row, so a
normal visible session a user happens to call 'Bot Chat' stays freely
renameable; re-asserting the same canonical title stays a no-op.
2026-08-26 03:21:42 -07:00
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
halaprix d3e4b50e68 fix(tui): sweep orphaned tui/desktop/subagent session rows at gateway startup
Close session rows left ended_at IS NULL when the in-process websocket
orphan timer dies with the process (#65194). Dual-clock staleness
(started_at AND newest message), desktop included, live in-memory
sessions excluded, scheduled once from both entry.main and the WS
sidecar so desktop/dashboard boots also run the sweep.
2026-08-23 18:58:40 -07:00
poisdahl abf87e7248 Merge current main into composite-carrier fix 2026-08-21 15:56:45 +02:00
Ojas Sharma 7095e23eb2 fix(agent): attribute background-review usage and add cost controls
Persist fork token usage under session_model_usage task=background_review,
emit a per-fork completion log line, and expose enabled/max_iterations/
prompt_file so operators can see and bound the automatic review cost.

Address review feedback: load auxiliary.background_review once per spawn,
classify completion logs by summarize action prefixes, treat explicit
api_call_count=None as the documented default of 1, and WARNING on the
fail-open enabled-gate path.
2026-08-16 06:38:38 -07:00
poisdahl 7ca1987459 Merge upstream main into PR 81234 2026-08-16 12:20:50 +02:00
Teknium 933ef69470 feat: session picker lifecycle status + delete 2026-08-15 18:15:22 -07:00
poisdahl 3f075d41dd Merge remote-tracking branch 'origin/main' into codex/81234-live-main-final
# Conflicts:
#	tui_gateway/methods_prompt.py
#	tui_gateway/server.py
2026-08-15 12:05:14 +02:00
Teknium 62d3ef683c fix(tests): subprocess-surviving isolation marker closes the #82770 fixture escape
The hermetic conftest now exports HERMES_TEST_ISOLATION (value = the tmp
isolation root) before any test module imports, and re-pins it per test in
_hermetic_environment. hermes_state._running_under_pytest() honors the
marker as a test-context signal alongside PYTEST_CURRENT_TEST /
PYTEST_VERSION.

Why a third layer: PYTEST_* belongs to pytest, and tests that spawn
children routinely rebuild the child env and strip it ("the subprocess
must look like a real CLI" — tests/cli/test_exit_watchdog_signal_arm.py,
tests/hermes_cli/test_config_loader_e2e.py do exactly this on purpose).
Such a child loses the HERMES_HOME redirect and the guard's arming signal
in one step, which is how 700+ zero-message fixture rows (dm:123, chat-1,
wx-chat, ...) landed in a developer's production state.db. The marker is
OURS: stripping it is never required to make a child "look real" (no
production code branches on it except the guard), it inherits by default,
and children that genuinely need a real DB use the sanctioned
HERMES_STATE_DB_GUARD_BYPASS=1 hatch instead.

The ancestry-walk layer (previous commits) stays: it covers children whose
env was rebuilt from a completely empty dict. The marker layer covers the
common **os.environ-derived rebuilds cheaply (one dict lookup, no psutil),
and — unlike ancestry — also covers detached/daemonized children that
escape the process tree.

tests/hermes_state/test_isolation_marker_env.py pins: the conftest export,
the marker-alone arming, the rebuilt-env child refusing the production
path, and the bypass hatch. Sabotage-verified: 3/6 fail without the fix.
test_live_db_guard_ancestry._scrubbed_env now strips the marker too, so
the ancestry tests keep proving ancestry rather than riding the marker.
2026-08-15 02:20:13 -07:00
joaomarcos 3f026fe210 test(state): assert the live-DB guard against the real platform root
`REAL_ROOT` hardcoded `Path.home() / ".hermes"`, but the guard resolves
`%LOCALAPPDATA%\hermes` on Windows. The paths under test were therefore
*correctly* classified as non-production, the guard never fired, and all
five TestProductionPathRefused cases failed for the wrong reason — the
guard was effectively unasserted on Windows.

Derive the root from `_real_platform_state_root()` — the same function the
guard uses — so the tests follow the implementation across platforms, and
build the unnormalized-spelling case from it instead of a second hardcoded
`~/.hermes`. Skips at module level if no platform root resolves.

Refs #82770
2026-08-15 02:20:13 -07:00
joaomarcos 174ce8770d fix(state): arm the live-DB guard by process ancestry, not env alone
Production `state.db` files accumulate zero-message "open" gateway session
rows carrying test-fixture identities (`chat-1` / `user-1` / `wx-chat`), with
matching `gateway_routing` scopes pointing at `pytest-of-*` temp directories.

The escape is structural. Hermetic isolation rides entirely on the process
environment: `HERMES_HOME` says *where* to write, `PYTEST_CURRENT_TEST` /
`PYTEST_VERSION` say *whether the guard is armed*. Both travel in the same
carrier, so a child spawned with a rebuilt environment loses them together —
it resolves the developer's real `state.db` *and* silences the only check
that would have stopped it, in one step. The guard is a no-op in precisely
the situation it was written for.

Back the env probe with process ancestry, which survives an env rebuild:

* `_process_looks_like_pytest()` matches a pytest launcher by argv token
  basename, so `/tmp/pytest-of-dev/...` paths in real argv cannot
  false-positive, and an unreadable process is never assumed to be a test.
* `_has_pytest_ancestor()` walks parents via psutil, memoised, and fails
  open when psutil is unavailable — a real `hermes` run pays for at most
  one walk and keeps the previous behaviour if the walk errors.
* `_in_test_context()` checks env first (two dict lookups, covers the
  in-process case) and only then ancestry.

`_STATE_DB_GUARD_BYPASS` is a module global and cannot cross a process
boundary, so ancestry-armed children would have had no way to opt out at
all; `HERMES_STATE_DB_GUARD_BYPASS=1` is the env-carried twin.

Also sweeps the rows already written. Bulk prune/archive cannot reach them:
their shared selector is pinned to `ended_at IS NOT NULL` so a live session
is never picked, which permanently excludes every never-closed row. Adds a
narrower selector — keyed, still open, and with no messages, tokens, tool
calls, API calls, activity or title — behind
`hermes sessions prune --never-active` (default floor 30 days, honours
--dry-run/--yes). Routing entries naming a deleted row go with it, so the
gateway is never left resuming a session id that no longer exists; `pinned`
and `archived` rows are excluded as explicit user intent.

Closes #82770
2026-08-15 02:20:13 -07:00
Teknium fbaea9bddc feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable) (#86797)
* feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable)

Adds a source-orthogonal, archive-orthogonal 'hidden' session flag meaning
'don't show in the global Sessions sidebar, but stay fully resumable by the
surface that owns it'. Mirrors the existing archived/pinned capability end to
end, so it's a generic widening (any plugin that owns its own session lifecycle
- kanban, Bot Mode, future plugins - can keep its sessions out of the shared
recents list) rather than a per-plugin special-case.

- Schema: hidden INTEGER NOT NULL DEFAULT 0 on sessions (additive; lands on
  existing DBs via the declarative _reconcile_columns ADD COLUMN path, same as
  archived/pinned - no version-gated migration).
- DB: SessionDB.set_session_hidden(session_id, hidden) (clones set_session_pinned
  incl. the compression-lineage recursive CTE); list_sessions_rich gains
  include_hidden=False, appending 's.hidden = 0' by default so hidden rows drop
  from every listing path (and the REST sidebar endpoints inherit it with no
  change).
- Gateway: session.set_hidden RPC (mirrors session.title); session.create accepts
  hidden=true, deferred via pending_hidden and applied in _ensure_session_db_row
  when the row is lazily created (mirrors pending_title).
- REST parity: PATCH /api/sessions/{id} accepts+bool-validates 'hidden' ->
  set_session_hidden; _session_response exposes it.

Enables Hermes-Bot-Mode to hide canonical 'Bot Chat' sessions from the sidebar
(NousResearch/Hermes-Bot-Mode#46) WITHOUT retagging source (which would mis-set
the agent platform). Bot Chats keep source=desktop. Gateway RPC needs a
SERVE-backend restart to take effect live. 1 focused test (default-exclude /
include_hidden / unhide round-trip).

* fix: teach lost-and-found recovery about the 55-column sessions layout

Adding the 'hidden' column makes the current sessions table 55 columns. The
SQLite lost-and-found recovery classifier keys off the physical field count
(SESSIONS_LAYOUT_NFIELDS) to identify a salvaged sessions row, so a recovered
current-layout row (nfield=55) would otherwise be unrecognized and dropped.
Add 55 to the frozenset (54/52 stay as historical prefixes) and update the
column-count assertions + synthetic current-layout insert in the recovery test.

---------

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-15 00:31:37 -07:00
zuowen7 bbb6cc7e99 test: cover tool-message dedupe and latest paging for include_compacted (#80680)
Dedupe key now includes tool_call_id/tool_calls/tool_name: compaction
copies carry those fields verbatim, so identical tool messages across
generations still collapse, while distinct tool calls sharing
role/content/timestamp are never merged. Add endpoint-level coverage
for the desktop's real read path (limit + order=latest +
include_compacted=true).
2026-08-14 21:08:14 -07:00
zuowen7 f71f91a39b fix(desktop): surface compaction-archived messages in transcript reads (#80680) 2026-08-14 21:08:14 -07:00
poisdahl 3e5e4c5d20 fix(agent): preserve live turns in compaction carriers 2026-08-13 21:38:30 +02:00
joaomarcos 60645f8a53 fix(state): make a rewind truncation recoverable instead of a hard DELETE (#82756)
Guarding the *aim* of a rewind still leaves every other way of aiming it
wrong terminal. All three reported incidents (#70516, #80763, #82756) ended
at the same write — `replace_messages()` in the `prompt.submit` truncation
path — and all three were unrecoverable for the same reason: the rows are
DELETEd, which also evicts them from the FTS index, so there is no `active=0`
archive and nothing to restore from.

The codebase already draws this distinction and already has the safe half of
it. `archive_and_compact` is documented as "the durability-preserving
alternative to replace_messages"; `rewind_to_message` — the `/undo` path —
soft-deletes to `active=0, compacted=0` and keeps the rows "on disk for audit
/ forensic inspection". The desktop rewind is the same user-facing operation
as `/undo` and was the one taking the destructive branch.

`replace_messages(..., archive_dropped=True)` flips the DELETE to a
content-preserving `UPDATE messages SET active = 0`, reusing the existing
transaction and the existing `active=0, compacted=0` marking so the dropped
turns stay readable via `get_messages(..., include_inactive=True)` and stay
out of session search (`compacted=0` = "the user took it back", vs
compaction's `compacted=1` = "summarized away, still discoverable").

The live transcript is byte-identical either way — only the durability of the
dropped turns changes. The parameter defaults to False, so the fork handler,
the ACP adapter and `gateway/session.py` keep their current semantics
untouched; a test pins that.

`active_only=True` stays on the call: #80216 still applies, and archiving must
not disturb rows an earlier compaction deliberately archived.

Test doubles for `replace_messages` in the gateway suite are widened to the
real signature — they are stand-ins for SessionDB, and a double that does not
accept what production passes silently converts this write into a 5008.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 11:01:15 +05:30
joaomarcos c790ed2a5d fix(state): recover gateway sessions stranded without a routing identity
When state.db's write path fails (corrupt FTS, or a crash landing between
routing publication and row creation), the live gateway conversation can end
up in a session row that never received its identity columns: session_key,
chat_id, chat_type and origin_json are all NULL. In-memory routing hides the
damage for as long as the gateway stays up. After a restart the chat is
resolved from the DB, and find_latest_gateway_session_for_peer cannot see
that row — both of its queries match on the very columns it lacks — so the
chat resumes the last keyed sibling instead, days older. The messages were
never lost, only unreachable.

Hardening the write side cannot reach a row that is already damaged, so add
the offline repair path the tracking issue asks for:

- SessionDB.find_orphaned_gateway_sessions() reports message-bearing rows
  with no session_key, and names the predecessor each one continues only
  when the evidence is unambiguous — a recorded parent_session_id
  ("lineage"), or exactly one keyed row of the same source and compatible
  user_id that fell quiet within 15 minutes of the orphan's start
  ("contiguity"). Contested pairs are reported with a reason and left alone:
  a wrong adoption would splice one person's conversation into another
  person's chat. Branch, delegate and tool rows are excluded — they are
  unkeyed by design, not by damage.
- SessionDB.adopt_orphaned_gateway_session() stamps the orphan from the
  predecessor (never overwriting a column that already has a value), records
  the lineage, and retires the predecessor under end_reason
  'superseded_by_repair' — a reason recovery does not treat as resumable, so
  the repaired row wins the chat from then on. The pair is re-verified inside
  the write transaction, making a concurrent heal a no-op rather than a
  conflicting write.
- `hermes sessions repair-routing` drives both. It reports without touching
  the database; --apply confirms first and warns that a running gateway
  still holds the old mapping in memory.

Refs #82616.
2026-08-09 14:06:06 -07:00
cryptoyasenka f2d03c1f2a fix(state,cli,tui-gateway): keep reasoning fields intact across forks and branches
get_messages() only deserializes content and tool_calls; the structured
reasoning columns (reasoning_details, codex_reasoning_items,
codex_message_items) come back as the raw TEXT they were stored as.
Feeding those rows straight back into a write, which is exactly what
the POST /api/sessions/{id}/fork handler does by piping get_messages()
into replace_messages(), hit an unguarded json.dumps() and stored the
already-serialized string encoded a second time. On replay of the fork,
json.loads() then yields the inner string instead of a list, and every
consumer's isinstance(..., list) gate silently drops it: preserved
Anthropic thinking blocks, Codex encrypted-reasoning/message-item
replay, and OpenRouter multi-turn reasoning context are all lost after
a fork, with one more encoding layer added per fork.

The /branch copy loop had the same defect from the other side: it
forwarded reasoning but none of the structured columns, and both TUI
branch writers persisted role/content alone, dropping reasoning and
reasoning_content along with them.

Route the six dumps sites in append_message and _insert_message_rows
through a shared guard that keeps already-serialized strings as-is;
structured values from the live runtime are dumped exactly as before.
Forward the reasoning fields in all three branch writers, matching the
set gateway/slash_commands.py already forwards on its own /branch path.
2026-08-08 17:37:26 -07:00
kshitij a82910c37b test: fold review findings — plain fixture, call-shaped probe guard, public-API row counting
- state_db fixture: drop the HERMES_HOME setenv + sys.modules purge
  (SessionDB takes an explicit path; tests/conftest.py already sandboxes
  HERMES_HOME; the purge risks split-class identity for other modules
  holding the old hermes_state reference) — matches the plain
  tests/hermes_state/ sibling fixture pattern.
- probe guard asserts on the CALL ('has_archived_messages(') instead of a
  local-variable name — a reintroduced probe under any rename now trips
  it (mutation-checked: renamed-probe reintroduction fails the guard;
  restored stack green).
- _archived_count uses the public get_messages(include_inactive=True)
  instead of poking db._lock/_conn.
2026-08-07 15:42:49 +05:30
kshitij ee6d79648a fix(state): finish the #80216 bug class — archive-preserving rewrites at the two remaining sibling sites
#80216 fixed /retry (and a follow-up fixed yuanbao recall) destroying
soft-archived active=0/compacted=1 in-place-compaction rows via the
destructive replace_messages default. Two sibling sites still carried the
same class:

- acp_adapter/session.py _persist (non-owned-agent branch): probed
  has_archived_messages and FAILED OPEN into the destructive full replace
  on any probe error; the probe can also race a concurrent
  archive_and_compact. Now passes active_only=True unconditionally — on a
  fresh create/fork every row is active=1 so behavior is identical, and
  the probe (its only production caller) is deleted.
- tui_gateway/methods_prompt.py edit/regenerate truncation: bare
  replace_messages() deleted the archived transcript of a compacted
  session on every edit/regenerate. Now active_only=True.

hermes_state.has_archived_messages docstring updated (probe is now
test/diagnostic-only). Test stubs in test_tui_gateway_server.py accept the
new kwarg. New regression tests: real-SQLite archive-survival for both
write shapes, fresh-session equivalence (the claim the unconditional
switch rests on), and source-level guards pinning that neither site
re-grows the fail-open probe (both mutation-checked: revert either fix and
its guard fails).
2026-08-07 14:42:32 +05:30
Teknium 19fc9c103e fix(tests): fail hard when pytest resolves the production state.db (live-DB isolation guard)
Forensics on a live developer machine found pytest fixture rows inside the
REAL ~/.hermes/state.db — sessions with chat_id 'chat-1', '123', 'wx-chat',
and gateway_routing rows whose scope was literally under /tmp/pytest-of-*/.
A pytest-spawned process also opened the live DB and flipped its journal
mode (journal_mode=DELETE fallback on SQLite 3.50.4) under the WAL-mode
gateway writer, destroying committed transcripts ("Persisted transcript
lagged live cached history ... possible FTS write corruption", 15+
occurrences). The existing live-system guard covers kill primitives but not
the SessionDB/SessionStore write paths.

Root cause (leak vector): the session-level HERMES_HOME sandbox in
tests/conftest.py only created a tempdir when HERMES_HOME was UNSET. On a
machine where the shell (e.g. gateway-launched, or an exported
HERMES_HOME=~/.hermes) hands pytest the production home, the sandbox was
skipped entirely — every argless SessionDB()/SessionStore() and every
collection-time DEFAULT_DB_PATH froze onto the real state.db.

Fixes (fail the class, one owner):

* hermes_state._ensure_test_isolation(): single choke point wired into
  SessionDB.__init__ (every construction, incl. read_only). Under pytest
  (PYTEST_CURRENT_TEST / PYTEST_VERSION — inherited by subprocess
  children), a db path resolving to <real-root>/state.db or
  <real-root>/profiles/<name>/state.db raises RuntimeError('live-system
  guard: ...') before any connection, mkdir, or journal-mode pragma.
* tests/conftest.py: session sandbox now also tempdir-redirects a pre-set
  HERMES_HOME that points at the production root (the actual escape
  vector); kanban deny-list capture updated to match. New autouse
  _state_db_write_guard fixture honors the existing
  @pytest.mark.live_system_guard_bypass marker as the escape hatch and
  feeds custom (non-~/.hermes) production roots into the guard deny-list.
* gateway/session.py: SessionStore.__init__ no longer swallows the guard's
  RuntimeError into the JSONL fallback — guard trips are loud.
* tests/hermes_state/test_live_db_isolation_guard.py: behavioral
  regression tests — production paths (direct, profile, read-only,
  unnormalized, default-resolution) raise; tmp HERMES_HOME works; bypass
  marker works; SessionStore re-raises guard errors but still degrades on
  ordinary failures; subprocess child without HERMES_HOME is refused while
  a hermetic child succeeds.

No new HERMES_* env vars; no hardcoded ~/.hermes (platform root comes from
hermes_constants._get_platform_default_hermes_home()).
2026-08-06 07:49:50 -07:00
Ryder Freeman 84e93ffefb fix(state): stop delegate/tool children corrupting compression lineage
get_compression_lineage's forward walk accepted any non-branch child as
the compression continuation. Delegate subagent rows (_delegate_from)
and tool-tagged rows (source=tool) created before the real continuation
were picked as the lineage successor, so the lineage — and session .md
export built on it — followed a subagent's transcript instead of the
actual conversation continuation.

Rename _is_branch_child_row to _is_explicit_fork_child_row, treat
_delegate_from and source=tool rows as explicit forks alongside
_branched_from, and require _is_compression_child_row in the forward
walk instead of merely excluding branches.

Sliced from PR #79024 by @RyderFreeman4Logos (the cache-scope portion
of that PR is tracked separately in #79017).
2026-08-05 13:50:26 +05:30
Brooklyn Nicholson ec0c8d9c20 feat(state): sessions carry read/unread state
Adds a last_read_at watermark to the sessions table so surfaces (CLI,
TUI, desktop) can badge unread conversations. Read state derives from
the watermark vs latest activity, so new messages flip a conversation
back to unread with zero writes on the message path. NULL means never
tracked, so shipping the column doesn't badge pre-existing history.

set_session_read() stamps the whole compression lineage, matching the
archive/pin semantics; list_sessions_rich() rows carry a derived
`unread` key. DB layer only — no surface exposes it yet.
2026-08-04 12:32:27 -06:00
kshitij da6d9604dd refactor(state): fold simplify findings — reuse _insert_message_rows, share guards, chunk seeds
Simplify-pass folds on the #23254 salvage:

- REUSE (HIGH): append_messages_batch now delegates row serialization to
  the pre-existing _insert_message_rows helper (already shared by
  replace_messages / archive_and_compact / portability import) instead
  of adding a third serialization path (_prepare_message_row +
  _MESSAGE_INSERT_SQL are gone). One row-writer for every multi-row
  path; the row-ID return was consumed by no production caller, so the
  batch returns the inserted count.

- QUALITY (HIGH): the compression-lock + compression-closed admission
  guards are extracted into _check_transcript_write_guards, shared by
  append_message and append_messages_batch (previously duplicated 23
  lines that had already needed targeted fixes, #74478). The role-gated
  reasoning filtering is no longer duplicated in run_agent.py — it
  lives at its one site inside _insert_message_rows.

- EFFICIENCY (MEDIUM, measured): unbounded seed copies hold one BEGIN
  IMMEDIATE for seconds (10k rows ~= 2.4s; FTS triggers dominate) and
  monopolize the in-process write lock. append_messages_batch grows a
  chunk_rows param; all seed/copy call sites use chunk_rows=500. Same
  recovery semantics as the old per-row loops, bounded lock holds.

- REUSE (MEDIUM): the two remaining per-row branch-copy loops found by
  the pass (gateway/slash_commands.py /branch, hermes_cli
  cli_commands_mixin.py branch) are converted to chunked batches too
  (AsyncSessionDB's generic to_thread forwarder covers the async site).

Turn-flush benchmark unchanged after the refactor: 2.43 -> 0.87 ms
median per 5-message flush (64% faster).
2026-08-03 20:43:38 +05:30
devsart95 06ae5b6faa perf(state): batch the turn flush into one SQLite transaction
Re-derivation of #23254 (@devsart95) on today's flush loop. The turn
flush in _flush_messages_to_session_db wrote one BEGIN IMMEDIATE
transaction per message row; a typical agent turn (user + assistant +
tool results) paid 3-8 transactions -- and, off WAL (the default on
macOS while the WAL-reset guard is active), 3-8 fsyncs -- per turn.

Adds SessionDB.append_messages_batch: same row shape as append_message
(shared _prepare_message_row serializer + _MESSAGE_INSERT_SQL column
list, so the two writers cannot drift), same compression-lock and
compression-closed guards, one aggregated session-counter UPDATE, one
transaction for the whole batch. Row serialization stays outside the
write lock.

The flush loop now collects the turn's new rows and writes them in one
call. All-or-nothing pairs exactly with the persisted-marker stamping:
on failure no rows landed and no markers were stamped, so the next
flush re-writes the whole tail (same recovery contract as before,
minus the partial-prefix case that could double-count).

Measured (same harness, 5-message turn, journal_mode=DELETE,
synchronous=FULL): 2.32ms -> 0.83ms median per turn flush (64% faster,
5 fsyncs -> 1). On WAL the win is smaller but the atomicity fix holds.
2026-08-03 20:43:38 +05:30
Teknium 39975613b1 test: prune wave 2 + speed fixes — 28,106 → 19,757 test functions, suite wall 315s → 294s
Second, deeper pass over tools/gateway/hermes_cli plus first pass over
the trees wave 1 missed (acp, acp_adapter, skills, computer_use, docker,
dashboard, conformance, monitoring, secret_sources, hermes_state,
providers). Same rubric as wave 1 (AGENTS.md test policy); security,
alternation/caching invariants, issue-number regressions, and E2E kept.

Real test-quality fixes found and rooted out along the way:
- tests/tools/test_command_guards.py made real auxiliary-LLM HTTPS calls
  (DEFAULT_CONFIG smart-approval leaked in) — pinned approval
  mode=manual via autouse fixture: 17.4s → 0.4s.
- test_model_switch_custom_providers.py / test_user_providers_model_switch.py
  silently probed live provider catalogs (~2s/test) — stubbed
  cached_provider_model_ids/provider_model_ids/fetch_api_models.
- test_telegram_noise_filter.py: 15-platform copy-paste matrix over
  shared gateway.run logic → 3 representative platforms (55s → 3.9s).
- test_gateway_shutdown.py: stop()'s 5s interrupt-deadline loop spun on
  MagicMock agents — interrupt.side_effect now clears _running_agents
  (22s → 1.0s).
- test_gateway_inactivity_timeout.py poll-harness timings shrunk 3-5x
  (24s → 1.1s); test_mcp_stability.py backoff/SIGTERM-grace sleeps
  patched (15.4s → 2.5s); test_async_delegation.py negative-drain wait
  5s → 0.5s.
- test_telegram_init_deadline.py: loop-block margin restored to 1.0s
  with rationale comment — the watchdog-dump assertion needs the loop
  blocked well past deadline+grace under parallel load (flaked once in
  the 40-worker verification run at a 0.2s margin).

Verification: full hermetic suite via scripts/run_tests.sh —
2,438 files, 21,718 tests passed, 0 failed, 293.9s wall.
Suite totals vs original baseline: 46,820 → 19,757 test functions
(−57.8%), wall 583.5s → 293.9s (−50%), subprocess CPU 13,564s → 11,623s.
2026-07-29 13:39:40 -07:00
Teknium 6b81590c55 test: prune low-value tests suite-wide (wave 1) — 46,820 → 28,106 test functions
Systematic prune per AGENTS.md test policy, one pass over every major
test tree (gateway, hermes_cli, tools, agent, run_agent, plugins, cli,
cron, tui_gateway, honcho/openviking, root-level):

- DELETE: source-reading tests (read_text/getsource on prod files),
  change-detector tests (exact catalog counts, model-name snapshots,
  config version literals), mock-echo tests (assert a mock returns what
  it was told), assertion-free/trivial tests, near-duplicate
  parametrizations (boundaries + one representative kept), async/sync
  twin duplicates, cosmetic within-file variations.
- KEEP (mandatory): security/redaction/approval guards, message-role
  alternation invariants, prompt-caching/deterministic-call-id
  invariants, issue-number regression tests (deduped), E2E tests.
- 6 test files deleted outright (script-style/no-assert or fully
  redundant); conftest.py, fakes/, fixtures/ untouched.
- tests/acp/conftest.py added: autouse fixture stubs the live
  models.dev/GitHub/Copilot/Anthropic inventory fetches that ACP server
  tests performed on every session create — test_server.py 147s → 3.4s,
  and the tests are now genuinely hermetic.
- Sleep-based slowness shrunk where safe (codex_ttfb_watchdog,
  compression_concurrent_fork, etc.); no wall-clock assertion tightened.

Verification: full hermetic suite via scripts/run_tests.sh —
2439 files, 31,130 tests passed, 0 failed, 0 flaky retries, 315s wall
(baseline: 583s wall, 13,564s subprocess CPU).
2026-07-29 13:10:23 -07:00