hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Refs #100368. The forensics thread established that a sqlite3 CLI with
the WAL-reset opener bug (fixed 3.51.3+ / backports 3.50.7 / 3.44.6;
Debian/Ubuntu system shells 3.45.1/3.46.1 are in the vulnerable band)
unlinks the live -wal/-shm pair when pointed at a live state.db whose
writer's DMS lock has been cancelled, splitting the store into two
concurrent generations whose acknowledged writes vanish while both
report integrity_check ok. Hermes' own corruption banners instructed
exactly that command.
- gateway corruption broadcast, run_agent corrupt-cause explanation,
hermes_state repair-budget and forensic-backup refusals, and the
kanban manual-recovery hint now route operators to
`hermes sessions recover --source <db>` (which snapshots the damaged
bundle before any shell touches it) and warn against a raw sqlite3
shell on the live file
- find_sqlite3_cli() now refuses a WAL-reset-vulnerable shell for the
page-level salvage lane even on the snapshot, reusing the canonical
gate from hermes_cli.sqlite_runtime so the embedded runtime and the
salvage shell can never disagree
- find_sqlite3_cli_refusal() records why a shell was refused so the
lost_and_found lane can tell the operator exactly what to install
instead of a generic "not found"
- regression tests cover the version gate (vulnerable/fixed matrix, the
mirror check), every refusal reason, and each guidance site
Test plan:
- scripts/run_tests.sh tests/hermes_cli/test_sqlite3_cli_salvage_gate.py
tests/test_state_db_repair_loop_cap.py
tests/run_agent/test_corruption_recovery_guidance.py
tests/hermes_cli/test_session_recovery_lost_and_found.py
tests/hermes_cli/test_session_recovery.py tests/test_sqlite_wal_reset_gate.py
tests/hermes_cli/test_sqlite_runtime.py - 91 passed, 1 skipped locally
(cherry picked from commit e62940d1021e80e9b7d6423ced1cbdfe7dd0c37d)
The plausibility gate excluded stub rows by a hard-coded title literal that
session_lost_and_found builds inline at two sites; rewording either would have
turned every stub into a mis-mapped row. One STUB_TITLE_PREFIX constant.
Follow-ups on the #101423 salvage (#101409):
- `stub_missing_parent_sessions` legitimately writes `started_at = 0.0` when
no timestamped message survived; a salvage where only stubs remain is
depleted, not mis-mapped. Stub rows leave the sessions denominator.
- Reuse `_EPOCH_LOW` from session_lost_and_found instead of a second
2001-epoch constant; collapse the two per-table blocks into one loop.
- Tests: a real upgraded-layout source (started_at physically appended) with
page 1 zeroed goes through the real `.recover` lane and the report comes
back `verified: False`; a stub-only output is not flagged; a mapped row
with a NULL title (the mis-mapped shape) still counts as mapped.
The lost_and_found salvage lane maps cells positionally onto the
destination template's declared column order, but a source upgraded
via ALTER TABLE has its columns in the order they were added, which
differs from SCHEMA_SQL whenever a column was inserted mid-definition
(#101409). Every row still inserts, so integrity/FK/FTS/count checks
stay green and the report ends up verified: true — while all 1,875
sessions in the reporter's DB got started_at = 0.0 with counters and
URLs shifted into the wrong columns.
Add a semantic plausibility gate to the salvage lane: when every
salvaged sessions.started_at or messages.timestamp is NULL or before
2001-09, the cells were mapped onto the wrong columns and the
recovery is reported with errors and healthy: false, so verified no
longer claims a mis-mapped output. A partially damaged column (torn
cells on some rows) does not trip the gate — only a systematic
violation does.
Fixing the positional mapping itself needs a historical-layouts
table derived from schema-version history; that design decision is
left to maintainers (suggested fix 2 in the issue).
Follow-up to the salvaged #100350 commits: replace the per-table
'if table == "delivery_obligations"' branches in session_recovery.py and
session_lost_and_found.py with a single _AUXILIARY_TABLE_SCHEMAS registry
(table -> destination DDL initializer) that both the SQL-level and the
lost_and_found lanes consume, so the next lazily-created state.db table is
one entry, not three code paths. The .recover lane now iterates
_CANONICAL_TABLES + _AUXILIARY_TABLES instead of a duplicated literal list.
Tests: the .recover direct-copy lane creates the missing ledger on the
destination; a source-vs-destination obligation count mismatch fails
verification (complete=False) instead of reporting a clean salvage.
Docs: state.db table inventory lists delivery_obligations.
Addresses #100313
The lazy gateway outbox was missing from the recovery inventory, so a
verified salvage could drop owed replies even when the rows were still
readable. Initialize the destination schema and copy the table.
Fixes#80205: when one ordered rowid-edge probe failed,
_salvage_rowid_bounds() substituted the whole SQLite rowid domain and
_copy_table_salvage() burned the 10,000-query budget bisecting a
synthetic tail that could not contain rows, silently omitting readable
boundary rows (field case: message 76882 of 76882). Two-part fix:
* _probe_populated_edge(): gallop outward from the surviving edge with
doubling offsets; a clean 'no rows beyond X' probe caps the domain in
O(log range) queries instead of exhausting the budget on it.
* exact-key singleton salvage: a one-row range scan must advance the
cursor past the hit into the damaged sibling page to prove exhaustion,
which discards the already-produced row; 'WHERE rowid = ?' stops at
the hit, recovering the boundary row exactly like sqlite3 .recover.
* the strict-path refusal now points users at --allow-partial.
New last-resort lane for --allow-partial when the sessions/messages
table schemas themselves are unreadable (previously a hard refusal even
though page-level salvage recovers the rows fine). If a sqlite3 CLI is
on PATH, shell out to '.recover --ignore-freelist' into a scratch
lost_and_found DB, then heuristically map rows back into a fresh
SessionDB-schema database (hermes_cli/session_lost_and_found.py):
classification keyed on nfield counts + sentinel columns (session ids
matching ^\d{8}_\d{6}_, roles in user/assistant/tool/system, known
source strings), covering the current 54-col sessions layout, the
52-col historical layout, a 14-col legacy identity-only salvage,
rowid-alias messages rows and 18-col session_model_usage rows. Missing
parent sessions are stubbed (children are never deleted for FK
cleanup), FTS is rebuilt at the end, and output is labeled BEST-EFFORT
everywhere. Without the CLI the error names the sqlite3 requirement
with actionable guidance. Mirrors a successful manual recovery of a
real corrupt state.db (2026-08-12), and this lane was validated against
that preserved file: 32 sessions / 7 messages / 4 usage rows mapped,
integrity_check ok, opens via SessionDB.
Also fixes#72291: the source-fingerprint 'bundle changed while it was
being copied' error now enumerates that the parent interactive CLI
session itself counts as a Hermes process and suggests a fresh shell or
an immutable snapshot.
Tests use real physical page corruption (flipped b-tree/schema header
bytes), skip the CLI-dependent path cleanly when sqlite3 is absent, and
keep the mapper unit tests binary-independent via a synthetic
lost_and_found DB. Sabotage-verified: reverting the fixes makes the
regression tests fail with the exact field failure shape.
Reported in Discord by @spherohero: `sessions recover --allow-partial` copied
20,817 of 20,824 message rows, then orphan cleanup deleted every one of them
because no session row was salvageable. Final output: 0 sessions, 0 messages.
The salvage worked and then threw the result away -- the exact opposite of
what --allow-partial exists to do.
_cleanup_partial_orphans() removed dependent rows whose session_id had no
matching sessions row. That is correct when a few sessions are lost; it is
catastrophic when the sessions b-tree is damaged worse than messages, which
is the common shape (sessions is small and hot, messages is large).
Reproduced exactly: 500 readable messages, unreadable sessions -> 500
copied, 500 removed, empty output.
Now _reconstruct_missing_sessions() runs FIRST, inside the same transaction,
synthesizing a minimal row per orphaned session_id (only id/source/started_at
are NOT NULL). started_at comes from the earliest surviving message.
Placeholders carry source='recovered' and an explicit title so a fabricated
session can never be mistaken for an original. Same repro now retains all
500 messages under 1 reconstructed session.
Reconstruction is reported as LOSS, not a clean recovery: session metadata
(title, model, timestamps, cost) is genuinely gone even though the
conversation text survived. loss_detected=True, partial=True,
complete=False, with a warning naming the counts.
Also fixes total_removed_or_relinked, which summed every dict value and would
now have counted retained messages as removed.
Sabotage-verified: removing the reconstruction call restores the wipe and
fails the test. 942 targeted tests green.
Second round of @helix4u review on #71779. Both findings reproduced before
fixing.
1. My previous fix turned a crash into SILENT DATA LOSS. Returning
status="missing" for a present-but-unusable state_meta looked like a safe
degrade, but _verify_recovered_database only escalates "failed"/"partial"
into a warning + loss_detected. Measured on the branch: a run that dropped
a real metadata table reported warnings=[], loss_detected=False,
partial=False, complete=True. Strictly worse than the ValueError it
replaced -- that at least failed loudly.
Now "failed" when the table exists but lacks key/value, "missing" only
when genuinely absent. The damaged case yields
warnings=['state_meta copy status is failed'], loss_detected=True,
partial=True, complete=False, while staying verified=True so the output
is still installable-with-review.
2. The race test I wrote had its own scheduling race: after the guard
released the lock, the racer could win before the main thread set the
release event, failing on a correct implementation. Rewritten per
helix4u's design -- copy runs in a worker parked inside the patched
copy, a second worker attempts connect_tracked(), assert it stays blocked,
release, assert it then opens. Deterministic and ~1.1s instead of 10s;
12/12 stable.
Sabotage-verified. Note the third scenario only failed after adding a
unit-level test: recover_session_database short-circuits on the inspection
result when state_meta is entirely absent, so the helper's absent-branch is
unreachable end-to-end and a regression there was invisible. Both statuses
are now pinned directly.
939 targeted tests green.
Post-merge follow-up to #71770. Both defects were found by @helix4u in review
and reproduced against merged main before fixing.
1. Check/use race in _copy_source_bundle (my bug, from the #71770 follow-up
commit). It called has_live_connection(), released the registry lock, and
only then ran shutil.copy2() over the bundle. A tracked connection could
open in that window; the copy's close() then cancels its POSIX advisory
locks -- the exact class #71724 closed. Measured on main: a racer thread
opened a connection mid-copy after blocking 0.000s.
Adds sqlite_safe_read.offline_file_access(), a context manager that holds
the connection-lifecycle lock across an entire multi-step raw access, and
routes the bundle copy through it. Same racer now blocks 10.0s until every
raw descriptor is closed. Any future raw read of a database file (hashing,
moving a bundle aside) should use this rather than a bare pre-check.
2. _copy_state_meta_salvage assumed a 'key' column. A damaged state_meta can
keep 'value' and lose 'key'; columns.index("key") then raised ValueError
and aborted the whole partial recovery. The mirror case (key without
value) would have copied key-only rows and reported the table complete.
Now requires both, matching the non-partial _copy_state_meta, so an
unusable optional table is recorded as missing/failed and --allow-partial
still recovers sessions and messages.
Both regression tests verified by sabotage: reinstating the bare pre-check
fails the race test, removing the key/value requirement fails the other.
937 targeted tests green.
Follow-up to the partial-recovery salvage. _copy_source_bundle() raw-copied
the source state.db and its -wal/-shm sidecars with no live-connection check.
Copying a database file is an open()/close() on it, and close() cancels every
POSIX advisory lock the process holds on that file -- including a running
VACUUM's EXCLUSIVE lock (fe431651c5). hermes_state._backup_db_file already
refuses that situation; this path did not, leaving two policies for one
hazard. Verified against the merged guard: with a tracked connection open,
_backup_db_file returned None (refused) while _copy_source_bundle copied
anyway.
Recovery normally runs as its own short-lived CLI process against an
offline/quarantined file, so this should never fire in practice. It is a
consistency fix, not a live corruption path -- but it is exactly the drift
that reintroduces the bug later.
Regression test verified by sabotage: removing the guard fails it.