hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
Addresses two data-integrity gaps @andrexibiza flagged reviewing #88425.
1. Forensic dedupe no longer reuses the repair-epoch fingerprint.
_db_fingerprint masks SQLite's commit counters and samples only head/tail
so an ordinary write does not re-key the repair budget — the right
predicate for 'same damage epoch', the WRONG one for 'same recovery
image'. A live writer committing rows into an interior page (size
preserved, head/tail untouched) collided under it, so _backup_db_file
handed back a STALE backup that predates real user data. New
_backup_content_identity() digests the whole file + every sidecar; the
dedupe uses it. The O(n) read is cheaper than the O(n) copy it avoids on a
hit.
2. Backup bundle is now published atomically. The promotion loop replaced
files one at a time (main first) and cleanup unlinked only staging srcs,
so a sidecar os.replace failure after the main promotion left the
final-prefix main backup on disk — a countable-but-incomplete bundle that
passed the #69603 hard stop and deduped as legitimate next pass. Now
sidecars publish first and the main DB last (its name is the commit
marker _existing_malformed_backups counts), and cleanup rolls back every
already-published destination.
Two regressions added (both mutation-checked — each fails on pre-fix code):
- test_backup_not_deduped_after_interior_page_write
- test_publication_failure_leaves_no_countable_partial_bundle
tests/test_state_db_repair_loop_mtime.py: 28 passed.
The pre-repair copy took only -wal/-shm. In rollback-journal (DELETE) mode --
Hermes's fallback on NFS/SMB/FUSE/ZFS and on WAL-reset-vulnerable SQLite builds
-- a hot <db>-journal exists on disk whenever a transaction was open, and that
file is what rolls the damaged bytes back to a consistent state. A forensic copy
without it cannot be recovered by hand, which is the entire purpose of taking
the copy before destructive surgery.
Verified the journal is really there:
files while a txn is open: ['state.db', 'state.db-journal']
files after commit: ['state.db']
Add _DB_SIDECAR_SUFFIXES = ("-wal", "-shm", "-journal") and use it at the four
sites that must agree: the disk-guard sizing, the staging copy, the
backup-count exclusion in _existing_malformed_backups (so a copied journal is
not itself counted as a forensic backup), and _prune_malformed_backups (which
otherwise leaks one journal per pruned backup, quietly defeating the retention
cap this PR is partly about).
Matches the spelling hermes_cli/session_recovery.py:61 already uses for the
same concept.
Third self-review pass found the content fingerprint was still defeated on
rollback-journal deployments, by the same mechanism as the original mtime bug.
The head sample starts at byte 0, so it covers the database header's file
change counter (bytes 24-27) and version-valid-for (92-95). In DELETE mode a
commit writes the main file directly and bumps both. A malformed-SCHEMA DB
still accepts writes -- that is the whole premise of this PR -- so any ordinary
session write between passes re-keyed the ledger:
DELETE, 18MB db, one peer UPDATE between passes (before this commit)
pass 1..6: attempts=1 every pass, exhausted=False -> unbounded loop
after
pass 1..3: attempts=1,2,3 pass 4: BLOCKED
WAL is unaffected (commits land in -wal; the main header only moves on
checkpoint), so this was invisible on a WAL host and reproducible on every
NFS/SMB/FUSE/ZFS or WAL-reset-vulnerable host -- exactly the deployments the
earlier lock-safety commit was written for.
Mask the two volatile ranges out of the sample. Page 1's sqlite_master b-tree
sits after byte 100 and stays in, so genuine recovery still resets the budget:
verified schema rewrite, index rebuild, VACUUM and truncation all change the
key, while a bare utime and an ordinary commit do not.
Test-cost cleanup in the same file, since the new tests needed a
larger-than-sample fixture and the file was already slow:
- the two guard tests that allocated 450MB of os.urandom now use sparse
truncate (both only ever read st_size), and the new fixtures use 600 rows
rather than 40k;
- file runtime 127s -> 35s.
Self-review of the previous commit found it reintroduced the bug this PR
exists to fix, by a different route.
`_db_fingerprint` fell back to `size:mtime_ns` when a live connection made the
content read unsafe. The ledger compares keys for EQUALITY, and the two keys
have different SHAPES, so a gateway peer connecting between passes flipped the
shape and the counter reset to 1 every time:
pass 1 [offline] attempts=1 fp=8192:58c7924f0fba...
pass 2 [LIVE ] attempts=1 fp=8192:1786972039271402096
pass 3 [offline] attempts=1 fp=8192:58c7924f0fba...
... never reaches _MAX_PERSISTENT_REPAIR_ATTEMPTS
Return None instead, and teach the two ledger helpers to cope:
- `_persistent_repair_attempts_exhausted` falls back to the recorded key's
SIZE prefix (the one component both shapes share and that needs no raw
read) rather than reading as "not exhausted" — otherwise a peer connection
hides an exhausted budget on every pass, same loop.
- `_record_repair_outcome` keeps the key already on record and still
increments, rather than dropping the pass.
pass 1 [offline] attempts=1 pass 2 [LIVE] attempts=2
pass 3 [offline] attempts=3 pass 4 [LIVE] BLOCKED
Intra-pass flips were already safe (the probe and the record are both reached
with the same liveness within one `repair_state_db_schema` call); it is the
cross-pass change that desynced.
Also drops two `type: ignore` directives `ty` flagged as unused, and replaces
the `LiveConnectionError = ()` / `nullcontext()` shim with a real no-op
contextmanager + exception class so the scaffold-install path is honest.
The staging name was derived from the backup name
(`<db>.malformed-backup-<stamp>.incomplete`), which still matches the prefix
`_existing_malformed_backups` selects on -- it excludes only `-wal`/`-shm`.
Three consequences, all reproduced:
- it is COUNTED as a forensic backup;
- it sorts NEWEST (`.incomplete` > the bare stamp), so prune's
keep-3-newest slice retained partials and deleted intact copies -- the
exact inversion the staging change was meant to prevent;
- worst, the dedupe ran BEFORE the sweep, and a staging file orphaned by a
kill mid-copy is a byte-identical copy of the damaged DB, so its
fingerprint MATCHES and it was handed back as the official `backup_path`.
Repair then passed the #69603 hard-stop gate and ran destructive surgery
believing a forensic copy existed, and the next pass's sweep deleted that
very file.
Move staging outside the prefix (`<db>.backup-staging-<stamp>`) and sweep
before the dedupe. The sweep also matches the pre-merge `.incomplete`
spelling so a host that ran the earlier build does not keep prefix-matching
debris that sorts newest and survives prune forever.
Before / after on the same fixture (orphaned staging + a later pass):
before backup_path = ...malformed-backup-<stamp>.incomplete (staging!)
pass-1 forensic copy deleted by the next sweep
after backup_path = ...malformed-backup-<stamp> (real copy)
debris swept, pass-1 forensic copy preserved
The content fingerprint takes a raw descriptor, and close() on ANY descriptor
cancels every POSIX advisory lock the process holds on that file. The
exhaustion probe runs before _backup_db_file's has_live_connection guard, so
the read happened even when a peer SessionDB held a write lock.
Verified end-to-end (journal_mode=DELETE, gateway mid-turn write, peer in a
subprocess):
before peer BLOCKED -> repair -> peer BLOCKED, holder COMMIT ok
unfixed peer BLOCKED -> repair -> peer STOLE the lock,
holder COMMIT: disk I/O error
WAL is immune (it coordinates through -shm), but DELETE is what Hermes falls
back to on NFS/SMB/FUSE/ZFS and on SQLite builds vulnerable to the WAL-reset
bug, so this is a real deployment shape.
Run the read under offline_file_access and fall back to size:mtime_ns when a
connection is live. That keeps the ledger counting instead of returning None
(which reads as "not exhausted" and would restore the unbounded loop), and the
content key stays load-bearing on the offline repair path -- the only path
where surgery actually runs.
Also fail the free-space guard CLOSED: a nearly-full volume is exactly where
statvfs is likeliest to fail, and proceeding is the multi-GB copy that finishes
off the disk.
Follow-up to adversarial review of the first commit. Three findings, two
confirmed by test and fixed here, one disproven and left alone.
CONFIRMED — the free-space guard was a threshold, not cleanup. Prune runs
only on the success path, so any copy that failed partway (ENOSPC, sidecar
copy failure, kill mid-copy) left a file matching the `malformed-backup-`
prefix that nothing ever removed. Measured on the unpatched tree: backups
capped at 3 while copies succeed, but 13+ and climbing once copy2 raises —
self-reinforcing, since each partial consumes the space that guarantees the
next failure. Worse, partials sort newest-by-name, so a later successful
prune KEPT the garbage and deleted the intact forensic copies.
Fix: copy to a `.incomplete` staging name that does not match the backup
prefix, os.replace into place only after every copy succeeds, unlink staging
on failure, and sweep stale staging debris on entry.
CONFIRMED — the 2GiB floor was a small-volume regression. A 50MB DB on a
10GB volume with 1.5GB free (30x headroom) was refused, and since a refused
backup is a HARD STOP (#69603) that silently converts "repair loops" into
"repair never runs". Fix: require the copy itself (now including its
-wal/-shm sidecars, which the old check ignored) plus proportional headroom
— max(256MiB, 2% of volume).
DISPROVEN — the review claimed a refused backup skips _record_repair_outcome
so the loop never terminates. It does not: repair_state_db_schema records the
outcome on the result returned by _repair_state_db_schema_locked, which is
where the hard stop returns. Verified on a simulated low-disk host: terminal
at pass 4 with zero backups written. No change made.
Tests: 5 new (small-volume allow, proportional headroom, sidecar accounting,
failed-copy leaves no countable debris + staging swept). 23 pass with the
#86747 suite; test_hermes_state.py 252 passed. Pre-existing unrelated
failures unchanged.
A malformed-schema state.db sent Hermes into a repair loop that wrote a
fresh full-size forensic backup every ~10s: 31 copies / 2.3GB in 20
minutes, free space heading to zero on a host running an agent fleet.
The #86747 guards for exactly this were already present and did not hold.
Both keyed on `size:mtime_ns`:
* `_db_fingerprint` -> the ledger's attempt counter reset to 1 on every
pass, so `_MAX_PERSISTENT_REPAIR_ATTEMPTS` was never reached and the
loop never terminated;
* `_backup_db_file`'s dedupe compared mtime, so it never matched and
each pass wrote another full-size copy.
The assumption behind that key -- "nothing can successfully write to a
damaged file" -- holds for the b-tree damage of #86747 but not for the
malformed-SCHEMA class: the DB still opens and accepts writes (only
sqlite_master is unreadable), so live writers, WAL checkpoints and the
in-place repair strategies themselves all move mtime between passes.
Fixes:
* fingerprint on size + a bounded head/tail content sample instead of
mtime. Stable across passes that merely touch the file, still changes
on genuine repair/truncation/restore (so recovery resets the budget),
and stays O(1) on a multi-GB DB.
* dedupe the forensic backup on that same fingerprint.
* add the missing free-space guard: refuse the pre-repair copy when it
would leave under 2GiB free, with an actionable error. The backup is a
full raw copy of the damaged DB, so a repair loop is a disk amplifier
that can take down every process on the host -- and the refusal path
already hard-stops the repair (#69603) rather than mutating the only
remaining copy.
Tests fail on the unfixed tree and pass here; the pre-existing failures in
test_state_db_malformed_repair.py and TestFTS5Search are unrelated and
reproduce on the base commit.