Commit Graph

9 Commits

Author SHA1 Message Date
Teknium 53db597201 simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
2026-09-03 13:46:50 -07:00
kshitijk4poor 1fe8683e58 fix(state): split forensic-backup identity from repair-epoch fingerprint; publish backup bundle atomically
Addresses two data-integrity gaps @andrexibiza flagged reviewing #88425.

1. Forensic dedupe no longer reuses the repair-epoch fingerprint.
   _db_fingerprint masks SQLite's commit counters and samples only head/tail
   so an ordinary write does not re-key the repair budget — the right
   predicate for 'same damage epoch', the WRONG one for 'same recovery
   image'. A live writer committing rows into an interior page (size
   preserved, head/tail untouched) collided under it, so _backup_db_file
   handed back a STALE backup that predates real user data. New
   _backup_content_identity() digests the whole file + every sidecar; the
   dedupe uses it. The O(n) read is cheaper than the O(n) copy it avoids on a
   hit.

2. Backup bundle is now published atomically. The promotion loop replaced
   files one at a time (main first) and cleanup unlinked only staging srcs,
   so a sidecar os.replace failure after the main promotion left the
   final-prefix main backup on disk — a countable-but-incomplete bundle that
   passed the #69603 hard stop and deduped as legitimate next pass. Now
   sidecars publish first and the main DB last (its name is the commit
   marker _existing_malformed_backups counts), and cleanup rolls back every
   already-published destination.

Two regressions added (both mutation-checked — each fails on pre-fix code):
- test_backup_not_deduped_after_interior_page_write
- test_publication_failure_leaves_no_countable_partial_bundle

tests/test_state_db_repair_loop_mtime.py: 28 passed.
2026-08-22 14:33:20 +05:30
kshitij 5777e68b3d fix(state): include the rollback journal in the forensic backup
The pre-repair copy took only -wal/-shm. In rollback-journal (DELETE) mode --
Hermes's fallback on NFS/SMB/FUSE/ZFS and on WAL-reset-vulnerable SQLite builds
-- a hot <db>-journal exists on disk whenever a transaction was open, and that
file is what rolls the damaged bytes back to a consistent state. A forensic copy
without it cannot be recovered by hand, which is the entire purpose of taking
the copy before destructive surgery.

Verified the journal is really there:

  files while a txn is open: ['state.db', 'state.db-journal']
  files after commit:        ['state.db']

Add _DB_SIDECAR_SUFFIXES = ("-wal", "-shm", "-journal") and use it at the four
sites that must agree: the disk-guard sizing, the staging copy, the
backup-count exclusion in _existing_malformed_backups (so a copied journal is
not itself counted as a forensic backup), and _prune_malformed_backups (which
otherwise leaks one journal per pruned backup, quietly defeating the retention
cap this PR is partly about).

Matches the spelling hermes_cli/session_recovery.py:61 already uses for the
same concept.
2026-08-22 14:33:20 +05:30
kshitij 8779b782b3 fix(state): exclude SQLite's commit counters from the repair fingerprint
Third self-review pass found the content fingerprint was still defeated on
rollback-journal deployments, by the same mechanism as the original mtime bug.

The head sample starts at byte 0, so it covers the database header's file
change counter (bytes 24-27) and version-valid-for (92-95). In DELETE mode a
commit writes the main file directly and bumps both. A malformed-SCHEMA DB
still accepts writes -- that is the whole premise of this PR -- so any ordinary
session write between passes re-keyed the ledger:

  DELETE, 18MB db, one peer UPDATE between passes (before this commit)
    pass 1..6: attempts=1 every pass, exhausted=False -> unbounded loop

  after
    pass 1..3: attempts=1,2,3   pass 4: BLOCKED

WAL is unaffected (commits land in -wal; the main header only moves on
checkpoint), so this was invisible on a WAL host and reproducible on every
NFS/SMB/FUSE/ZFS or WAL-reset-vulnerable host -- exactly the deployments the
earlier lock-safety commit was written for.

Mask the two volatile ranges out of the sample. Page 1's sqlite_master b-tree
sits after byte 100 and stays in, so genuine recovery still resets the budget:
verified schema rewrite, index rebuild, VACUUM and truncation all change the
key, while a bare utime and an ordinary commit do not.

Test-cost cleanup in the same file, since the new tests needed a
larger-than-sample fixture and the file was already slow:
  - the two guard tests that allocated 450MB of os.urandom now use sparse
    truncate (both only ever read st_size), and the new fixtures use 600 rows
    rather than 40k;
  - file runtime 127s -> 35s.
2026-08-22 14:33:20 +05:30
kshitij 602c45e45e fix(state): never let a peer connection reset the repair budget
Self-review of the previous commit found it reintroduced the bug this PR
exists to fix, by a different route.

`_db_fingerprint` fell back to `size:mtime_ns` when a live connection made the
content read unsafe. The ledger compares keys for EQUALITY, and the two keys
have different SHAPES, so a gateway peer connecting between passes flipped the
shape and the counter reset to 1 every time:

  pass 1 [offline] attempts=1  fp=8192:58c7924f0fba...
  pass 2 [LIVE   ] attempts=1  fp=8192:1786972039271402096
  pass 3 [offline] attempts=1  fp=8192:58c7924f0fba...
  ... never reaches _MAX_PERSISTENT_REPAIR_ATTEMPTS

Return None instead, and teach the two ledger helpers to cope:

- `_persistent_repair_attempts_exhausted` falls back to the recorded key's
  SIZE prefix (the one component both shapes share and that needs no raw
  read) rather than reading as "not exhausted" — otherwise a peer connection
  hides an exhausted budget on every pass, same loop.
- `_record_repair_outcome` keeps the key already on record and still
  increments, rather than dropping the pass.

  pass 1 [offline] attempts=1  pass 2 [LIVE] attempts=2
  pass 3 [offline] attempts=3  pass 4 [LIVE] BLOCKED

Intra-pass flips were already safe (the probe and the record are both reached
with the same liveness within one `repair_state_db_schema` call); it is the
cross-pass change that desynced.

Also drops two `type: ignore` directives `ty` flagged as unused, and replaces
the `LiveConnectionError = ()` / `nullcontext()` shim with a real no-op
contextmanager + exception class so the scaffold-install path is honest.
2026-08-22 14:33:20 +05:30
kshitij 8de64b1634 fix(state): stop backup staging from posing as a forensic copy
The staging name was derived from the backup name
(`<db>.malformed-backup-<stamp>.incomplete`), which still matches the prefix
`_existing_malformed_backups` selects on -- it excludes only `-wal`/`-shm`.
Three consequences, all reproduced:

  - it is COUNTED as a forensic backup;
  - it sorts NEWEST (`.incomplete` > the bare stamp), so prune's
    keep-3-newest slice retained partials and deleted intact copies -- the
    exact inversion the staging change was meant to prevent;
  - worst, the dedupe ran BEFORE the sweep, and a staging file orphaned by a
    kill mid-copy is a byte-identical copy of the damaged DB, so its
    fingerprint MATCHES and it was handed back as the official `backup_path`.
    Repair then passed the #69603 hard-stop gate and ran destructive surgery
    believing a forensic copy existed, and the next pass's sweep deleted that
    very file.

Move staging outside the prefix (`<db>.backup-staging-<stamp>`) and sweep
before the dedupe. The sweep also matches the pre-merge `.incomplete`
spelling so a host that ran the earlier build does not keep prefix-matching
debris that sorts newest and survives prune forever.

Before / after on the same fixture (orphaned staging + a later pass):

  before  backup_path = ...malformed-backup-<stamp>.incomplete   (staging!)
          pass-1 forensic copy deleted by the next sweep
  after   backup_path = ...malformed-backup-<stamp>              (real copy)
          debris swept, pass-1 forensic copy preserved
2026-08-22 14:33:20 +05:30
kshitij b3f14c8534 fix(state): keep the repair fingerprint from cancelling POSIX advisory locks
The content fingerprint takes a raw descriptor, and close() on ANY descriptor
cancels every POSIX advisory lock the process holds on that file. The
exhaustion probe runs before _backup_db_file's has_live_connection guard, so
the read happened even when a peer SessionDB held a write lock.

Verified end-to-end (journal_mode=DELETE, gateway mid-turn write, peer in a
subprocess):

  before   peer BLOCKED -> repair -> peer BLOCKED, holder COMMIT ok
  unfixed  peer BLOCKED -> repair -> peer STOLE the lock,
                                    holder COMMIT: disk I/O error

WAL is immune (it coordinates through -shm), but DELETE is what Hermes falls
back to on NFS/SMB/FUSE/ZFS and on SQLite builds vulnerable to the WAL-reset
bug, so this is a real deployment shape.

Run the read under offline_file_access and fall back to size:mtime_ns when a
connection is live. That keeps the ledger counting instead of returning None
(which reads as "not exhausted" and would restore the unbounded loop), and the
content key stays load-bearing on the offline repair path -- the only path
where surgery actually runs.

Also fail the free-space guard CLOSED: a nearly-full volume is exactly where
statvfs is likeliest to fail, and proceeding is the multi-GB copy that finishes
off the disk.
2026-08-22 14:33:20 +05:30
jirathip-k c914a9ac4b fix(state): make backup atomic and the disk guard proportional
Follow-up to adversarial review of the first commit. Three findings, two
confirmed by test and fixed here, one disproven and left alone.

CONFIRMED — the free-space guard was a threshold, not cleanup. Prune runs
only on the success path, so any copy that failed partway (ENOSPC, sidecar
copy failure, kill mid-copy) left a file matching the `malformed-backup-`
prefix that nothing ever removed. Measured on the unpatched tree: backups
capped at 3 while copies succeed, but 13+ and climbing once copy2 raises —
self-reinforcing, since each partial consumes the space that guarantees the
next failure. Worse, partials sort newest-by-name, so a later successful
prune KEPT the garbage and deleted the intact forensic copies.
Fix: copy to a `.incomplete` staging name that does not match the backup
prefix, os.replace into place only after every copy succeeds, unlink staging
on failure, and sweep stale staging debris on entry.

CONFIRMED — the 2GiB floor was a small-volume regression. A 50MB DB on a
10GB volume with 1.5GB free (30x headroom) was refused, and since a refused
backup is a HARD STOP (#69603) that silently converts "repair loops" into
"repair never runs". Fix: require the copy itself (now including its
-wal/-shm sidecars, which the old check ignored) plus proportional headroom
— max(256MiB, 2% of volume).

DISPROVEN — the review claimed a refused backup skips _record_repair_outcome
so the loop never terminates. It does not: repair_state_db_schema records the
outcome on the result returned by _repair_state_db_schema_locked, which is
where the hard stop returns. Verified on a simulated low-disk host: terminal
at pass 4 with zero backups written. No change made.

Tests: 5 new (small-volume allow, proportional headroom, sidecar accounting,
failed-copy leaves no countable debris + staging swept). 23 pass with the
#86747 suite; test_hermes_state.py 252 passed. Pre-existing unrelated
failures unchanged.
2026-08-22 14:33:20 +05:30
jirathip-k 27d661e171 fix(state): stop unbounded state.db repair loop from filling the disk
A malformed-schema state.db sent Hermes into a repair loop that wrote a
fresh full-size forensic backup every ~10s: 31 copies / 2.3GB in 20
minutes, free space heading to zero on a host running an agent fleet.

The #86747 guards for exactly this were already present and did not hold.
Both keyed on `size:mtime_ns`:

  * `_db_fingerprint` -> the ledger's attempt counter reset to 1 on every
    pass, so `_MAX_PERSISTENT_REPAIR_ATTEMPTS` was never reached and the
    loop never terminated;
  * `_backup_db_file`'s dedupe compared mtime, so it never matched and
    each pass wrote another full-size copy.

The assumption behind that key -- "nothing can successfully write to a
damaged file" -- holds for the b-tree damage of #86747 but not for the
malformed-SCHEMA class: the DB still opens and accepts writes (only
sqlite_master is unreadable), so live writers, WAL checkpoints and the
in-place repair strategies themselves all move mtime between passes.

Fixes:

  * fingerprint on size + a bounded head/tail content sample instead of
    mtime. Stable across passes that merely touch the file, still changes
    on genuine repair/truncation/restore (so recovery resets the budget),
    and stays O(1) on a multi-GB DB.
  * dedupe the forensic backup on that same fingerprint.
  * add the missing free-space guard: refuse the pre-repair copy when it
    would leave under 2GiB free, with an actionable error. The backup is a
    full raw copy of the damaged DB, so a repair loop is a disk amplifier
    that can take down every process on the host -- and the refusal path
    already hard-stops the repair (#69603) rather than mutating the only
    remaining copy.

Tests fail on the unfixed tree and pass here; the pre-existing failures in
test_state_db_malformed_repair.py and TestFTS5Search are unrelated and
reproduce on the base commit.
2026-08-22 14:33:20 +05:30