Commit Graph

3 Commits

Author SHA1 Message Date
kshitijk4poor 29c9066577 chore(state): tidy post-salvage residue in state.db durability test
Two leftovers from the #91852 descope (integrity-check tests removed but their
scaffolding stayed):

- Drop the now-unused `import pytest` (orphaned when the verify_state_db_integrity
  tests that used it were removed; no markers/raises/fixtures remain in this file).
- Rename the `# Defect 1:` section label to just `# Repair-path write durability`
  — the sibling "Defect 2" section was descoped out, leaving the numbering dangling.

Test-only, no behavior change. tests/test_state_db_write_durability.py: 4 passed.
2026-08-22 04:13:17 +05:30
kshitijk4poor 3bdc2165c3 fix(state): scope salvage to repair-connection durability + live-writer guard
Follow-up to the salvaged repair-durability commit. Scope corrections so this
PR ships only the reachable, non-competing, WAL-mode-correct half:

- Drop verify_state_db_integrity() + its 4 tests. Zero production callers here
  (dead code); the caller lives in the follow-up that wires it into
  SessionStore._open_session_db_for_active_scope() (PR #91754). The function
  moves with its wiring.
- Drop the _db_fingerprint change (size:mtime_ns -> dev:ino:size) + its 3
  ledger tests. This is competing work: PR #88425 (salvage of @jirathip-k's
  #88224) already fixes the same size:mtime_ns budget-reset bug with a
  content-sample + volatile-header-mask that also handles the DELETE-mode
  commit-counter case, and carries @jirathip-k's diagnosis/credit. Landing a
  second, divergent fingerprint contract would stomp that lineage. Fingerprint
  stays with #88425; this PR reverts _db_fingerprint to main's form.
- Mark test_repair_refuses_while_another_connection_holds_the_db requires_wal.
  _live_writer_holds_db detects an out-of-process holder via the WAL-index
  exclusive lock, absent in journal_mode=DELETE (used on WAL-reset-vulnerable
  SQLite <3.51.3 incl. CI's 3.50.4, and on NFS/SMB). The test failed there;
  the conftest requires_wal gate auto-skips it. DELETE-mode limitation is now
  documented on the guard docstring: repair is serialised only by the
  cross-process repairer lock there. The reported incident was in WAL mode.
- Map dhanesh@users.noreply.github.com -> dhanesh (contributors/emails) so the
  attribution CI gate passes.

Net: this PR is repair-connection durability barriers + the live-writer guard.
addresses @andrexibiza's #90747 review (dead-code verifier + fingerprint
interlock with #88425).
2026-08-22 04:02:29 +05:30
Dhanesh Purohit ca28a69ada fix(state): apply macOS write barriers on every state.db repair connection
state.db corrupted twice in two days with the torn-b-tree signature —
repeated "2nd reference to page", "Rowid out of order", and long runs of
"never used" pages in messages (rootpage 5) and idx_messages_session.

macOS fsync() guarantees neither data-on-platter nor write ordering, which
_enforce_macos_synchronous_full already documents: a rewrite interrupted by
process or OS termination leaves half-written b-tree pages. The mitigation
is per-connection (synchronous=FULL + checkpoint_fullfsync=1) and was
applied only through apply_wal_with_fallback(). The repair path opened
state.db with a bare sqlite3.connect() six times and then ran REINDEX,
VACUUM and writable_schema surgery through it — the operations that rewrite
nearly every page of the file — with no barrier at all.

- _connect_repair_durable() routes every repair/probe connection through the
  barriers. Applying them is best-effort by necessity: SQLite loads the
  schema before any statement, so on a malformed schema even
  PRAGMA synchronous=FULL raises DatabaseError, and a malformed database is
  precisely this helper's input. _reapply_durability_barriers() retakes them
  before REINDEX and VACUUM, once the schema parses and they can stick.
- verify_state_db_integrity() adds the proactive check that was missing.
  Repair only ever ran reactively, after a caller already hit a malformed
  error, so a database torn in pages no query happened to touch stayed live
  and kept accepting writes. On 2026-08-19 that gap was 11 hours across two
  restarts that both reported a clean start. Size-aware: degrades to an O(1)
  probe above 2 GiB rather than pegging a CPU at startup.

Also restores two fixes lost when `hermes update` reset the tree to
origin/main before they were committed:

- _db_fingerprint keys the repair ledger on dev+inode+size instead of
  size+mtime_ns. The old form was justified as "stable for a file nothing
  can successfully write to"; that premise is false, because on FTS
  corruption this module deliberately keeps canonical writes enabled with
  FTS detached. mtime churned on every write, so each pass re-keyed the
  ledger and reset the counter to 1 — the cap could never be reached and the
  damaging surgery could retry forever.
- _live_writer_holds_db() refuses surgery while another connection holds the
  database. The cross-process lock only serialises repairers against each
  other; it says nothing about the gateway, Desktop or a CLI. Rewriting
  b-tree pages under a concurrent writer is what spread the 2026-08-18/19
  damage out of the FTS shadow tables and into the canonical ones. Fails
  open, so it cannot strand the self-heal path it protects.

The guard's own tests built a two-table toy schema, so every repair aborted
on "no such table: sessions" before reaching the guards under test — the
assertions were passing over a code path that never ran. They now build
through a real SessionDB.

Targeted state/repair suites: 330 passed, 1 pre-existing unrelated failure.
Broader sweep: 50 failed/1221 passed -> 46 failed/1225 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Signed-off-by: Dhanesh Purohit <dhanesh@users.noreply.github.com>
2026-08-22 04:02:29 +05:30