ca28a69ada
state.db corrupted twice in two days with the torn-b-tree signature — repeated "2nd reference to page", "Rowid out of order", and long runs of "never used" pages in messages (rootpage 5) and idx_messages_session. macOS fsync() guarantees neither data-on-platter nor write ordering, which _enforce_macos_synchronous_full already documents: a rewrite interrupted by process or OS termination leaves half-written b-tree pages. The mitigation is per-connection (synchronous=FULL + checkpoint_fullfsync=1) and was applied only through apply_wal_with_fallback(). The repair path opened state.db with a bare sqlite3.connect() six times and then ran REINDEX, VACUUM and writable_schema surgery through it — the operations that rewrite nearly every page of the file — with no barrier at all. - _connect_repair_durable() routes every repair/probe connection through the barriers. Applying them is best-effort by necessity: SQLite loads the schema before any statement, so on a malformed schema even PRAGMA synchronous=FULL raises DatabaseError, and a malformed database is precisely this helper's input. _reapply_durability_barriers() retakes them before REINDEX and VACUUM, once the schema parses and they can stick. - verify_state_db_integrity() adds the proactive check that was missing. Repair only ever ran reactively, after a caller already hit a malformed error, so a database torn in pages no query happened to touch stayed live and kept accepting writes. On 2026-08-19 that gap was 11 hours across two restarts that both reported a clean start. Size-aware: degrades to an O(1) probe above 2 GiB rather than pegging a CPU at startup. Also restores two fixes lost when `hermes update` reset the tree to origin/main before they were committed: - _db_fingerprint keys the repair ledger on dev+inode+size instead of size+mtime_ns. The old form was justified as "stable for a file nothing can successfully write to"; that premise is false, because on FTS corruption this module deliberately keeps canonical writes enabled with FTS detached. mtime churned on every write, so each pass re-keyed the ledger and reset the counter to 1 — the cap could never be reached and the damaging surgery could retry forever. - _live_writer_holds_db() refuses surgery while another connection holds the database. The cross-process lock only serialises repairers against each other; it says nothing about the gateway, Desktop or a CLI. Rewriting b-tree pages under a concurrent writer is what spread the 2026-08-18/19 damage out of the FTS shadow tables and into the canonical ones. Fails open, so it cannot strand the self-heal path it protects. The guard's own tests built a two-table toy schema, so every repair aborted on "no such table: sessions" before reaching the guards under test — the assertions were passing over a code path that never ran. They now build through a real SessionDB. Targeted state/repair suites: 330 passed, 1 pre-existing unrelated failure. Broader sweep: 50 failed/1221 passed -> 46 failed/1225 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Dhanesh Purohit <dhanesh@users.noreply.github.com>