ca28a69ada
state.db corrupted twice in two days with the torn-b-tree signature — repeated "2nd reference to page", "Rowid out of order", and long runs of "never used" pages in messages (rootpage 5) and idx_messages_session. macOS fsync() guarantees neither data-on-platter nor write ordering, which _enforce_macos_synchronous_full already documents: a rewrite interrupted by process or OS termination leaves half-written b-tree pages. The mitigation is per-connection (synchronous=FULL + checkpoint_fullfsync=1) and was applied only through apply_wal_with_fallback(). The repair path opened state.db with a bare sqlite3.connect() six times and then ran REINDEX, VACUUM and writable_schema surgery through it — the operations that rewrite nearly every page of the file — with no barrier at all. - _connect_repair_durable() routes every repair/probe connection through the barriers. Applying them is best-effort by necessity: SQLite loads the schema before any statement, so on a malformed schema even PRAGMA synchronous=FULL raises DatabaseError, and a malformed database is precisely this helper's input. _reapply_durability_barriers() retakes them before REINDEX and VACUUM, once the schema parses and they can stick. - verify_state_db_integrity() adds the proactive check that was missing. Repair only ever ran reactively, after a caller already hit a malformed error, so a database torn in pages no query happened to touch stayed live and kept accepting writes. On 2026-08-19 that gap was 11 hours across two restarts that both reported a clean start. Size-aware: degrades to an O(1) probe above 2 GiB rather than pegging a CPU at startup. Also restores two fixes lost when `hermes update` reset the tree to origin/main before they were committed: - _db_fingerprint keys the repair ledger on dev+inode+size instead of size+mtime_ns. The old form was justified as "stable for a file nothing can successfully write to"; that premise is false, because on FTS corruption this module deliberately keeps canonical writes enabled with FTS detached. mtime churned on every write, so each pass re-keyed the ledger and reset the counter to 1 — the cap could never be reached and the damaging surgery could retry forever. - _live_writer_holds_db() refuses surgery while another connection holds the database. The cross-process lock only serialises repairers against each other; it says nothing about the gateway, Desktop or a CLI. Rewriting b-tree pages under a concurrent writer is what spread the 2026-08-18/19 damage out of the FTS shadow tables and into the canonical ones. Fails open, so it cannot strand the self-heal path it protects. The guard's own tests built a two-table toy schema, so every repair aborted on "no such table: sessions" before reaching the guards under test — the assertions were passing over a code path that never ran. They now build through a real SessionDB. Targeted state/repair suites: 330 passed, 1 pre-existing unrelated failure. Broader sweep: 50 failed/1221 passed -> 46 failed/1225 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Dhanesh Purohit <dhanesh@users.noreply.github.com>
211 lines
8.0 KiB
Python
211 lines
8.0 KiB
Python
"""Regression: state.db repair-path writes must be durable on macOS, and a
|
|
torn database must be detected proactively rather than 11 hours later.
|
|
|
|
Incident (2026-08-19, recurrence of 2026-08-18/19): `state.db` was recovered
|
|
clean at 01:02, tore again in the pages holding rows written 02:18-02:22, and
|
|
the damage went undetected until 13:36 when a write finally landed on a
|
|
damaged page (`append_message failed: constraint failed`). `PRAGMA
|
|
integrity_check` on the file reported the torn-b-tree signature:
|
|
|
|
Tree 5 page 47256 cell 423..429: 2nd reference to page ...
|
|
Tree 5 page 60788 cell 4: Rowid 34637 out of order
|
|
Page 50549..52587: never used
|
|
|
|
Two defects:
|
|
|
|
1. hermes_state already knows macOS `fsync()` does not guarantee write
|
|
ordering, and mitigates it with `synchronous=FULL` +
|
|
`checkpoint_fullfsync=1` (see `_enforce_macos_synchronous_full`, whose
|
|
docstring names this exact failure: "a WAL checkpoint race with process
|
|
termination ... can leave the main DB with half-written btree pages").
|
|
Those pragmas are per-connection and were applied only via
|
|
`apply_wal_with_fallback()`. The repair path opened `state.db` with a bare
|
|
`sqlite3.connect()` five times and then ran REINDEX, VACUUM and
|
|
`writable_schema` surgery through it — the operations that rewrite nearly
|
|
every page of the file — with no barrier at all.
|
|
|
|
2. Repair only ran reactively, when a caller already hit a malformed error
|
|
(`SessionDB` open). Nothing checked the file proactively, so a torn
|
|
database stayed live and accepted writes for hours before anyone noticed.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import re
|
|
import sqlite3
|
|
import sys
|
|
from pathlib import Path
|
|
|
|
import pytest
|
|
|
|
import hermes_state
|
|
from hermes_state import (
|
|
_connect_repair_durable,
|
|
repair_state_db_schema,
|
|
verify_state_db_integrity,
|
|
)
|
|
|
|
|
|
def _make_db(tmp_path: Path) -> Path:
|
|
db = tmp_path / "state.db"
|
|
conn = sqlite3.connect(str(db))
|
|
conn.execute("PRAGMA journal_mode=WAL")
|
|
conn.execute("CREATE TABLE sessions (session_id TEXT PRIMARY KEY)")
|
|
conn.execute("CREATE TABLE messages (id INTEGER PRIMARY KEY, body TEXT)")
|
|
conn.execute("INSERT INTO messages (body) VALUES ('seed')")
|
|
conn.commit()
|
|
conn.close()
|
|
return db
|
|
|
|
|
|
# ── Defect 1: repair-path write durability ──────────────────────────────
|
|
|
|
|
|
def test_connect_repair_durable_sets_macos_barriers(tmp_path: Path) -> None:
|
|
"""The repair connection must carry both macOS durability barriers."""
|
|
db = _make_db(tmp_path)
|
|
conn = _connect_repair_durable(db)
|
|
try:
|
|
synchronous = conn.execute("PRAGMA synchronous").fetchone()[0]
|
|
checkpoint_fullfsync = conn.execute(
|
|
"PRAGMA checkpoint_fullfsync"
|
|
).fetchone()[0]
|
|
finally:
|
|
conn.close()
|
|
|
|
if sys.platform == "darwin":
|
|
# SQLite: 0=OFF, 1=NORMAL, 2=FULL, 3=EXTRA. NORMAL is what tore the
|
|
# b-tree pages; FULL is what _enforce_macos_synchronous_full sets.
|
|
assert synchronous == 2, (
|
|
f"repair connection opened with synchronous={synchronous}; on "
|
|
"Darwin this lets REINDEX/VACUUM leave half-written b-tree pages"
|
|
)
|
|
assert checkpoint_fullfsync == 1, (
|
|
"repair connection has no F_FULLFSYNC barrier at checkpoint "
|
|
"boundaries; macOS fsync() does not flush the drive cache"
|
|
)
|
|
else:
|
|
# Elsewhere the helper is a plain connect — no behaviour change.
|
|
assert synchronous in (0, 1, 2, 3)
|
|
|
|
|
|
def test_connect_repair_durable_is_autocommit(tmp_path: Path) -> None:
|
|
"""Must preserve isolation_level=None — repair runs DDL and VACUUM."""
|
|
db = _make_db(tmp_path)
|
|
conn = _connect_repair_durable(db)
|
|
try:
|
|
assert conn.isolation_level is None
|
|
# VACUUM is only legal outside an implicit transaction.
|
|
conn.execute("VACUUM")
|
|
finally:
|
|
conn.close()
|
|
|
|
|
|
def test_repair_path_has_no_bare_connects() -> None:
|
|
"""No repair/probe site may bypass the durability helper.
|
|
|
|
Source-level guard: the bare form is exactly what regressed, and a unit
|
|
test on the helper alone would not notice a sixth site being added.
|
|
"""
|
|
source = Path(hermes_state.__file__).read_text()
|
|
pattern = r"^\s*conn = sqlite3\.connect\(str\(db_path\), isolation_level=None\)"
|
|
|
|
# The one legitimate bare connect is inside the helper itself; everything
|
|
# after that definition must go through it.
|
|
helper = source.index("def _connect_repair_durable(")
|
|
body_end = source.index("\ndef ", helper + 1)
|
|
inside_helper = re.findall(pattern, source[helper:body_end], flags=re.MULTILINE)
|
|
assert len(inside_helper) == 1, (
|
|
"_connect_repair_durable no longer opens the connection itself"
|
|
)
|
|
|
|
elsewhere = re.findall(
|
|
pattern, source[:helper] + source[body_end:], flags=re.MULTILINE
|
|
)
|
|
assert elsewhere == [], (
|
|
f"{len(elsewhere)} repair-path connection(s) still bypass "
|
|
"_connect_repair_durable() and write state.db without the macOS "
|
|
"fsync barriers"
|
|
)
|
|
|
|
|
|
def test_repair_still_works_through_durable_connection(tmp_path: Path) -> None:
|
|
"""Routing every strategy through the helper must not break the path.
|
|
|
|
The helper is entered once per strategy, so a plumbing fault (recursion,
|
|
a leaked connection, a refused pragma) surfaces as an exception rather
|
|
than a report. Whether this fixture's minimal schema is *repairable* is
|
|
beside the point — the assertion is that the path runs to completion.
|
|
"""
|
|
db = _make_db(tmp_path)
|
|
report = repair_state_db_schema(db, backup=False)
|
|
assert isinstance(report, dict)
|
|
assert set(report) >= {"repaired", "strategy", "backup_path"}
|
|
# The file must still open afterwards — repair may fail, but it must not
|
|
# leave the database less usable than it found it.
|
|
conn = sqlite3.connect(str(db))
|
|
try:
|
|
assert conn.execute("SELECT COUNT(*) FROM messages").fetchone()[0] == 1
|
|
finally:
|
|
conn.close()
|
|
|
|
|
|
# ── Defect 2: proactive verification gate ───────────────────────────────
|
|
|
|
|
|
def test_verify_state_db_integrity_passes_on_healthy_db(tmp_path: Path) -> None:
|
|
db = _make_db(tmp_path)
|
|
result = verify_state_db_integrity(db)
|
|
assert result["ok"] is True
|
|
assert result["problems"] == []
|
|
|
|
|
|
def test_verify_state_db_integrity_detects_torn_btree(tmp_path: Path) -> None:
|
|
"""A torn database must be reported, not silently accepted."""
|
|
db = _make_db(tmp_path)
|
|
# Grow past one page, then corrupt an interior/leaf page directly — the
|
|
# same class of damage as "2nd reference to page" / "never used".
|
|
conn = sqlite3.connect(str(db))
|
|
conn.execute("PRAGMA journal_mode=DELETE")
|
|
conn.executemany(
|
|
"INSERT INTO messages (body) VALUES (?)",
|
|
[(f"row-{i}" * 40,) for i in range(500)],
|
|
)
|
|
conn.commit()
|
|
conn.close()
|
|
|
|
page_size = 4096
|
|
raw = bytearray(db.read_bytes())
|
|
# Scribble over a data page (page 3+), leaving the header page intact so
|
|
# the file still opens — that is what makes this class so long-lived.
|
|
start = page_size * 4
|
|
raw[start:start + page_size] = b"\xff" * page_size
|
|
db.write_bytes(bytes(raw))
|
|
|
|
result = verify_state_db_integrity(db)
|
|
assert result["ok"] is False
|
|
assert result["problems"], "torn pages reported no problems"
|
|
|
|
|
|
def test_verify_state_db_integrity_skips_pragma_when_oversized(
|
|
tmp_path: Path,
|
|
) -> None:
|
|
"""Must degrade to an O(1) probe rather than pegging a CPU for minutes.
|
|
|
|
`PRAGMA integrity_check` walks every page, so an unbounded check at
|
|
startup would hang the gateway on a multi-GB state.db.
|
|
"""
|
|
db = _make_db(tmp_path)
|
|
result = verify_state_db_integrity(db, max_bytes=1)
|
|
assert result["ok"] is True
|
|
assert result["checked"] == "probe"
|
|
|
|
|
|
def test_verify_state_db_integrity_missing_file_is_not_a_failure(
|
|
tmp_path: Path,
|
|
) -> None:
|
|
"""A first run has no state.db yet; that must not look like corruption."""
|
|
result = verify_state_db_integrity(tmp_path / "absent.db")
|
|
assert result["ok"] is True
|
|
assert result["checked"] == "absent"
|