fix(state): second-process maintenance on state.db refuses ANY foreign holder

`hermes doctor --fix`'s WAL checkpoint and `repair_state_db_schema`'s
preflight documented themselves as fail-OPEN: `live_writer_holds_db` only
refused on unknown/deleted/uninspectable holders and then trusted a
`BEGIN IMMEDIATE` probe, which is blind to a `journal_mode=DELETE` reader
(SHARED only) and cannot run on a malformed file — exactly the states repair
and checkpoint get invoked in. A repair in a second process then REINDEXed /
VACUUMed a file the gateway still held (#103339 item 2).

- `hermes_state_holders.live_writer_holds_db`: any foreign holder of the DB or
  a sidecar is a live holder; the probe is only an additional positive signal.
- doctor `--fix`: the checkpoint runs on `_exclusive_repair_db_guard`'s
  connection instead of a bare writable `sqlite3.connect`, so an opener
  arriving after the scan is refused, not joined; `_session_count` is a
  `mode=ro` reader.
- Normal SessionDB writers are untouched: gateway + dashboard in two processes
  both keep writing (a process-wide flock on the write path — PR #109270's
  shape — would break that).

Tests: the two-process repair race test releases the test process's own
header-probe fd (it is a genuine holder now); the mid-repair writer fixture
opens its connection after staging starts (a pre-existing holder is refused up
front, which is the point).

Refs #103339 #100896
This commit is contained in:
teknium1
2026-09-14 07:20:38 -07:00
committed by Teknium
parent 4da41c92b4
commit 12173db5b7
7 changed files with 162 additions and 41 deletions
+5 -8
View File
@@ -824,15 +824,12 @@ def _db_opens_cleanly(db_path: Path) -> Optional[str]:
def _live_writer_holds_db(db_path: Path) -> bool:
"""True when a connection outside this call still holds ``db_path`` open.
"""True when another process (or a connection outside this call) still holds ``db_path``.
Asks SQLite for what a repair needs and a live holder cannot grant: ``locking_mode=EXCLUSIVE`` then
``BEGIN IMMEDIATE`` — in WAL mode that needs exclusive WAL-index locks, so any other open connection fails
it with SQLITE_BUSY; neither statement parses the schema, so it works on malformed DBs. Fails **open**
(False) on anything but a positive busy/locked signal: refusing to repair a DB nobody holds would strand
the self-heal path. In ``journal_mode=DELETE`` a held reader takes only SHARED and this returns False;
repair is then serialised only by the cross-process repairer lock. Before probing, the foreign-holder scan
(``hermes_state_holders``) fails closed on deleted-WAL-generation, uninspectable, or unknown holders."""
The foreign-holder scan (``hermes_state_holders``) is the authority: any other process with the DB or a
WAL sidecar open, a deleted WAL generation, or an unknown/uninspectable holder fails CLOSED. The SQLite
probe (``locking_mode=EXCLUSIVE`` + ``BEGIN IMMEDIATE``) is only an additional positive signal — it cannot
see a ``journal_mode=DELETE`` reader and cannot run on a malformed file, which is why the scan comes first."""
import hermes_state_holders as _state_holders
return _state_holders.live_writer_holds_db(db_path, connect_repair_durable=_connect_repair_durable)