fix(dashboard): startup schema reconcile opens state.db read-only first

The dashboard's `_eager_reconcile_own_session_db` did an unconditional
writable `acquire()` at every startup. When the gateway shares that
state.db the dashboard became a second long-lived writable SessionDB
owner: a close-time WAL checkpoint plus a possible FTS rebuild in
`_init_fts`, the two-writer vector behind the corruption reports in
#107688 and #100896 ("5 live SessionDB handles" precursor, gateway +
dashboard both holding the WAL).

Route the startup reconcile through `_open_session_db_at_path(...,
read_only=True)`, which already bootstraps a missing store and heals a
stale/malformed schema through exactly ONE writable open before
reopening read-only. A healthy store now gets zero writable opens from
the dashboard while the #79531/#80037 "bring schema current before the
first poll" contract is kept (existing heal test unchanged).

Live repro (healthy store, count writable SessionDB.__init__ calls from
the startup worker): before=1 after=0.

Reported-by: #107688, #100896 (@kokhlo diagnosis)
Refs #107688 #100896
This commit is contained in:
Teknium
2026-09-11 02:02:17 -07:00
parent b02f3aa00b
commit c9b71ebaf7
2 changed files with 38 additions and 7 deletions
+12 -7
View File
@@ -171,18 +171,23 @@ def _resolve_restart_drain_timeout() -> float:
def _eager_reconcile_own_session_db() -> None:
"""One writable open of this process's own state.db at startup.
"""Bring this process's own state.db schema current at startup — read-only first.
``SessionDB.__init__`` runs ``_init_schema`` → ``_reconcile_columns`` with
open-time lock patience. Never raises: an unfixable store still gets the
per-poll read-probe heal in :func:`_open_session_db_at_path`.
The dashboard is a view layer; the gateway owns the writer. A healthy store
must never see a second writable ``SessionDB`` from this process (its
close-time checkpoint and a possible FTS rebuild in ``_init_fts`` are the
two-writer corruption vector, #107688 / #100896). The read-only path still
bootstraps a missing store and heals a stale/malformed schema through ONE
writable open, so the #79531 contract holds. Never raises: an unfixable
store still gets the per-poll read-probe heal.
"""
try:
from hermes_state import _default_db_path
from hermes_state_registry import acquire, release_or_close
db = acquire(Path(_default_db_path()))
release_or_close(db)
from hermes_cli.web_server_sessions import _open_session_db_at_path
db = _open_session_db_at_path(Path(_default_db_path()), read_only=True)
db.close()
except Exception as exc:
_log.warning(
"startup schema reconcile of state.db failed (%s); session "
+26
View File
@@ -566,6 +566,32 @@ class TestWebServerEndpoints:
db.close()
assert [r["id"] for r in rows] == ["eager-stale"]
def test_startup_eager_reconcile_is_read_only_on_a_healthy_store(self, monkeypatch):
"""A current-schema store gets NO writable open from the dashboard (#107688).
The gateway owns the writer; a second writable SessionDB from the
dashboard (close-time checkpoint, possible FTS rebuild) is the
two-writer corruption vector. Only the stale-schema heal may write.
"""
import hermes_state
from hermes_constants import get_hermes_home
from hermes_state import SessionDB
SessionDB(db_path=get_hermes_home() / "state.db").close()
writable_opens = []
real_init = SessionDB.__init__
def spy(self, *args, **kwargs):
if not kwargs.get("read_only"):
writable_opens.append(kwargs)
return real_init(self, *args, **kwargs)
monkeypatch.setattr(hermes_state.SessionDB, "__init__", spy)
_web_server_lifecycle._eager_reconcile_own_session_db()
assert writable_opens == []
def test_startup_eager_reconcile_never_raises(self, monkeypatch):
"""A store the eager reconcile cannot open must not break startup."""
import sqlite3 as sqlite3_module