fix(state): quarantine SessionDB handle after structural corruption

A bare SQLITE_CORRUPT/NOTADB on a live write (not FTS-scoped, not a
replaced file) now sets a sticky per-instance flag: later writes fail
fast with StateDbCorruptError, the handle never reopens after close(),
and close() skips its explicit PASSIVE WAL checkpoint. Gateway and agent
flush paths divert pending transcripts to JSONL/spool like the replaced
case instead of retrying forever.

Field evidence: a handle that kept writing for ~50 minutes after the
first structural error checkpointed 15 pages under the wrong page
numbers on shutdown (page 1 <- messages_fts_trigram_data leaf), turning
"malformed" into "file is not a database".

Refs #90837, #90950, #97940, #89332, #45383

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CNX8rNYHqA5pT4tAGSzXtb
This commit is contained in:
leomcamilo
2026-09-02 05:22:41 -03:00
committed by kshitij
parent d8616f1c88
commit bcc2e65818
10 changed files with 565 additions and 11 deletions
+7 -2
View File
@@ -2670,13 +2670,17 @@ class AIAgent:
# ("storage was busy, send it again") from disk-full/read-only.
from hermes_state import (
CompressionSessionClosedError,
StateDbCorruptError,
StateDbReplacedError,
classify_persistence_error,
divert_session_transcript_jsonl,
)
self._last_persistence_error_cause = classify_persistence_error(e)
if isinstance(e, StateDbReplacedError):
if isinstance(e, (StateDbReplacedError, StateDbCorruptError)):
# Replaced generation or quarantined (structurally corrupt)
# handle: SQLite will not take this batch again, so keep it
# on disk instead of only in RAM.
try:
divert_session_transcript_jsonl(
getattr(self, "session_id", "") or "",
@@ -2684,7 +2688,8 @@ class AIAgent:
)
except Exception:
logger.warning(
"JSONL divert failed after state.db replace for %s",
"JSONL divert failed after state.db %s for %s",
self._last_persistence_error_cause,
getattr(self, "session_id", None),
exc_info=True,
)