fix(state): fail fast on non-contention flock errors and retry deferred FTS rebuilds in-process (salvage #100130)

Two pieces of PR #100130 (@HexLab98) re-applied on top of the orphaned-flock
break (894fc35337) and fail-closed admission (#100895) that landed since:

* `is_advisory_lock_contention` (hermes_state_common): only EAGAIN /
  EWOULDBLOCK / EACCES / EDEADLK mean "another process holds the lock".
  ESTALE / ENOTSUP / ENOLCK / EIO from flock or msvcrt.locking are
  environment failures that polling cannot fix — `_acquire_db_flock` and
  both Windows msvcrt loops (FTS rebuild admission, state.db repair lock)
  now defer immediately with the real errno instead of burning the full
  120s / holder timeout and then logging a fake "held by another process".

* `retry_deferred_fts_recovery` (hermes_state_schema): a SessionDB whose
  open-time `_recover_stale_fts` deferred (foreign holders or busy rebuild
  lock) stayed `_fts_stale` — LIKE-only search — until the process
  reopened state.db. Short-lived CLIs reopen every run; the gateway opens
  once and stays up for days, so the deferral was effectively permanent
  (#100108). The retry runs from the EXISTING gateway housekeeping tick
  (`_start_gateway_housekeeping`, 60s) against the shared SessionDB
  instances via `hermes_state_registry.live_shared_session_dbs()`:
  non-blocking admission (`fts_rebuild_admission(timeout_seconds=0)`),
  bounded backoff 60s -> 1h, no new thread, still fails closed on live
  holders. `fts_rebuild_admission` gains the `timeout_seconds` kwarg.

* WAL-reset warning names `sys.executable` so a "linked SQLite 3.45.1"
  line can be matched to the interpreter that actually linked it
  (#100108 point 3).

Deliberately NOT carried from #100130: the "leftover lock file = holder"
premise (a 0-byte lock file never blocked flock; the real cause was the
fork-inherited fd, fixed in 894fc35337) and the `_rebuild_fts_once`
one-shot rework.

Co-authored-by: HexLab98 <liruixinch@outlook.com>
This commit is contained in:
teknium1
2026-09-02 03:57:36 -07:00
committed by Teknium
parent 238b6c1ab9
commit fd05029430
6 changed files with 231 additions and 25 deletions
+81 -4
View File
@@ -43,6 +43,15 @@ logger = logging.getLogger("hermes_state")
_FTS_HOLDER_ESCALATE_ATTEMPTS = 3
_FTS_HOLDER_ESCALATE_SECONDS = 60.0
# Minimum spacing between in-process retries of a deferred stale-FTS rebuild
# (``retry_deferred_fts_recovery``). The startup open already paid the full
# admission wait once; later retries are non-blocking probes on this cadence
# so a live holder never stalls a long-lived writer.
_FTS_STALE_RETRY_SECONDS = 60.0
# Each failed retry doubles the spacing up to this cap, so a holder that never
# goes away (a second long-lived writer) costs one deferral warning per hour,
# not one per minute. A successful rebuild clears the stale state entirely.
_FTS_STALE_RETRY_MAX_SECONDS = 3600.0
# Cache for schema_read_probe_statements() — parsing SCHEMA_SQL spins up an
# in-memory SQLite database, so derive the statements once per process.
@@ -422,8 +431,14 @@ class SessionSchemaMixin:
)
return None
def _recover_stale_fts(self, cursor: sqlite3.Cursor, *, legacy: bool) -> bool:
"""Atomically rebuild stale base/trigram indexes and resume syncing."""
def _recover_stale_fts(
self, cursor: sqlite3.Cursor, *, legacy: bool, timeout_seconds=None
) -> bool:
"""Atomically rebuild stale base/trigram indexes and resume syncing.
*timeout_seconds* bounds the cross-process admission wait; None uses
the full startup budget, ``0`` is the non-blocking in-process retry.
"""
foreign_holders = self._foreign_state_db_holders()
if foreign_holders:
now = time.time()
@@ -502,7 +517,9 @@ class SessionSchemaMixin:
# authority (fail closed). Losing the race means another process is
# already performing this exact recovery; the stale breadcrumb stays
# set, so this process simply keeps FTS detached and retries later.
with fts_rebuild_admission(getattr(self, "db_path", None)) as admitted:
with fts_rebuild_admission(
getattr(self, "db_path", None), timeout_seconds=timeout_seconds
) as admitted:
if not admitted:
logger.warning(
"Deferred stale state.db FTS rebuild: another process "
@@ -512,6 +529,65 @@ class SessionSchemaMixin:
return False
return self._recover_stale_fts_locked(cursor, legacy=legacy)
def retry_deferred_fts_recovery(self) -> bool:
"""Retry a deferred stale-FTS rebuild on this open SessionDB.
``_recover_stale_fts`` runs at open and fails closed when foreign
holders or the rebuild lock are busy, leaving ``_fts_stale`` set and
search on the LIKE fallback. Live write/search paths must never start
a full rebuild (#97940), so on a short-lived CLI that deferral is
cleared by the next process open — but a gateway opens state.db
once and stays up for days, so "next open" never came (#100108).
This is the in-process retry: bounded backoff from
``_FTS_STALE_RETRY_SECONDS`` doubling to ``_FTS_STALE_RETRY_MAX_SECONDS``,
non-blocking admission (``timeout=0``) so a live holder is skipped and
tried again later, no new thread — the caller is an existing periodic
tick (gateway housekeeping).
Returns True only when the index was rebuilt and sync triggers
restored. Never raises.
"""
if not getattr(self, "_fts_stale", False):
return False
if getattr(self, "read_only", False) or getattr(self, "_conn", None) is None:
return False
now = time.monotonic()
if now < getattr(self, "_fts_stale_retry_after", 0.0):
return False
interval = float(
getattr(self, "_fts_stale_retry_interval", 0.0)
) or _FTS_STALE_RETRY_SECONDS
self._fts_stale_retry_after = now + interval
self._fts_stale_retry_interval = min(
interval * 2.0, _FTS_STALE_RETRY_MAX_SECONDS
)
try:
with self._lock:
if self._conn is None or not self._fts_stale:
return False
cursor = self._conn.cursor()
legacy = self._db_has_legacy_inline_fts(cursor)
recovered = self._recover_stale_fts(
cursor, legacy=legacy, timeout_seconds=0.0
)
if recovered:
# CJK was detached alongside the base indexes; its own
# ensure path decides when it comes back online.
self._ensure_fts_cjk_schema(cursor)
self._fts_stale_retry_interval = 0.0
try:
self._conn.commit()
except sqlite3.Error:
pass
return recovered
except Exception: # noqa: BLE001 - background retry must never raise
logger.warning(
"In-process retry of the deferred stale state.db FTS rebuild "
"failed; will retry later.",
exc_info=True,
)
return False
def _recover_stale_fts_locked(
self, cursor: sqlite3.Cursor, *, legacy: bool
) -> bool:
@@ -1510,7 +1586,8 @@ class SessionSchemaMixin:
breadcrumb is persisted, mirroring ``_enter_fts_fail_open``'s
ordering contract: triggers must never be live over an index with an
unrebuilt gap. FTS stays detached for this instance; the winner's
rebuild — or ``_recover_stale_fts`` at the next startup — restores
rebuild — or ``retry_deferred_fts_recovery`` from the gateway
housekeeping tick, or ``_recover_stale_fts`` at the next startup — restores
the index and triggers atomically.
"""
with fts_rebuild_admission(getattr(self, "db_path", None)) as admitted: