fix(state): fail fast on non-contention flock errors and retry deferred FTS rebuilds in-process (salvage #100130)
Two pieces of PR #100130 (@HexLab98) re-applied on top of the orphaned-flock break (894fc35337) and fail-closed admission (#100895) that landed since: * `is_advisory_lock_contention` (hermes_state_common): only EAGAIN / EWOULDBLOCK / EACCES / EDEADLK mean "another process holds the lock". ESTALE / ENOTSUP / ENOLCK / EIO from flock or msvcrt.locking are environment failures that polling cannot fix — `_acquire_db_flock` and both Windows msvcrt loops (FTS rebuild admission, state.db repair lock) now defer immediately with the real errno instead of burning the full 120s / holder timeout and then logging a fake "held by another process". * `retry_deferred_fts_recovery` (hermes_state_schema): a SessionDB whose open-time `_recover_stale_fts` deferred (foreign holders or busy rebuild lock) stayed `_fts_stale` — LIKE-only search — until the process reopened state.db. Short-lived CLIs reopen every run; the gateway opens once and stays up for days, so the deferral was effectively permanent (#100108). The retry runs from the EXISTING gateway housekeeping tick (`_start_gateway_housekeeping`, 60s) against the shared SessionDB instances via `hermes_state_registry.live_shared_session_dbs()`: non-blocking admission (`fts_rebuild_admission(timeout_seconds=0)`), bounded backoff 60s -> 1h, no new thread, still fails closed on live holders. `fts_rebuild_admission` gains the `timeout_seconds` kwarg. * WAL-reset warning names `sys.executable` so a "linked SQLite 3.45.1" line can be matched to the interpreter that actually linked it (#100108 point 3). Deliberately NOT carried from #100130: the "leftover lock file = holder" premise (a 0-byte lock file never blocked flock; the real cause was the fork-inherited fd, fixed in894fc35337) and the `_rebuild_fts_once` one-shot rework. Co-authored-by: HexLab98 <liruixinch@outlook.com>
This commit is contained in:
+81
-4
@@ -43,6 +43,15 @@ logger = logging.getLogger("hermes_state")
|
||||
|
||||
_FTS_HOLDER_ESCALATE_ATTEMPTS = 3
|
||||
_FTS_HOLDER_ESCALATE_SECONDS = 60.0
|
||||
# Minimum spacing between in-process retries of a deferred stale-FTS rebuild
|
||||
# (``retry_deferred_fts_recovery``). The startup open already paid the full
|
||||
# admission wait once; later retries are non-blocking probes on this cadence
|
||||
# so a live holder never stalls a long-lived writer.
|
||||
_FTS_STALE_RETRY_SECONDS = 60.0
|
||||
# Each failed retry doubles the spacing up to this cap, so a holder that never
|
||||
# goes away (a second long-lived writer) costs one deferral warning per hour,
|
||||
# not one per minute. A successful rebuild clears the stale state entirely.
|
||||
_FTS_STALE_RETRY_MAX_SECONDS = 3600.0
|
||||
|
||||
# Cache for schema_read_probe_statements() — parsing SCHEMA_SQL spins up an
|
||||
# in-memory SQLite database, so derive the statements once per process.
|
||||
@@ -422,8 +431,14 @@ class SessionSchemaMixin:
|
||||
)
|
||||
return None
|
||||
|
||||
def _recover_stale_fts(self, cursor: sqlite3.Cursor, *, legacy: bool) -> bool:
|
||||
"""Atomically rebuild stale base/trigram indexes and resume syncing."""
|
||||
def _recover_stale_fts(
|
||||
self, cursor: sqlite3.Cursor, *, legacy: bool, timeout_seconds=None
|
||||
) -> bool:
|
||||
"""Atomically rebuild stale base/trigram indexes and resume syncing.
|
||||
|
||||
*timeout_seconds* bounds the cross-process admission wait; None uses
|
||||
the full startup budget, ``0`` is the non-blocking in-process retry.
|
||||
"""
|
||||
foreign_holders = self._foreign_state_db_holders()
|
||||
if foreign_holders:
|
||||
now = time.time()
|
||||
@@ -502,7 +517,9 @@ class SessionSchemaMixin:
|
||||
# authority (fail closed). Losing the race means another process is
|
||||
# already performing this exact recovery; the stale breadcrumb stays
|
||||
# set, so this process simply keeps FTS detached and retries later.
|
||||
with fts_rebuild_admission(getattr(self, "db_path", None)) as admitted:
|
||||
with fts_rebuild_admission(
|
||||
getattr(self, "db_path", None), timeout_seconds=timeout_seconds
|
||||
) as admitted:
|
||||
if not admitted:
|
||||
logger.warning(
|
||||
"Deferred stale state.db FTS rebuild: another process "
|
||||
@@ -512,6 +529,65 @@ class SessionSchemaMixin:
|
||||
return False
|
||||
return self._recover_stale_fts_locked(cursor, legacy=legacy)
|
||||
|
||||
def retry_deferred_fts_recovery(self) -> bool:
|
||||
"""Retry a deferred stale-FTS rebuild on this open SessionDB.
|
||||
|
||||
``_recover_stale_fts`` runs at open and fails closed when foreign
|
||||
holders or the rebuild lock are busy, leaving ``_fts_stale`` set and
|
||||
search on the LIKE fallback. Live write/search paths must never start
|
||||
a full rebuild (#97940), so on a short-lived CLI that deferral is
|
||||
cleared by the next process open — but a gateway opens state.db
|
||||
once and stays up for days, so "next open" never came (#100108).
|
||||
This is the in-process retry: bounded backoff from
|
||||
``_FTS_STALE_RETRY_SECONDS`` doubling to ``_FTS_STALE_RETRY_MAX_SECONDS``,
|
||||
non-blocking admission (``timeout=0``) so a live holder is skipped and
|
||||
tried again later, no new thread — the caller is an existing periodic
|
||||
tick (gateway housekeeping).
|
||||
|
||||
Returns True only when the index was rebuilt and sync triggers
|
||||
restored. Never raises.
|
||||
"""
|
||||
if not getattr(self, "_fts_stale", False):
|
||||
return False
|
||||
if getattr(self, "read_only", False) or getattr(self, "_conn", None) is None:
|
||||
return False
|
||||
now = time.monotonic()
|
||||
if now < getattr(self, "_fts_stale_retry_after", 0.0):
|
||||
return False
|
||||
interval = float(
|
||||
getattr(self, "_fts_stale_retry_interval", 0.0)
|
||||
) or _FTS_STALE_RETRY_SECONDS
|
||||
self._fts_stale_retry_after = now + interval
|
||||
self._fts_stale_retry_interval = min(
|
||||
interval * 2.0, _FTS_STALE_RETRY_MAX_SECONDS
|
||||
)
|
||||
try:
|
||||
with self._lock:
|
||||
if self._conn is None or not self._fts_stale:
|
||||
return False
|
||||
cursor = self._conn.cursor()
|
||||
legacy = self._db_has_legacy_inline_fts(cursor)
|
||||
recovered = self._recover_stale_fts(
|
||||
cursor, legacy=legacy, timeout_seconds=0.0
|
||||
)
|
||||
if recovered:
|
||||
# CJK was detached alongside the base indexes; its own
|
||||
# ensure path decides when it comes back online.
|
||||
self._ensure_fts_cjk_schema(cursor)
|
||||
self._fts_stale_retry_interval = 0.0
|
||||
try:
|
||||
self._conn.commit()
|
||||
except sqlite3.Error:
|
||||
pass
|
||||
return recovered
|
||||
except Exception: # noqa: BLE001 - background retry must never raise
|
||||
logger.warning(
|
||||
"In-process retry of the deferred stale state.db FTS rebuild "
|
||||
"failed; will retry later.",
|
||||
exc_info=True,
|
||||
)
|
||||
return False
|
||||
|
||||
def _recover_stale_fts_locked(
|
||||
self, cursor: sqlite3.Cursor, *, legacy: bool
|
||||
) -> bool:
|
||||
@@ -1510,7 +1586,8 @@ class SessionSchemaMixin:
|
||||
breadcrumb is persisted, mirroring ``_enter_fts_fail_open``'s
|
||||
ordering contract: triggers must never be live over an index with an
|
||||
unrebuilt gap. FTS stays detached for this instance; the winner's
|
||||
rebuild — or ``_recover_stale_fts`` at the next startup — restores
|
||||
rebuild — or ``retry_deferred_fts_recovery`` from the gateway
|
||||
housekeeping tick, or ``_recover_stale_fts`` at the next startup — restores
|
||||
the index and triggers atomically.
|
||||
"""
|
||||
with fts_rebuild_admission(getattr(self, "db_path", None)) as admitted:
|
||||
|
||||
Reference in New Issue
Block a user