test(state): non-contention errno table, repair-lock sibling, in-process deferred-FTS retry via housekeeping tick

Regression coverage for the #100130 salvage, all against real SessionDB
files and a real child process holding the flock:

* errno table for `is_advisory_lock_contention` (EAGAIN/EWOULDBLOCK/EACCES
  contend; ESTALE/ENOTSUP/ENOLCK/EIO fail fast); no misleading "held by
  another process" line on the fast-fail path; `_cross_process_repair_lock`
  shares the filter (sibling site).
* `retry_deferred_fts_recovery`: open under a live holder -> stale; retry
  returns in <2s with a 30s admission budget (timeout=0); rate limit +
  60s->120s backoff engaged; holder dies -> same instance recovers, triggers
  restored, breadcrumb cleared; no-op when not stale / read-only.
* `_start_gateway_housekeeping` tick (real loop, 50ms interval) recovers a
  stale shared-registry SessionDB with no direct call and no extra thread.

Backoff floor: a monkeypatched 0s base interval must not zero the doubled
interval (min 1s), so the cap math is testable.

Sabotage run (source at origin/main, these tests): 16 failed / 35 passed,
including 30s timeouts on the fast-fail tests.
This commit is contained in:
Teknium
2026-09-02 04:06:45 -07:00
parent c5138618f7
commit dbb6acd333
2 changed files with 197 additions and 4 deletions
+5 -4
View File
@@ -554,12 +554,13 @@ class SessionSchemaMixin:
now = time.monotonic()
if now < getattr(self, "_fts_stale_retry_after", 0.0):
return False
interval = float(
getattr(self, "_fts_stale_retry_interval", 0.0)
) or _FTS_STALE_RETRY_SECONDS
interval = float(getattr(self, "_fts_stale_retry_interval", 0.0))
if interval <= 0.0:
interval = _FTS_STALE_RETRY_SECONDS
self._fts_stale_retry_after = now + interval
self._fts_stale_retry_interval = min(
interval * 2.0, _FTS_STALE_RETRY_MAX_SECONDS
max(interval, _FTS_STALE_RETRY_SECONDS, 1.0) * 2.0,
_FTS_STALE_RETRY_MAX_SECONDS,
)
try:
with self._lock: