Commit Graph

44 Commits

Author SHA1 Message Date
leomcamilo bcc2e65818 fix(state): quarantine SessionDB handle after structural corruption
A bare SQLITE_CORRUPT/NOTADB on a live write (not FTS-scoped, not a
replaced file) now sets a sticky per-instance flag: later writes fail
fast with StateDbCorruptError, the handle never reopens after close(),
and close() skips its explicit PASSIVE WAL checkpoint. Gateway and agent
flush paths divert pending transcripts to JSONL/spool like the replaced
case instead of retrying forever.

Field evidence: a handle that kept writing for ~50 minutes after the
first structural error checkpointed 15 pages under the wrong page
numbers on shutdown (page 1 <- messages_fts_trigram_data leaf), turning
"malformed" into "file is not a database".

Refs #90837, #90950, #97940, #89332, #45383

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CNX8rNYHqA5pT4tAGSzXtb
2026-09-02 16:57:21 +05:30
Teknium dbb6acd333 test(state): non-contention errno table, repair-lock sibling, in-process deferred-FTS retry via housekeeping tick
Regression coverage for the #100130 salvage, all against real SessionDB
files and a real child process holding the flock:

* errno table for `is_advisory_lock_contention` (EAGAIN/EWOULDBLOCK/EACCES
  contend; ESTALE/ENOTSUP/ENOLCK/EIO fail fast); no misleading "held by
  another process" line on the fast-fail path; `_cross_process_repair_lock`
  shares the filter (sibling site).
* `retry_deferred_fts_recovery`: open under a live holder -> stale; retry
  returns in <2s with a 30s admission budget (timeout=0); rate limit +
  60s->120s backoff engaged; holder dies -> same instance recovers, triggers
  restored, breadcrumb cleared; no-op when not stale / read-only.
* `_start_gateway_housekeeping` tick (real loop, 50ms interval) recovers a
  stale shared-registry SessionDB with no direct call and no extra thread.

Backoff floor: a monkeypatched 0s base interval must not zero the doubled
interval (min 1s), so the cap math is testable.

Sabotage run (source at origin/main, these tests): 16 failed / 35 passed,
including 30s timeouts on the fast-fail tests.
2026-09-02 04:15:02 -07:00
HexLab98 c5138618f7 test(state): cover deferred FTS retry, leftover lock files, and WAL interpreter identity 2026-09-02 04:15:02 -07:00
Teknium 238b6c1ab9 fix(compression): persist the anti-thrash recovery deadline so gateway agent rebuilds cannot block a session forever
The #14694 recovery clock (`_anti_thrash_recovery_deadline`) was a
process-local `time.monotonic()` value zeroed in `bind_session_state()`.
The gateway rebuilds the AIAgent (and its ContextCompressor) on every
cache eviction, so each fresh compressor bound to a durably tripped
session row (#69872) re-armed a full 300s window and the half-open probe
never fired — a long messaging conversation above the threshold stayed
blocked permanently.

Persist the deadline as a wall-clock epoch in a new
`sessions.compression_recovery_deadline REAL` column (declarative column
reconciliation; SCHEMA_VERSION 26 -> 27) with
`SessionDB.get/set_compression_recovery_deadline`. The compressor loads it
in `bind_session_state()` and writes it on change only via
`_set_anti_thrash_recovery_deadline()`. A fresh compressor with no stored
deadline still starts a full window blocked (#54923 restart contract); one
that loads an armed deadline resumes that window. Backward clock jumps are
bounded to one window. The 300s window is unchanged.

Minimal salvage of #100185 (the probe-lease/fencing state machine and
model_config-blob storage were not carried).

Refs #100185
Co-authored-by: Komzpa <me@komzpa.net>
2026-09-02 04:14:10 -07:00
HexLab98 b4d691183b test(state): cover fail-closed deferral on an unopenable admission lock file
Drives a real unopenable lock path — a directory where the code expects a
regular file, so open() raises a genuine kernel OSError — rather than
monkeypatching the helpers, standing in for the ENOSPC/EMFILE the field reports
hit without needing to fill a disk.

Both authorities are covered at the primitive and the behavior level:
fts_rebuild_admission refuses admission and rebuild_fts() reports no progress
(asserted against a preceding successful rebuild, so the 0 is the deferral and
not an unrelated no-op); _cross_process_repair_lock refuses the authority and
repair_state_db_schema runs no writable_schema surgery, takes no forensic
backup, and leaves the damaged image byte-identical for the next authorised
pass. A guardrail test pins that a pathless in-memory store is still admitted,
so the fix cannot turn that legitimate no-op into a permanent deferral.

All four deferral assertions fail on the pre-fix code, where the repair test
shows the surgery really did proceed without cross-process authority.
2026-09-01 23:58:51 -07:00
Teknium 894fc35337 fix(state): break provably-orphaned repair/FTS-rebuild locks left by dead holders (#100108) 2026-09-01 10:52:06 -07:00
the3asic 18ac3c4fb6 fix(state): defer corrupt FTS rebuilds past live operations 2026-08-31 12:08:30 -07:00
Teknium 51773a7733 test(state): physical-corruption acceptance tests for the fail-closed classifier
Real byte-flip fixtures (no mocks) proving the #96038/#98090-class fix
end to end, closing the acceptance gate on issue #97940:

- test_canonical_btree_corruption_fails_closed: checkpoint the WAL,
  clobber every messages-table B-tree leaf page header, then assert a
  live append raises the genuine bare SQLITE_CORRUPT, the classifier
  refuses the FTS route, no rebuild/detach/stale-marker side effects
  occur, and the field incident's misdiagnosis log line ('canonical
  message rows are preserved') never appears.
- test_fts_only_corruption_still_self_heals: contrast case — a real
  messages_fts_data shadow-table stomp raises SQLITE_CORRUPT_VTAB (267),
  is classified as FTS-scoped, and the write path still self-heals with
  canonical rows intact.

Sabotage-verified: reverting the classifier fix (96739033c4) makes the
canonical-corruption test fail by entering the FTS self-heal route.

Credits @fangliquanflq (PR #98090) for the production timeline analysis
and @diatche (PR #96038) for the classifier fix these tests gate on.
Refs #97940, #98077.
2026-08-31 12:02:58 -07:00
Pavel Diatchenko 96739033c4 fix(state): fail closed on unscoped corruption 2026-08-31 11:42:23 -07:00
teknium1 9db48053e8 fix(state): self-heal SessionDB writes after close() races an in-flight worker
Subagent/cron sessions died mid-run with "Session DB append_message
failed: 'NoneType' object has no attribute 'execute'": a teardown owner
(cron run_job finally, delegate timeout owner, agent close()) called
SessionDB.close() — nulling _conn — while a still-unwinding worker had
one more transcript flush to land. The flush then hit None.execute, the
turn force-ended as session_persistence_failed, and the session tail was
silently dropped while cron delivery reported last_status: ok.

Fix at the shared persistence boundary: _execute_write and the _read_ctx
writer-lock fallback detect the closed handle under self._lock and
reopen a connection to the same database file with a loud WARNING naming
the race. Read-only handles never reopen — they raise an explicit
'was closed' error. A failed reopen raises an OperationalError naming
the teardown race so classify_persistence_error gets a real cause.

Closes #94736
2026-08-31 10:51:19 -07:00
Teknium 22dcbdece6 fix(state): contain post-commit FTS maintenance errors + lock-audit the writer conn (salvage #90734)
Salvaged from PR #90734 by @Kyzcreig onto current main:
- hermes_state_search.py: post-commit FTS incremental merge failures
  (including the bare SystemError CPython's sqlite3 layer raises under
  cross-thread errmsg scrambling) are contained and logged instead of
  escaping and making the caller replay an ambiguous, possibly-durable
  write (exactly-once refinement by @yuzilongleif-collab).
- tests/state/test_writer_conn_thread_safety.py: live reader/writer race
  hammer + AST sweep freezing the no-unlocked-writer-conn invariant.

On top: the sweep now also flags self._conn PASSED to helpers, which
caught one more live site on main — get_session_delete_targets handed
the shared writer connection to _collect_delegate_child_ids inside a
_read_ctx block, executing on it without self._lock. Routed to the
borrowed read connection.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
2026-08-31 09:56:04 -07:00
VVV 087cc49a26 fix(compression): refresh lease in-transaction before publish; arm cooldown on split failure
Two narrow repairs for #97948 symptom B (large-session rotation aborts with
'Compression lease lost before publication' / session_split_failed, then the
next turn re-runs the identical doomed compression):

1. publish_compression_child gains require_lease_refresh: the lease is
   extended inside the same transaction as the expiry check (same conn, no
   TOCTOU), giving a worker whose refresher thread died from transient DB
   failures one final chance to keep its completed work.

2. A failed compression split now records a 60s failure cooldown, so the
   next turn cannot immediately re-trigger the same compression.

Salvaged from #98137 (author: vsd2807). The timeout-reconciliation half of
that PR is NOT carried: it has a blocking review (runtime sid vs persisted
session_key, one-shot check cannot observe a 6-minute commit, no identity
projection) and needs a redesign.
2026-08-31 14:09:42 +05:30
kshitijk4poor 112baae665 fix(state): close gate blind spots — alias + variable-SQL readers (simplify findings)
Reviewer findings from the formal /simplify-code pass, all verified:

1. Scanner blind spots (HIGH): five more locked pure readers were
   invisible to the v1 gate — get_compression_fallback_streak and
   get_compression_ineffective_count hid behind `conn = self._conn`
   aliasing; list_gateway_sessions, find_session_by_origin and
   search_sessions hid behind SQL held in variables/f-strings (the
   scanner required unknown == 0 to flag). All five converted to
   _read_ctx(); the scanner now (a) tracks self._conn aliases and
   (b) flags lock blocks with NO proven write instead of silently
   skipping unprovable SQL. Sabotage self-check extended to pin all
   three detection classes (literal, alias, variable-SQL) plus a
   mixed variable-SQL writer that must NOT fire.

2. get_meta reverted to the writer lock: its inline comment (present
   on main) documents a real read-your-writes dependency —
   fts_rebuild_step reads rebuild progress that a pooled WAL reader
   cannot see mid-transaction. The blanket conversion had overridden
   a documented design decision; it is now the single justified
   _ALLOWED_LOCKED_READERS entry, replacing the dead
   _enter_fts_fail_open entry (whose lock block counts 3 writes and
   never needed allowlisting).

Strengthened scanner on pre-conversion main: 43 violations
(39 pure-read + 4 no-proven-write). This branch: zero.

Suites: gate 2/2; tests/test_hermes_state.py + tests/state/ 331
passed (same 3 pre-existing FTS-rebuild reds as clean main); 433
passed across the 12 consumer suites of the five newly-converted
methods (compression anti-thrash, session search, status, scheduler).
2026-08-29 10:38:53 +05:30
kshitijk4poor 0534f1033b perf(state): route 39 pure-read SessionDB methods off the writer lock + gate
Pattern-C architectural fix (write-lock contention), completing what
#90734 started: that PR fixed the four UNLOCKED readers racing the
writer connection; this one fixes the 39 LOCKED pure readers convoying
every concurrent turn's persistence behind the global writer lock, and
adds the gate that stops the class from re-entering.

The gateway shares ONE SessionDB across every agent. A read-only query
under `with self._lock:` blocks all concurrent writers for its
duration; under WAL, _read_ctx() serves the same read from a pooled
read-only connection with no lock at all (non-WAL falls back to the
locked writer byte-for-byte, so DELETE-journal installs are unchanged).

Converted (SELECT-only bodies, mechanical `with self._lock:` →
`with self._read_ctx() as conn:`): gateway routing loaders, session/
message counters, titles, compression tip/lineage/cooldown readers,
telegram topic bindings, prune candidate scans, meta readers, resume
resolution — 39 methods, verified per-method that every statement is a
SELECT and every conn use stays inside the with-block. Read-modify-
write methods and checkpoint/maintenance PRAGMAs stay on the writer
lock (their read is ordered against their own write).

Gate: tests/state/test_no_locked_readers_gate.py — an AST scanner over
SessionDB that fails CI when a pure-read method body takes the writer
lock, with a sabotage self-check proving the scanner fires. On
pre-conversion main it reports 38 violations; on this branch, zero.

Measured (2 writer threads + 3 reader threads, 3s, WAL):
  before: 78.5k reads, p99 2.49ms, max 160.7ms  (readers convoy)
  after: 200.4k reads, p99 0.53ms, max  25.5ms  (2.6x throughput,
         4.7x better p99, 6x better tail; writes unchanged)
Honest caveat: on runtimes where WAL is refused (the currently-bundled
SQLite 3.46 trips the WAL-reset-vulnerability gate → journal=DELETE),
_read_ctx() falls back to the identical locked-writer path and this
change is behavior-neutral by construction; the win applies to WAL
installs (legacy WAL DBs, fixed runtimes, wal-configured operators).

Suites: tests/test_hermes_state.py 243 passed; combined state sweep
540 passed — the only 3 reds are the pre-existing
test_fts_runtime_rebuild failures, verified failing on clean
origin/main before this change.
2026-08-29 10:26:21 +05:30
fangliquanflq fe24525605 fix(state): reap only proven database holders 2026-08-23 20:01:41 -07:00
fangliquanflq 8f3a82f96a fix(state): recover FTS after orphan holder deferrals 2026-08-23 20:01:41 -07:00
Teknium 9d0727d49b fix(state): single fail-closed cross-process authority for all full FTS rebuilds
Follow-up to the salvaged #93200 commit. Factors the portable
_cross_process_repair_lock ownership pattern (msvcrt on Windows, flock on
POSIX, bounded 120s wait) into a cycle-safe shared primitive,
fts_rebuild_admission() in hermes_state_common, and routes EVERY full
structural FTS rebuild entry point through it:

- SessionSearchMixin.rebuild_fts() (replaces the POSIX-only, fail-open
  30s flock from the original commit)
- _init_schema's trigger-repair rebuilds (_rebuild_fts_indexes /
  _rebuild_legacy_fts_indexes) via _run_admitted_startup_rebuild
- _recover_stale_fts()

Fail closed: a caller that cannot acquire the authority DEFERS the rebuild
(FTS detached + durable stale breadcrumb, retried at next startup) instead
of proceeding into the exact concurrent-rebuild interleaving that
structurally corrupted state.db in production. Chunked deferred backfill
(fts_rebuild_step) intentionally stays outside the authority.

Adds spawned-process regression tests (real child process holding the real
lock file): holder blocks contender, deferral fails closed on both the
runtime and schema paths, release/holder-death permits the next owner, and
stale recovery completes after contention clears. Sabotage-verified: 4/6
tests fail with the admission forced open.
2026-08-23 19:00:36 -07:00
kshitij 987064caa4 fix: restore generic corruption match in FTS self-heal
The PR narrowed _is_fts_write_corruption_error to only match FTS5-specific
'fts5: corrupt structure record' errors, dropping the generic 'database disk
image is malformed' match. But FTS shadow table corruption (the common case)
raises the generic error on SQLite < 3.53, not the FTS5-specific one. This
broke FTS self-heal for 10 existing tests and for users on older SQLite.

Restore the generic match via is_malformed_db_error. Safety is preserved
because the FTS rebuild only touches derived indexes — if the damage is
actually in a canonical B-tree, the rebuild itself fails and the write
propagates.

Also restore the original test assertion and remove the
test_generic_malformed_write_fails_closed test whose premise (generic
corruption should not trigger FTS rebuild) was wrong for the FTS self-heal
path.
2026-08-23 02:55:03 +05:30
Bruce Xu 50bbcbf2b4 fix(state): fail closed on unscoped SQLite corruption 2026-08-23 02:55:03 +05:30
kshitijk4poor 1c59daaace fix(state): use /proc readlinks + cmdline fallback for holder detection
Address review feedback from @jackulau on PR #90871:

1. psutil.open_files() silently drops '(deleted)' WAL sidecar entries
   on Linux because isfile_strict() stats the literal path including
   the suffix and fails. Switch to direct /proc/<pid>/fd readlinks
   which preserve the '(deleted)' suffix so _canonical can match.

2. psutil.process_iter() converts AccessDenied to None, which
   or-() skips silently — the fail-closed branch never runs. For the
   root-gateway vs user-desktop topology in the issue, the fd table is
   unreadable but /proc/<pid>/cmdline is world-readable. Add a cmdline
   fallback that flags uninspectable processes.

Also keep the psutil path for macOS/BSD (no '(deleted)' convention).
2026-08-22 03:56:13 +05:30
fangliquanflq fc72d6c716 fix(state): defer FTS rebuild under foreign WAL holders 2026-08-22 03:56:13 +05:30
Victor Kyriazakos 8b03e65804 chore: generalize field-report attribution in code comments 2026-08-17 17:20:06 -07:00
Victor Kyriazakos e997435004 fix(state): v25 prompt dedupe degrades gracefully on a contended DB
Only the initial SELECT of _dedupe_legacy_system_prompts was guarded;
a 'database is locked' on any per-row write propagated out, aborted
schema init, left the schema version below 25, and made every later
SessionDB.__init__ re-enter the same migration against the same
contended DB - the second half of the enterprise crash-loop report.

The per-row loop now catches OperationalError, logs once, and returns.
Partial migration is safe by design: the legacy system_prompt column
is the documented read fallback for unmigrated rows, and the next
schema init resumes where the contention stopped. Tests prove rows
migrated before the failure stay migrated, the remainder stays
readable, and a later run completes it.
2026-08-17 17:20:06 -07:00
Teknium 652f5c2ebb fix(compression): rotation path clones the concurrent tail into the child
CI caught the sibling site the in-place fix missed: legacy (non-in-place)
compression rotates via publish_compression_child, where a mid-summary
append previously stranded in the closed parent. Same watermark + pure-SQL
column clone as archive_and_compact, with session_id rewritten to the child.
Lineage-guard test flipped to pin the appends-flow-freely contract; rotation
watermark tests added (tail follows the child; None = historical behavior).
2026-08-15 23:37:22 -07:00
embwl0x e89532d97e fix(desktop): order async session git metadata 2026-08-15 00:33:11 -07:00
fangliquanflq 19b1204392 fix(sessions): revive uncontested expired turn leases 2026-08-15 03:22:16 +05:30
fangliquan f1025b2c00 fix(sessions): fence transcript writes with the turn-lease holder
Refresh-loss interrupt is cooperative, so a stalled writer could still flush after another process reclaimed the conversation. Carry the holder into append_message / append_messages_batch and reject the write in the same SQLite transaction when the lease row is missing, expired, or owned by someone else.
2026-08-15 03:22:16 +05:30
fangliquan c21efeeb52 fix(sessions): keep turn lease across inherited-marker compressions
Presence-only _delegate_from/_branched_from checks stopped the lease walk on
continuations that copied a delegate's model_config, so the first refresh
after rotation missed the parent-key lease and hard-interrupted. A failed
get_session probe also skipped acquire entirely. Walk the lineage inside
the write transaction and treat a probe error as contended, not a fresh
session.
2026-08-15 03:22:16 +05:30
fangliquan 5e2be43fd4 fix(sessions): harden cross-process turn lease wait and refresh
Honor interrupts while waiting for admission, stop the turn when refresh
loses the lease, poll once per second under contention, and test dead-PID
reclaim.
2026-08-15 03:22:16 +05:30
fangliquan 3b09456019 fix(sessions): surface wait status for cross-process turn leases
Emit lifecycle notices while waiting on another process, and return a
resend-friendly timeout result instead of a bare TimeoutError.
2026-08-15 03:22:16 +05:30
fangliquan 6e929a9694 fix(sessions): serialize turns across processes 2026-08-15 03:22:16 +05:30
izumi0uu 1527a81b5e fix(state): keep canonical writes available when FTS is corrupt 2026-08-09 14:10:21 -07:00
kshitij a0801b878a fix: bind continuation-marker exclusions to the queried parent (fail-open fix)
Adversarial review of the salvaged recovery found a reachable fail-open:
compression continuations inherit the rotated agent's model_config
verbatim (publish_compression_child callers pass
agent._session_init_model_config), so a delegate subagent's continuation
carries _delegate_from=<the delegate's own parent>. The marker-PRESENCE
filters in reopen_orphaned_compression_session and
find_live_compression_child misclassified such a REAL continuation as a
delegate child:

- reopen: parent 'orphaned' -> reopened while a live continuation exists
  -> two live heads in one lineage (verified with a live repro)
- find_live: adoption misses the continuation (fail-closed, masked the
  fork pre-PR; the PR made it active)

Fix: markers only disqualify a child when they point at the queried
parent (shared _NON_CONTINUATION_CHILD_FILTER_SQL fragment, also
resolving the duplicated-SQL drift risk flagged by the reuse reviewer).
Both directions regression-tested: reopen fails closed on an
inherited-marker continuation; find_live adopts it.

Also from review: reopen-failure log raised debug->warning (the failure
hard-fails the turn moments later), commit-semantics hardening comment
on the lease DELETE path, blank-line nit.

The three read-only projection walks (get_compression_tip,
list_sessions_rich chain, resume walk) share the marker-presence shape
but fail closed (skip a continuation -> resume shows the parent), and
the fixed adoption path self-heals that case at turn start; left as-is.
2026-08-07 13:24:56 +05:30
izumi0uu 95a7058e4b fix(sessions): fence expired orphan recovery leases 2026-08-07 13:24:56 +05:30
izumi0uu 988f2baaf8 fix(sessions): recover compression parents without continuations 2026-08-07 13:24:56 +05:30
RelaxJonh b6ca4fc856 fix(state): heal session_model_usage PK unconditionally to restore token/cost accounting
Installs whose state.db reached schema_version >= 22 before the task
dimension was added carry a 5-column PRIMARY KEY on
session_model_usage. The column reconciler ADDs task as a bare
nullable, but SQLite cannot ALTER a primary key, and the version-gated
v22 rebuild is unreachable (current_version < 22 already false), so
the composite 6-column key never lands. Every upsert in
_record_model_usage then fails with 'ON CONFLICT clause does not match
any PRIMARY KEY or UNIQUE constraint', aborting the enclosing write
transaction — token/cost accounting permanently dead (#73823).

Add an idempotent _heal_session_model_usage_pk() modeled on
_heal_gateway_routing_pk(), run unconditionally from _init_schema on
every open. Salvaged from #73838 with fix-ups:

- ported to SessionSchemaMixin in hermes_state_schema.py (the schema
  code moved out of hermes_state.py in 21c7ae8563; the PR targeted the
  old location)
- rebuild wrapped in a PRAGMA foreign_keys=OFF/ON window: the
  connection enables FKs before _init_schema and OR IGNORE does NOT
  suppress FK violations, so a single orphaned usage row (session
  pruned while accounting was broken) would have aborted the heal
- COALESCE('') on the nullable reconciler-added task column (and the
  billing columns) during the copy
- stale-v22+ regression tests: rebuilt PK + restored upsert, orphan
  rows survive the FK window, healthy-DB no-op, no legacy leftover

Fixes #73823
2026-07-31 23:18:12 -07:00
Dannoob 14eca89779 fix(state): retry transient 'no more rows available' across all sqlite3.Error classes
Under dual gateway/agent WAL contention (FTS5 trigram sync holding the
write lock on large appends) the SQLite engine can raise a transient
'no more rows available' error. The exception CLASS varies with the
build — some surface it as InterfaceError, a SIBLING of DatabaseError —
so it escaped both existing retry branches in _execute_write on attempt
0 and killed the turn as session_persistence_failed even though the
identical write succeeds standalone.

Port of #74934 onto the deadline-patience rewrite (8da8a7887d): the
PR's attempt-counted constants (60 retries / 300ms jitter / 2.0s engine
timeout) predate that rewrite and are superseded by the patience
budget, so they are intentionally NOT carried over. Instead the check
is message-scoped and rides the existing deadline/patience loop:

- extract the jittered-sleep-within-deadline logic into a shared
  _sleep_before_retry helper (behavior-preserving for locked/busy)
- retry 'no more rows available' from OperationalError, DatabaseError
  (checked BEFORE the FTS-corruption rebuild path so it is not
  misrouted), and a message-scoped sqlite3.Error catch-all
- any other error in any class propagates untouched on attempt 0

Tests: transient InterfaceError retried to success; unrelated
InterfaceError propagates immediately; DatabaseError variant retried;
exhausted patience surfaces the original error.
2026-07-31 23:18:12 -07:00
Brooklyn Nicholson 24f346ee77 fix(gateway): fail prompt.submit when session storage hits a full disk
Disk-full / ENOSPC / SQLITE_FULL on first-message session persist used to be
swallowed as a debug log while prompt.submit still returned streaming, so the
send vanished with no error. Re-raise those failures, return a real RPC error,
and stamp session_persistence_failed turns with error so clients get a terminal
error frame.
2026-08-01 00:27:17 -05:00
Teknium 8da8a7887d fix(state): time-based write-lock patience so busy sibling processes can't destroy turns
A shared state.db is legitimately held for multi-second stretches by
sibling Hermes processes: VACUUM after auto-prune, the TRUNCATE WAL
checkpoint at close on a large WAL, offline recovery, or an older
still-running process whose FTS maintenance predates the bounded-merge
protocol (every `hermes update` leaves mixed-version processes sharing
the DB until the old ones exit).

The old retry budget was attempt-counted: 15 attempts x 20-150ms jitter
gives up after ~1.3s of waiting. Any hold longer than that surfaced as:

- append_message failing -> the conversation loop aborts the turn as
  session_persistence_failed ('No reply: the turn was stopped because
  session storage could not be written') on a perfectly healthy store;
- SessionDB() open failing -> the CLI disables persistence for the
  entire run ('Failed to initialize SessionDB ... database is locked').

Both observed in production logs on 2026-07-29 (10.8 GB state.db, 9
concurrent hermes processes, three of them pre-dating the bounded-merge
fix pull).

Changes:

- _execute_write patience is now TIME-based with two budgets: routine
  writes wait up to 20s; transcript-critical writes (append_message,
  session-row creation — the ones whose failure aborts a user turn)
  wait up to 60s. Jitter stays 20-150ms for the first 2s, then backs
  off to 250ms-1s so a long hold isn't hammered with BEGIN IMMEDIATE.
- Exhausted patience raises an error that names the actual cause
  (another process held the write lock; the database is healthy)
  instead of a bare 'database is locked' that reads like disk damage —
  and the turn-abort explainer inherits that clarity.
- SessionDB open now applies the same jittered patience to the
  locked/busy class around connect+schema-init, instead of failing the
  whole open (and disabling persistence for the run) on the first 1s
  timeout. Non-lock errors, including the malformed-schema repair
  class, propagate immediately as before.

Fixes #74478
2026-07-29 17:59:02 -07:00
Teknium 6b81590c55 test: prune low-value tests suite-wide (wave 1) — 46,820 → 28,106 test functions
Systematic prune per AGENTS.md test policy, one pass over every major
test tree (gateway, hermes_cli, tools, agent, run_agent, plugins, cli,
cron, tui_gateway, honcho/openviking, root-level):

- DELETE: source-reading tests (read_text/getsource on prod files),
  change-detector tests (exact catalog counts, model-name snapshots,
  config version literals), mock-echo tests (assert a mock returns what
  it was told), assertion-free/trivial tests, near-duplicate
  parametrizations (boundaries + one representative kept), async/sync
  twin duplicates, cosmetic within-file variations.
- KEEP (mandatory): security/redaction/approval guards, message-role
  alternation invariants, prompt-caching/deterministic-call-id
  invariants, issue-number regression tests (deduped), E2E tests.
- 6 test files deleted outright (script-style/no-assert or fully
  redundant); conftest.py, fakes/, fixtures/ untouched.
- tests/acp/conftest.py added: autouse fixture stubs the live
  models.dev/GitHub/Copilot/Anthropic inventory fetches that ACP server
  tests performed on every session create — test_server.py 147s → 3.4s,
  and the tests are now genuinely hermetic.
- Sleep-based slowness shrunk where safe (codex_ttfb_watchdog,
  compression_concurrent_fork, etc.); no wall-clock assertion tightened.

Verification: full hermetic suite via scripts/run_tests.sh —
2439 files, 31,130 tests passed, 0 failed, 0 flaky retries, 315s wall
(baseline: 583s wall, 13,564s subprocess CPU).
2026-07-29 13:10:23 -07:00
Anthony Ruiz 0ee8d41878 fix(compression): recover rotated session lineage 2026-07-24 16:00:34 -07:00
Teknium 373ec23e37 fix(state): extend search-path FTS self-heal to the CJK/trigram branch
The trigram MATCH branch in search_messages() had the same
OperationalError-only catch that #66420 fixed on the main FTS5 branch: a
corrupt messages_fts_trigram shadow table raises the malformed /
'fts5: corrupt structure record' class (sqlite3.DatabaseError, parent of
OperationalError), which propagated straight out of search_messages and
crashed CJK session/history search for read-only sessions.

Route that class through the shared one-shot _try_runtime_fts_rebuild()
and retry the trigram query (catch moved outside self._lock so
rebuild_fts() can re-acquire it, mirroring the main branch). If the
rebuild is refused (guard consumed / FTS disabled / different error) or
the retry fails, fall through to the existing LIKE substring fallback —
which reads only the canonical messages table — instead of raising, so
CJK search degrades gracefully rather than crashing.

Adds two regression tests: trigram search self-heals in place after
shadow-table corruption (answers from the rebuilt trigram index, not the
LIKE fallback), and degrades to LIKE without raising when the one-shot
rebuild was already consumed.

Follow-up to #66420; refs #66296 #66724
2026-07-21 12:40:48 -07:00
Frowtek 11710c51fc fix(state): self-heal FTS corruption on the SessionDB search path too
Complements #66296 (self-heal on the write path): search_messages()'s main
FTS5 MATCH query caught only sqlite3.OperationalError (a query-syntax error →
return empty). A corrupt FTS index raises the malformed / "fts5: corrupt
structure record" class, which is a sqlite3.DatabaseError — the parent of
OperationalError, so it was NOT caught and propagated straight out of
search_messages, crashing session/history search.

The write path now rebuilds and retries on that class, but a read-only
session (cron/CLI history search, or a search issued before any write) never
triggers a write, so its search stayed broken until the next process restart
ran the offline repair.

Catch the DatabaseError corruption class on the search MATCH read too and
route it through the existing one-shot _try_runtime_fts_rebuild(), then retry
the query. The catch is moved outside `with self._lock` so rebuild_fts() can
re-acquire the lock (mirrors _execute_write). The one-shot guard is shared
with the write path, so a single instance never loops on a genuinely
unrecoverable index. OperationalError syntax handling is unchanged (caught
first).

Adds a regression test: with a corrupted messages_fts and no post-corruption
write, search_messages() rebuilds in place and returns the match; without the
fix it raises DatabaseError.
2026-07-21 12:40:48 -07:00
Teknium 9e1b1d7536 fix(state): self-heal FTS corruption on the SessionDB write path (#66296)
Complements the #65637 salvage (53d358838 + a9cc17fd8): the gateway
session store now retries transcript appends through its own queue, but
cron and CLI writers call SessionDB directly — a corrupt FTS index still
hard-failed their appends until the next process restart triggered the
offline repair.

_execute_write now detects the FTS-corruption error class (both the
generic 'database disk image is malformed' and newer SQLite's
'fts5: corrupt structure record' variant), performs a one-shot in-place
rebuild by delegating to the existing rebuild_fts(), and retries the
failed write. One-shot per instance so an unrecoverable database cannot
loop; lock/busy jitter-retry path untouched.

E2E-verified: corrupted messages_fts_data rejects appends; with this fix
the same append self-heals, persists, and FTS search works again.
2026-07-17 08:51:51 -07:00