Regression coverage for the #100130 salvage, all against real SessionDB
files and a real child process holding the flock:
* errno table for `is_advisory_lock_contention` (EAGAIN/EWOULDBLOCK/EACCES
contend; ESTALE/ENOTSUP/ENOLCK/EIO fail fast); no misleading "held by
another process" line on the fast-fail path; `_cross_process_repair_lock`
shares the filter (sibling site).
* `retry_deferred_fts_recovery`: open under a live holder -> stale; retry
returns in <2s with a 30s admission budget (timeout=0); rate limit +
60s->120s backoff engaged; holder dies -> same instance recovers, triggers
restored, breadcrumb cleared; no-op when not stale / read-only.
* `_start_gateway_housekeeping` tick (real loop, 50ms interval) recovers a
stale shared-registry SessionDB with no direct call and no extra thread.
Backoff floor: a monkeypatched 0s base interval must not zero the doubled
interval (min 1s), so the cap math is testable.
Sabotage run (source at origin/main, these tests): 16 failed / 35 passed,
including 30s timeouts on the fast-fail tests.
Two pieces of PR #100130 (@HexLab98) re-applied on top of the orphaned-flock
break (894fc35337) and fail-closed admission (#100895) that landed since:
* `is_advisory_lock_contention` (hermes_state_common): only EAGAIN /
EWOULDBLOCK / EACCES / EDEADLK mean "another process holds the lock".
ESTALE / ENOTSUP / ENOLCK / EIO from flock or msvcrt.locking are
environment failures that polling cannot fix — `_acquire_db_flock` and
both Windows msvcrt loops (FTS rebuild admission, state.db repair lock)
now defer immediately with the real errno instead of burning the full
120s / holder timeout and then logging a fake "held by another process".
* `retry_deferred_fts_recovery` (hermes_state_schema): a SessionDB whose
open-time `_recover_stale_fts` deferred (foreign holders or busy rebuild
lock) stayed `_fts_stale` — LIKE-only search — until the process
reopened state.db. Short-lived CLIs reopen every run; the gateway opens
once and stays up for days, so the deferral was effectively permanent
(#100108). The retry runs from the EXISTING gateway housekeeping tick
(`_start_gateway_housekeeping`, 60s) against the shared SessionDB
instances via `hermes_state_registry.live_shared_session_dbs()`:
non-blocking admission (`fts_rebuild_admission(timeout_seconds=0)`),
bounded backoff 60s -> 1h, no new thread, still fails closed on live
holders. `fts_rebuild_admission` gains the `timeout_seconds` kwarg.
* WAL-reset warning names `sys.executable` so a "linked SQLite 3.45.1"
line can be matched to the interpreter that actually linked it
(#100108 point 3).
Deliberately NOT carried from #100130: the "leftover lock file = holder"
premise (a 0-byte lock file never blocked flock; the real cause was the
fork-inherited fd, fixed in 894fc35337) and the `_rebuild_fts_once`
one-shot rework.
Co-authored-by: HexLab98 <liruixinch@outlook.com>
Addresses Enough1122's two non-blocking review notes on PR #92316:
1. Lease vs deadline arithmetic: the state.db init/migration/repair leases
(600-900s) are authoritative against the 300s default deadline by design.
Add explicit comments at all three lease sites pinning the trade-off:
single lease is deliberate (clamped to _MAX_LEASE_S=900), honest worst
case is up to the lease duration of zombie time on a wedged DB phase,
accepted over per-chunk renewal complexity in the migration loops.
2. Arm-site coverage: add a structural contract test asserting every
documented entry point (hermes_cli/main.py argv fast-path,
hermes_cli/gateway.py config-bridge re-arm, gateway/run.py backstop,
cli.py legacy --gateway) actually calls arm_startup_watchdog (or its
aliased import), so a future entry point can't silently ship unwatched.
Addresses the two class-level review blockers on PR #89750:
1. Bounded hard-exit seam (escort thread). The forensic fire path
(logger.critical, dump record, faulthandler, lifecycle ledger) can
itself wedge — the parked main thread may hold the logging handler
lock, or the disk may be full/hung. _fire() now starts an exit-escort
daemon thread BEFORE any forensics; it is free of log handlers,
filesystem access, module loads and application locks, and hard-exits
with the restart code after _FIRE_EXIT_BOUND_S unless the normal fire
path signals completion. Adversarial tests hold the logging handler
lock / hang the dump write at fire time and assert the exit seam is
still reached.
2. Phase-owned progress leases (report_startup_progress). Process CPU
time proves process activity, not startup progress: an unrelated busy
thread could extend forever while startup sits parked (false
negative), and I/O-bound repair/backup accrues ~zero CPU and would be
killed (false positive). Long synchronous startup phases now declare
authoritative, clamped (_MAX_LEASE_S), renewable progress leases:
state.db _init_schema + the version-gated data-migration chain
(hermes_state_schema) and repair_state_db_schema (hermes_state) are
wired. CPU progress remains only as a bounded fallback, capped at
_MAX_CPU_EXTENSIONS, with leases outranking the cap. Adversarial
tests cover both directions (lease saves zero-CPU legitimate work;
capped CPU noise no longer hides a parked deadlock).
Fire-path dump record now includes lease_count/last_lease_phase for
forensics. gateway/startup_watchdog.py shim re-exports
report_startup_progress.
OOF-298
Companion to #98935, which fixes _fts_table_probe itself. This covers the
surfaces that PR does not touch:
- web_server._open_session_db_at_path: the one-writable-open heal only
caught sqlite3.DatabaseError; a raw UnicodeDecodeError (pysqlite failing
to decode SQLite's own error message over corrupt file bytes) bypassed
it, so the heal documented for malformed schema never fired (#98924
Failure 1). Both catches widened; decode errors dispatch to the heal.
- SessionSchemaMixin._recover_stale_fts_locked: drop-and-recreate skipped
vtables whose probe raised UnicodeDecodeError, the same too-narrow
catch the issue identified in the probe.
- TUI gateway: _ensure_session_db_row returned silently when the store
could not open, so prompt.submit streamed the turn while persisting
nothing (#98924 Failure 2). It now returns False and prompt.submit
fails the RPC with code 5072 so desktop maps it to a toast, mirroring
the disk-full/5070 convention. session.create stays silent per its
pinned degraded-mode contract.
Invalid UTF-8 bytes in messages.content (e.g. 0x81 from hardware issues,
corrupted disk I/O, or manual DB edits) caused read-only SessionDB init
to die on a bare UnicodeDecodeError in _fts_table_probe, taking down
every read endpoint (GET /api/sessions, Desktop read-only opens of other
profiles' DBs). The probe caught only sqlite3.OperationalError and would
re-raise any other exception, including UnicodeDecodeError (a ValueError,
not an sqlite3.Error subclass).
On some Python/SQLite builds the decode failure surfaces as
UnicodeDecodeError; on others as OperationalError('Could not decode to
UTF-8 column ...'). The fix catches both and treats them the same:
the FTS index is degraded (search may return less or fail), but the store
itself stays accessible for writes and non-FTS reads. Writable init
schedules a rebuild or degrades to LIKE search until repaired.
Adds test_98924_readonly_fts_decode_error.py with a regression test that
injects invalid UTF-8 via CAST(x'...' AS TEXT) through the Python sqlite3
module, triggers an FTS rebuild, then confirms that read-only init succeeds
instead of raising.
`_init_schema` decided whether the FTS triggers needed repair by comparing
the live trigger count against `len(_FTS_TRIGGERS)`, the full six-name set.
Three of those six are the `messages_fts_trigram_*` triggers, and they are
declared only inside `FTS_TRIGRAM_SQL` / `LEGACY_FTS_TRIGRAM_SQL`, whose
`CREATE VIRTUAL TABLE ... tokenize='trigram'` needs a tokenizer SQLite only
gained in 3.34.
On an older build `_ensure_fts_schema` soft-fails that DDL by design (via
`_is_trigram_unavailable_error`) and returns False, so those three triggers
can never be created. The count is therefore pinned at 3, `3 < 6` is
permanently true, and the repair path ran on every single `SessionDB` open,
forever, while holding the SQLite write lock. It never converged: every
`hermes` command, gateway start, dashboard request and cron tick paid a full
re-index of the message corpus. That is ordinary LTS territory — Ubuntu
20.04 ships 3.31, RHEL/CentOS 8 and Alibaba Cloud Linux ship 3.26, and
Hermes has no minimum-SQLite gate precisely because it is supposed to
degrade gracefully here.
The v23 repair also ends by clearing `fts_rebuild_high_water` and
`fts_rebuild_progress`, which is correct after a genuine full rebuild but
means an interrupted `hermes sessions optimize-storage` silently lost its
resume point on the next open, restarting the chunked backfill from zero
every time.
Fix: keep `_FTS_TRIGGERS` as the single source of truth and derive two
subsets from it, then measure each half against the DDL that can actually
create it. `_fts_trigger_count` takes an optional `names` sequence
(defaulting to the full set, so no caller changes), and both branches gate
on `base_triggers_missing or (trigram_enabled and trigram_triggers_missing)`.
The counts are still taken before the DDL runs so they describe the
pre-repair state, while `trigram_enabled` is only known afterwards — hence
the combination at the `if` rather than at the assignment.
Behaviour is unchanged wherever the tokenizer exists: a genuinely missing
trigram trigger on a capable host still triggers the rebuild. Only the
permanently unsatisfiable comparison changes.
Follow-up to the salvaged #93200 commit. Factors the portable
_cross_process_repair_lock ownership pattern (msvcrt on Windows, flock on
POSIX, bounded 120s wait) into a cycle-safe shared primitive,
fts_rebuild_admission() in hermes_state_common, and routes EVERY full
structural FTS rebuild entry point through it:
- SessionSearchMixin.rebuild_fts() (replaces the POSIX-only, fail-open
30s flock from the original commit)
- _init_schema's trigger-repair rebuilds (_rebuild_fts_indexes /
_rebuild_legacy_fts_indexes) via _run_admitted_startup_rebuild
- _recover_stale_fts()
Fail closed: a caller that cannot acquire the authority DEFERS the rebuild
(FTS detached + durable stale breadcrumb, retried at next startup) instead
of proceeding into the exact concurrent-rebuild interleaving that
structurally corrupted state.db in production. Chunked deferred backfill
(fts_rebuild_step) intentionally stays outside the authority.
Adds spawned-process regression tests (real child process holding the real
lock file): holder blocks contender, deferral fails closed on both the
runtime and schema paths, release/holder-death permits the next owner, and
stale recovery completes after contention clears. Sabotage-verified: 4/6
tests fail with the admission forced open.
Only the initial SELECT of _dedupe_legacy_system_prompts was guarded;
a 'database is locked' on any per-row write propagated out, aborted
schema init, left the schema version below 25, and made every later
SessionDB.__init__ re-enter the same migration against the same
contended DB - the second half of the enterprise crash-loop report.
The per-row loop now catches OperationalError, logs once, and returns.
Partial migration is safe by design: the legacy system_prompt column
is the documented read fallback for unmigrated rows, and the next
schema init resumes where the contention stopped. Tests prove rows
migrated before the failure stay migrated, the remainder stays
readable, and a later run completes it.
After `hermes update`, an existing state.db on an old schema made every
GET /api/sessions poll fail with sqlite3.OperationalError "no such
column: s.last_read_at" (or s.last_activity_at) until something
unrelated forced a writable open — the desktop sidebar showed "No
sessions yet" while every row sat intact on disk (#79531, #80037).
Two remaining root causes (the stale hand-written read probe was
already replaced by the SCHEMA_SQL-derived probe on main, prototyped in
draft PR #80030 by @Tilly-YL):
1. Migrations ran lazily: _init_schema/_reconcile_columns only ran on a
writable open, typically the user's first NEW session. The dashboard
backend now schedules one writable open of its own state.db from the
lifespan (daemon thread, never blocks the ready-probe socket, never
raises), so the store is brought current before the first session-
list poll on every `hermes serve` / `hermes dashboard` / Desktop
headless entrypoint.
2. _reconcile_columns caught sqlite3.OperationalError around every
ALTER TABLE ADD COLUMN and logged at DEBUG. Lock contention from
orphaned sibling backends made the ALTER fail silently — startup
"succeeded" with a half-reconciled schema, and the open-time lock
patience (#74478) never saw the error because it was swallowed
inside first. Now: "duplicate column" races stay at DEBUG,
locked/busy re-raises so _connect_and_init_with_lock_patience
retries the whole idempotent init with jittered backoff, and any
other failure (e.g. un-ADDable NOT NULL) logs at WARNING.
Regression tests: a store missing sessions.last_read_at is healed by
the eager startup reconcile and serves list_sessions_rich; a locked
ALTER propagates and is retried to success by the open lock patience;
duplicate-column races stay quiet; other ALTER failures warn.
Fixes#79531Fixes#80037
Reported-by: @yenhunghuang (#79531) and @FLOW3R0111 (#80037)
Root-cause analysis: @wangyi0177-eng (stale read probe) and
@www654cc-pixel (_reconcile_columns DEBUG-swallow under lock
contention); draft PR #80030 by @Tilly-YL prototyped the probe fix.
Cold CLI time-to-banner was ~1.8s (hermes) / ~2.8s (hermes -w). The banner
path was paying for work the session doesn't need before first input:
- aux availability probes built REAL OpenAI/httpx clients (openai import
~0.3s + SSL context) just to answer check_fns. New aux_probe_mode()
returns a cache-excluded stub; resolution policy unchanged.
- tools/mcp_tool imported the mcp SDK (~260ms, mcp.types pydantic model
construction) at module import even with zero MCP servers configured.
SDK import is now lazy behind _ensure_mcp_sdk(); _MCP_AVAILABLE is a
find_spec probe so every existing gate/test keeps its semantics.
- banner blocked 500ms on the update-check prefetch; now waits 50ms and
defers the warning line to a daemon thread (prints above the prompt).
- banner recomputed get_tool_definitions + skills scan + git state every
launch; now snapshotted to ~/.hermes/cache/banner_snapshot.json keyed on
(config.yaml, .env, checkout rev, toolsets) and replayed on warm launches
with a background refresh. Agent tool list is still computed fresh.
- _resolve_active_context_length probed the Nous portal /models (~200ms
network) per launch; the tool-search gate now prefers the on-disk
context cache when present.
- schema reconciliation re-executed SCHEMA_SQL in a scratch SQLite DB
(~85ms) per SessionDB(); the reference parse is now disk-memoized by
DDL hash (live-DB diffing still runs every startup).
- bundled-skills sync (~120-170ms rglob/hash) moved off the startup path
to a daemon thread; plugin discovery starts in the background and every
synchronous consumer joins via discover_plugins().
- hermes_cli.auth imported httpx eagerly (~30ms); now a lazy proxy that
test monkeypatching still reaches (setattr forwards to the real module).
- fast chat launch: unambiguous 'hermes'/'hermes chat' invocations skip
building all ~40 subcommand parsers (bails to full dispatch on anything
else, incl. container mode).
- -w path: git worktree add runs with checkout.workers=8 (0.6s→0.2s) and
overlaps HermesCLI construction; --skills preload runs in the background
and is folded in at agent init (finalize_preloaded_skills, same
fail-loud contract for fully-unknown skill lists); stale-worktree prune
moved off the banner path.
Warm results (PTY time-to-banner, 5-run): hermes 1.80s → 0.38-0.40s;
hermes -w -s hermes-agent-dev --yolo 2.82s → 0.57-0.69s.
After `hermes update`, the desktop sidebar showed "No sessions yet" until
the user's first message. #72424 added sessions.last_activity_at, which
list_sessions_rich now selects — but column adds only land through
_reconcile_columns() in the writable _init_schema, and read-only opens
skip that by design. Every sidebar read path opens state.db read-only, so
each poll raised "no such column: s.last_activity_at" until the first
prompt's lazy session-row persist forced a writable open and reconciled.
A heal for exactly this class already existed (_open_session_db_for_profile
probes the read-only handle and does a one-time writable reopen on
staleness), but its probe was a hand-written four-column list that never
learned last_activity_at — it went stale three days after shipping. And the
batched sidebar route (/api/profiles/sessions/sidebar) bypassed the helper
entirely, swallowing per-profile failures into an errors array the desktop
never surfaces, so the incident produced an empty sidebar with clean logs.
The fix removes the maintenance burden instead of paying it once more:
- hermes_state_schema.schema_read_probe_statements() derives one
`SELECT <every declared column> FROM <table> LIMIT 0` per table from
SCHEMA_SQL via the existing _parse_schema_columns() — the same source of
truth the writable reconciler diffs against, so any future ADD COLUMN is
probed with no list to update. Column references are table-qualified:
an unqualified double-quoted identifier that fails to resolve silently
degrades to a string literal (SQLite's double-quoted-string misfeature)
and would make the probe pass on exactly the store it exists to catch.
- web_server splits the heal into a path-level _open_session_db_at_path
(semantics unchanged) so the cross-profile session routes can share it;
both profiles.py loops and _count_status_active_sessions (the remaining
raw read-only sibling) now open through it. The heal stays a helper
rather than a SessionDB classmethod on purpose: escalation-to-writable
must remain an explicit caller decision — update_cmd.py opens read-only
mid-update and must never write.
- Exhaustion guard: if the writable heal SUCCEEDS and the re-probe still
fails (a schema problem ADD COLUMN cannot express), the store is marked
exhausted — warn once, skip the probe, serve reads probe-less — instead
of re-running the full writable init on every poll against a possibly
live DB. A FAILED writable open (transient lock) is deliberately not
recorded, so the next poll retries the heal.
- The per-profile swallow sites in profiles.py now also log a deduplicated
warning, so a persistent read failure is loud in errors.log even though
the response errors array stays invisible to the sidebar.
Tests: probe/SCHEMA_SQL coverage invariants (tests/test_schema_read_probe.py),
last_activity_at added to the /api/sessions heal parametrize, a sidebar-route
heal test reproducing the shipped symptom (errors == [] and the session
returned against a store missing the column), and an exhaustion test pinning
exactly one writable open. The sidebar and last_activity_at tests fail on
main.
Simplify-pass fold: to_drop names come from the literal update_names\nallowlist via IN binding, so the [A-Za-z0-9_]+ fullmatch could never\nfail — and if it somehow did, its `continue` would miscount (the\nskipped trigger stayed in len(to_drop)/the log while CREATE TRIGGER\nIF NOT EXISTS silently kept the broad variant). Delete the guard and\nits function-local re import; keep the invariant as a comment.
_ensure_fts_cjk_schema never raises on OperationalError; post-condition
after dropping messages_fts_cjk_update now requires a narrowed UPDATE
trigger or durable fts_cjk_stale + unavailable. Covers the production
soft-fail path the raise-only handler missed.
Retarget #73639 onto the SessionDB mixin split (hermes_state_common /
hermes_state_schema). Fresh installs create UPDATE OF content/tool_*
triggers; existing broad AFTER UPDATE triggers are inspected and
replaced under schema init without an FTS rebuild (WHEN clauses already
guarded content correctness; OF skips non-content status writes that
saturated disk I/O on large state.db).
Tests: tests/test_fts_update_of_narrowing.py (4)
Demote wrote the empty v23 schema via executescript inside BEGIN IMMEDIATE,
which commits early and can leave trash + empty indexes without rebuild
markers. Re-run then tore down trash and stamped fts_storage_version with
docsize=0, permanently losing historical session search.
Stage markers with the demote, create schema only after they are durable,
heal empty-index bookkeeping on resume, and refuse settle until the base
index is populated. Settle refusal returns ok=False instead of raising,
and resume fails fast if the base v23 table cannot be re-created.
Orphan-marker repair only resets a missing fts_rebuild_progress to 0 once
the index is known empty: the chunk worker replays its whole selected id
range without an anti-join, so a partially indexed DB that lost only its
progress key is first reset to a known-empty surface, then rebuilt.
Ported onto the SessionDB mixin split (hermes_state_search.py /
hermes_state_schema.py).
Installs whose state.db reached schema_version >= 22 before the task
dimension was added carry a 5-column PRIMARY KEY on
session_model_usage. The column reconciler ADDs task as a bare
nullable, but SQLite cannot ALTER a primary key, and the version-gated
v22 rebuild is unreachable (current_version < 22 already false), so
the composite 6-column key never lands. Every upsert in
_record_model_usage then fails with 'ON CONFLICT clause does not match
any PRIMARY KEY or UNIQUE constraint', aborting the enclosing write
transaction — token/cost accounting permanently dead (#73823).
Add an idempotent _heal_session_model_usage_pk() modeled on
_heal_gateway_routing_pk(), run unconditionally from _init_schema on
every open. Salvaged from #73838 with fix-ups:
- ported to SessionSchemaMixin in hermes_state_schema.py (the schema
code moved out of hermes_state.py in 21c7ae8563; the PR targeted the
old location)
- rebuild wrapped in a PRAGMA foreign_keys=OFF/ON window: the
connection enables FKs before _init_schema and OR IGNORE does NOT
suppress FK violations, so a single orphaned usage row (session
pruned while accounting was broken) would have aborted the heal
- COALESCE('') on the nullable reconciler-added task column (and the
billing columns) during the copy
- stale-v22+ regression tests: rebuilt PK + restored upsert, orphan
rows survive the FK window, healthy-DB no-op, no legacy leftover
Fixes#73823