hermes_state emits a once-per-process SQLite-version advisory on CI's linked
3.50.4, which caplog captured as a second WARNING. Scope the assertion to the
hermes_cli.web_server logger the router actually writes to.
The dashboard polls /api/analytics/usage and /api/analytics/models every
few seconds. When state.db is malformed the read raised straight through
the handler, so uvicorn logged a full traceback at ERROR on every poll —
one fleet host wrote ~520K identical journal entries in 24 h.
Wrap both analytics handlers in `corrupt_store_as_status`: a corrupt-image
sqlite3.DatabaseError (is_malformed_db_error) becomes a 503 with an
explicit `state_db_corrupt` payload pointing at `hermes doctor`, and the
warning is gated per store path via `{path: monotonic}` (>=300 s), then
debug. Busy/locked and every other error propagate unchanged, and the file
is never renamed or quarantined from the dashboard — repair stays with
`hermes doctor` / `hermes sessions repair`. `_session_db_path_for_profile`
is split out of `_open_session_db_for_profile` so the router can name the
store without opening it.
Refs #96591
Reported-by: #96591
`hermes sessions repair --check-only` printed the corruption reason and
exited 0, so scripts and the console wrapper gating on the status read a
broken state.db as healthy. Return 1 from the CLI handler (main.py already
sys.exits a truthy return) and from the console handler, where
`_capture_output` turns the status into a ConsoleCommandError carrying the
printed reason.
Salvaged from PR #103321 (the check-only reporting part only; the probe
rewrite and connection-tracking changes were not taken). Refs #63386.
`_db_opens_cleanly` drove one probe row through the messages_fts* triggers
and rolled back. FTS5 only buffers that row in an in-memory segment until
commit, so the probe never wrote to `<fts>_idx`/`_data` and could not hit a
stale `messages_fts_trigram_idx` row waiting at the next segid — the class
where PRAGMA integrity_check, the FTS5 integrity-check command and MATCH all
report clean while every committed append fails with
`IntegrityError: constraint failed`. The probe also caught only
OperationalError; IntegrityError is a DatabaseError sibling, so even a
colliding probe would have escaped and been reported as healthy.
Now the probe issues `INSERT INTO <fts>(<fts>) VALUES('flush')` for every
FTS family inside the rolled-back transaction (capability / not-built errors
stay benign), catches sqlite3.DatabaseError, and always rolls back in a
finally. `hermes doctor` and `hermes sessions repair --check-only` surface
the corruption and `repair_state_db_schema` heals it via the FTS rebuild
strategy (verified with a real stale-segid fixture).
Refs #100227
Reported-by: #100227
The recovery commands rendered on structural corruption — the turn explainer's
`session_persistence_failed`/corrupt body, the gateway's home-channel state.db
warning, and hermes_state_repair._persistent_repair_exhausted_error — already
interpolate the active profile's state.db path, but every `hermes ...` verb in them
was bare. A bare `hermes` follows the sticky `active_profile` file, so an operator
running the pasted `hermes doctor --fix` (or `hermes sessions recover` with a
relative source) from a named-profile incident could inspect or repair a different
profile's database (#105887).
hermes_constants.profile_cli_selector() renders `-p <name> ` for a named profile
home (default home and custom roots outside the profile tree render nothing: the
default is what a bare `hermes` already means, and a custom root is only reachable
via HERMES_HOME). Every command in the three guidance sites now carries it, and
the new `fts_index` guidance inherits the same interpolation.
Live check with HERMES_HOME=<root>/profiles/research and active_profile=other:
before `1. Run \`hermes doctor --fix\`` (targets "other"); after
`1. Run \`hermes -p research doctor --fix\`` and `hermes -p research sessions
recover --source <root>/profiles/research/state.db --inspect-only`.
Refs #105887
Reported-by: Cuttingwater
`hermes doctor` reported every write-health-probe failure as "state.db FTS write
corruption" and `--fix` ran the FTS repair ladder — rebuild, REINDEX, sqlite_master
surgery + VACUUM — on the damaged file. When the damage is structural (canonical
tables/indexes), none of those rungs can fix it, each one writes to the torn file in
place, and the operator is then told to "restore from the backup copy beside
state.db": a `.malformed-backup` that is a snapshot of the same corrupt image.
`hermes sessions recover`, the tool that actually rebuilds canonical rows into a
fresh file, was never mentioned (#88587; the 1.7 GB field incident lost days to it).
Discriminate before mutating. hermes_state_repair.integrity_damage_is_structural
maps `PRAGMA integrity_check` output onto the file: a `Tree N` id resolved through
sqlite_master.rootpage, an index named in `row N missing from index X`, or a
`Freelist:` line is structural unless the object is a Hermes-owned messages_fts*
table/shadow (full-matched, so a user lookalike such as archive_fts_data is never
swept into the rebuildable set). state_db_has_structural_damage runs it read-only on
a fresh connection; an integrity_check that RAISES under the walk (torn root page)
is structural too — no FTS-only fixture does that while sessions/messages read
cleanly. doctor's state check consults it first: structural damage becomes a
manual issue naming `hermes [-p <profile>] sessions recover --source <this db>
--inspect-only` (profile pinned, #105887) and explicitly warning off the
.malformed-backup; nothing is mutated and no backup is written. FTS-only damage
keeps the existing in-place repair path.
Verified against real fixtures: a torn `sessions` root page (before: "FTS write
corruption", --fix wrote a 1:1 malformed-backup and failed; after: structural,
recover guidance, no writes) and the 16-byte DEADBEEF messages_fts_data stomp
(still repaired in place via rebuild_fts).
Salvaged from PR #88604 (liuhao1024) onto the split doctor_state.py; the
classifier lives beside the repair ladder in hermes_state_repair so the ladder
itself can consult it next.
Fixes#88587
classify_persistence_error bucketed every _DB_CORRUPTION_MARKERS hit as "corrupt",
so an error SQLite itself scoped to the FTS5 index layer (SQLITE_CORRUPT_VTAB, or an
`fts5: corrupt structure record for table "messages_fts"` report) that escaped the
write path — the detach in _enter_fts_fail_open refused (generation/lock check),
or a read/search path with no fail-open at all — reached the turn boundary and the
gateway startup notice as structural corruption: the turn ended with `.recover` /
restore-backup advice on a file whose canonical tables were provably healthy.
One provenance rule, hermes_state_errors.is_fts_scoped_corruption_error, now feeds
both the write-repair gate (SessionDB._is_fts_write_corruption_error delegates to it,
so the gateway transcript retry inherits it) and the classifier: a known result code
outranks prose (only SQLITE_CORRUPT_VTAB is FTS-scoped; bare SQLITE_CORRUPT/NOTADB
and any contradictory code fail closed), and without a code the text must both carry
a corruption marker and name a messages_fts* object. The new "fts_index" cause
renders index-scoped guidance (doctor --fix / restart, do not run recovery) in the
turn explainer and the home-channel notice. The structural fail-close is untouched:
bare malformed / not-a-database still quarantine and still classify "corrupt".
Salvaged from PR #97843 (SulthanZahran1), trimmed: the quick_check-backed
"corrupt_unconfirmed" tier is dropped — on a live handle that just observed an
unscoped SQLITE_CORRUPT, PRAGMA quick_check on a damaged shadow b-tree raises rather
than reports on 3.53.1, so the probe could never downgrade the exact shape it was
built for, and a verdict that softens quarantine guidance on prose alone weakens the
fail-close. #97841 (Finn763) reached the same fts_index cause via text markers
only; its LIKE-degradation intent already lives in _search_messages_impl (_fts_stale).
Fixes#97794
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
optimize_fts() and vacuum() already refuse to run against a quarantined
handle (_db_corrupt / _db_replaced / _db_wal_generation_lost): both would
rewrite index/file pages in place, turning contained, diagnosable
corruption into an amplified one. rebuild_fts() never got the same guard,
despite being the more destructive of the two ("discards and recreates the
index data entirely", per its own docstring, vs. optimize_fts's segment
merge).
It's also independently reachable outside _execute_write's own quarantine
check: gateway/session_transcript.py's _rebuild_fts_once() calls
db.rebuild_fts() directly from the FTS-corruption transcript-retry path,
with no quarantine check of its own (only a WAL split-brain / foreign-holder
check, a different concern). A quarantined handle hitting that retry path
would run a full FTS rebuild — and commit it — on a corrupt, replaced, or
split-WAL-generation file.
Add the same self._raise_if_db_corrupt()/self._raise_if_db_replaced() pair
optimize_fts() already has, at the top of rebuild_fts(), before it enters
the cross-process rebuild admission.
A SIGKILL mid-write (OOM killer, #106667) can leave state.db with a garbage
page-1 header. SQLite refuses the file outright ("file is not a database",
SQLITE_NOTADB) and the sqlite3 shell's .recover opens the file like any other
client, so the lost_and_found lane failed with the same rc=26 twice although
every data page after the header survived. The #106587 quarantine now
preserves such a file as state.db.notadb-<ts>-<pid>.bak; this makes that
preserved file recoverable with `hermes sessions recover --source <bak>
--allow-partial`.
When both .recover attempts fail with "not a database", zero the 100-byte
header of the lane's private snapshot copy and rerun them. .recover trips
only on the magic check and infers page size and layout from the pages
themselves, so a zeroed header is enough; a spliced donor header (the PR's
original mechanism) instead advertises a database size / freelist that
contradicts the file and yields "database disk image is malformed" on a
direct open — verified live on a 139-page fixture, which also showed the
zeroed header recovers 60/60 sessions and 300/300 messages whether the
damage covers 100 bytes or the whole first page. The user's file is never
written; the report carries `sqlite3_cli.header_zeroed` and a warning about
the WAL boundary.
Live repro (sqlite3 shell 3.53.1 on PATH, header overwritten with random
bytes): BEFORE "page-level .recover salvage failed: ... file is not a
database (26)"; AFTER "Recovered 60 sessions and 300 messages", source md5
unchanged.
Salvaged from PR #102808 (intent; trimmed from 302 to ~40 source LOC by
dropping the donor-header/page-size sweep and the redundant preopen probe —
the shell's own refusal is the detector). Independent review on the PR by
@strzhao.
Refs #106667
Refs #106587
Reported-by: TaoMasterCoder
Cross-referenced-by: kshitijk4poor
Review feedback on #103657: the sibling clauses in the same statement
declare ESCAPE, and the trash enumeration two blocks up escapes its
underscores via replace. Use the same escaped ESCAPE form so every
underscore in the statement is a literal match instead of a
single-character wildcard. Behavior on the fixed schema is unchanged
(same enumeration split); this is consistency plus defense against
future lookalike table names.
Also reword the exclusion comment to the verified failure mechanism:
fts5 xRename renames the whole shadow family in one step, so sweeping
the cjk vtable aborts the loop on the next shadow entry and drags the
_config table (read by the vtable constructor) into the trash family.
The demote enumeration (name LIKE 'messages_fts_%') sweeps the
messages_fts_cjk vtable and its shadow tables into the fts_v22_trash_*
renames. Renaming the cjk vtable cascades to its shadow tables and breaks
the vtable constructor chain, so 'hermes sessions optimize-storage'
aborts with 'vtable constructor failed: messages_fts_cjk' on every DB
that carries both a legacy inline FTS layout and an established cjk
index (#103647). The cjk family is an independent v23+ index, not part
of the demoted legacy layout: skip it in the enumeration.
When the leftmost (or rightmost) leaf of a table b-tree is damaged, the edge
probe `SELECT rowid ... ORDER BY rowid ASC LIMIT 1` walks the table tree and
raises, and _salvage_rowid_bounds fell back to INT64_MIN. The gallop from the
surviving edge cannot cap that side either (every probe crosses the damaged
leaf), so bisection burned the entire 10,000-query budget moving the bound
inward by a few thousand rowids out of 9.2e18 and the table was lost — a
4-row gateway_routing table in #98050, sessions + session_model_usage in
#100313.
`SELECT min(rowid), max(rowid)` is answered by the planner from any covering
index (every Hermes table has at least the PRIMARY KEY autoindex) without
touching the damaged leaf, which is exactly what the reporter verified by
hand. Ask it for the missing edge(s) first; only when it fails too does the
domain fallback + gallop run as before. Reported under `aggregate_edges` so
recovery.json still shows how the bound was obtained.
Live repro (real fixture: leftmost `sessions` leaf cell count overwritten,
400 rows): BEFORE bounds low=-9223372036854775808 copied=0 range_queries=10000
query_limit_reached=True status=failed; AFTER low=1 high=400 copied=391
range_queries=40 status=partial (only the damaged leaf's rows are lost).
Refs #98050
Refs #100313
Reported-by: Ace-Kelly
Corroborated-by: Proff506
`hermes sessions recover --allow-partial` walks damaged tables by rowid range
and falls back to an exact-rowid lookup for the last cell of a broken range.
A phantom row produced by page damage (a `sessions` row with a NULL
`started_at`, an `async_delegations` row with a NULL `state`) reads fine from
the source but violates the destination's NOT NULL constraint, and the
resulting sqlite3.IntegrityError escaped the exact-lookup boundary and aborted
the whole recovery — losing every healthy row behind it.
Catch IntegrityError at that boundary, count it under
`destination_rejected_rows` and record the rowid as a skipped singleton
("destination constraint rejected row: ...") so the run completes, verifies,
and the report shows exactly which rows were dropped.
Live repro (real fixture: source schema's NOT NULL relaxed via
writable_schema, phantom `sessions` row inserted): BEFORE
sqlite3.IntegrityError "NOT NULL constraint failed: sessions.started_at" at
session_recovery.py recover_exact_rowid; AFTER status=partial copied=3
destination_rejected_rows=1, verified=True. Reporter @i8ei confirmed the same
patch recovers the field database (22 sessions / 2,248 messages) in #102240.
Salvaged from PR #91413 (rebased onto the _RowidRangeSalvage refactor; test
reduced to one invariant).
Refs #102240
Reported-by: i8ei
`sqlite3 state.db .recover` re-emits the FTS5 shadow tables (messages_fts_data,
_idx, _docsize, _config, _content) as ordinary tables but cannot re-emit the
CREATE VIRTUAL TABLE row. The next SessionDB open ran the FTS DDL in
_ensure_fts_schema and died with "fts5: error creating shadow table
messages_fts_data: table 'messages_fts_data' already exists", so a recovered
database was unusable until someone hand-dropped the shadows.
_init_fts now runs _drop_orphan_fts_shadow_tables before any FTS DDL. It is
per-family and exact-name scoped: a family's shadows are dropped only when its
own vtable row is absent from sqlite_master (type='table' AND sql LIKE
'CREATE VIRTUAL TABLE%'), so a healthy messages_fts_trigram survives a
base-family repair untouched. A repaired base/trigram family is then treated
like a missing-trigger repair and rebuilt from the canonical messages table
under the cross-process rebuild admission; the shadows are derived index
state, nothing is lost.
Live repro: real `sqlite3 x.db .recover | sqlite3 y.db` on sqlite 3.50.4
keeps the vtable rows (the shell emits CREATE VIRTUAL TABLE), so the
deterministic fixture removes the vtable row via writable_schema leaving the
shadows behind: BEFORE OperationalError on open; AFTER opens, fts_enabled,
MATCH returns every message, trigram sqlite_master rowids unchanged.
Salvaged from PR #56824 (intent applied onto the current hermes_state_fts /
hermes_state_schema siblings). The ownership-safety point (never touch a live
family's shadows) was raised by @ggoldani in #103868 / #103840.
Refs #103840
Refs #56815
Co-authored-by: ggoldani <ggoldani@users.noreply.github.com>
The _sql_json_extract wrapper removed the `COALESCE(model_config` text the
alias string-replace keyed on, so the deferred-backfill SELECT joined an
unqualified `model_config`. Build the predicate from the alias directly and
derive the unaliased constant from it, so the two can never disagree.
Sibling widening of the #101726 salvage: FTS_TRIGRAM_SESSION_SQL (trigram view/triggers/backfill),
the v16 delegate-tagging data migration and reopen_session's legacy reset-child stamp still called
json_extract() on the raw model_config cell, so one malformed JSON row could still abort FTS
maintenance, a schema migration or /resume of a reset child. Zero raw model_config json_extract
reads remain in hermes_state_*.py.
Three invariants over real SQLite fixtures (TEXT and 8.4e252 timestamps written straight into the
REAL columns; 1200 ids under a 999-variable ceiling via setlimit): list/export/insights complete
and name the corrupt session in a WARNING; writers never persist an out-of-window timestamp;
prune/delete_sessions succeed with zero orphaned messages. All three fail on origin/main.
prune_sessions(), delete_empty_sessions() and prune_empty_ghost_sessions()
built one IN (?, ..., ?) clause containing every selected session id (the
parent-orphaning UPDATE), so cleaning more than SQLITE_MAX_VARIABLE_NUMBER
sessions failed atomically with "too many SQL variables". The per-row
DELETE loops that followed are folded into the same 900-id batches.
Same single _execute_write() transaction; only the binding is split.
Hand-ported from PR #100658 (targeted the pre-decomposition hermes_state.py
god file; the methods now live in hermes_state_maintenance.py /
hermes_state_sessions.py). The one-pass transcript-directory sweep from
that PR is not ported (out of scope for the variable-limit bug). Authored
by @Mi55ed; ported under --author.
`hermes sessions prune --source cron --older-than 14` on a store with ~60K
cron sessions (~50K matches) died with sqlite3.OperationalError: too many
SQL variables. SessionDB.delete_sessions, _collect_delegate_child_ids and
_delete_delegate_children each bound the full id list into a single
IN (?,?,...). SQLite caps bound parameters at SQLITE_MAX_VARIABLE_NUMBER
(999 on < 3.32, 32766 after), so any bulk delete above that failed outright.
Chunk every IN list (`_id_chunks` / `_SQL_IN_CHUNK` in hermes_state_common,
900 ids; the delegate walk binds each id twice so it chunks at half). Same
transaction, same cascade/orphan contract; only the parameter binding is
split.
Hand-ported from PR #102679 (targeted the pre-decomposition hermes_state.py
god file; the functions now live in hermes_state_sessions.py). Authored by
@mssteuer; ported under --author.
SQLite dynamic typing lets a TEXT cell ('not-a-timestamp'), inf/nan or a
garbage double (8.4e252 salvaged from a damaged page) sit in a REAL
timestamp column. Every reader called datetime.fromtimestamp()/float
arithmetic on the raw cell, so ONE bad row raised TypeError/OverflowError
out of the row loop and took down the whole `hermes sessions list`/browse
table (#102399), all three exporters — JSONL/MD, QMD, HTML (#102352) —
and `hermes insights` (#99959).
Fix the class with ONE helper, hermes_cli.timefmt.coerce_epoch(): a
stored cell becomes float epoch seconds inside a sane 1970..2103 window
or None after a WARNING that names the session id. Every reader routes
through it — relative_time (list/browse/resume picker), format_epoch
(prune/candidates tables), the three exporters' timestamp formatters,
insights' _get_sessions/_day/period range — so a bad row renders as
'?'/'N/A'/raw text for that one cell and the command completes.
Write side: hermes_state_messages._coerce_timestamp (append_message,
append_messages_batch, import) and the import path's started_at now use
the same window, so a new out-of-range timestamp falls back to now()
instead of being persisted — new bad rows cannot be written by Hermes.
Reported-by: #102399, #102352, #99959 reporters; kokhlo's insights
analysis pointed at every reporting site, not just line 860.
The schema initializer now replays SCHEMA_SQL through executescript (the
single-authority path from #94701's follow-up), which bypasses the
execute()-level DDL failure injection — the regression stopped raising.
Bind the same simulated failure onto the executescript path so the
connect-close-on-init-failure contract stays pinned for both replay
mechanisms.
Review follow-up on #94701: the delegation tool's _initialize_schema
still carried its own CREATE TABLE + ALTER column list for
async_delegations, leaving a second durable-shape authority even with
the column declared in SCHEMA_SQL. Its legacy ALTER added
origin_session_id as bare TEXT (nullable, no default); reconciliation
repairs missing column names only, so a database first opened through
the tool kept a non-canonical shape forever (#94691).
Remove the private DDL entirely. The tool's initializer now calls a new
reconcile_state_schema() in hermes_state_schema, which replays the
canonical SCHEMA_SQL (idempotent CREATE IF NOT EXISTS for every table,
canonical indexes included) and reuses SessionDB's declarative
_reconcile_columns for missing-column backfill — one reconciliation
implementation, one authority. Because _parse_schema_columns
reconstructs each column's full constraint expression (type, NOT NULL,
DEFAULT), the tool-first legacy path now adds origin_session_id as
TEXT NOT NULL DEFAULT '' — the canonical shape — and SQLite backfills
existing rows with the '' default.
Opening-order regressions compare FULL PRAGMA table_info metadata
(type, notnull, dflt_value, pk) plus the canonical index set across
fresh SessionDB→tool, legacy→SessionDB, and legacy→tool→SessionDB,
each preserving a pre-existing legacy delegation row.
The delegation tool carries its own CREATE TABLE for async_delegations
(tools/async_delegation.py _initialize_schema) plus a lazy ALTER TABLE
ADD COLUMN for the tables it finds already existing. Its column list
had drifted ahead of the canonical SCHEMA_SQL: origin_session_id
(raw api_server session id of the originating request, the wake
self-post target) existed only through the tool's lazy path, so two
databases at the same schema_version had different
async_delegations shapes depending solely on whether the delegation
tool had ever run. Rebuild/replay pipelines that reconstruct state.db
from the canonical schema then hit the column with no version gate to
explain it (#94691).
Declare the column in SCHEMA_SQL with the same TEXT NOT NULL DEFAULT ''
shape the tool uses. Fresh installs now carry it canonically; the
declarative _reconcile_columns backfills it into legacy databases on
the next writable open (same pattern as earlier additive columns); the
tool's lazy ALTER keeps serving pre-reconciliation databases. The two
schema authorities now agree, pinned by a test that runs the tool's
initializer over a canonical database and asserts the shape is
unchanged.
Fixes#94691
The corruption explainer filled `{db_path}` from `_default_db_path()`, the
process default. A Desktop `serve` backend launched on the root home hosts
named-profile sessions whose SessionDB is `profiles/<name>/state.db`, so the
operator was told to inspect/repair a different profile's database. Pass the
agent's own `_session_db.db_path` from the turn finalizer; the process
default remains the fallback for agents without a bound store.
Reported in #105887.
Reading a session by id that missed the caller's store fell through to
_locate_session_db(), which opened every profile's state.db read-only and
returned the first owner's full transcript — no opt-in, no profile named, and
the miss path even fired after an explicit non-matching profile= read. Any
caller holding an id (ids appear in logs and tool output) could read a
foreign profile's conversation. Profiles are isolated islands by design.
A miss now stays a miss, with a hint to name the owning profile
(profile=<name> / @session:<profile>/<id>), which remains the sanctioned,
explicit cross-profile read. The schema eval runner no longer needs to fake
the scan.
Reported by the #106761 filer; reproduced by @kokhlo. Refs #87779.
The salvaged #102534 test patched hermes_state.get_shared_session_db, a seam
main dropped (server._get_db now calls hermes_state_registry.acquire), so the
fixture errored at setup and the file reported 0 passed. Assert on the real
handle's db_path and on the foreign home staying untouched instead of on a
fake factory; red with the pin reverted, green with it.
The lazy _get_db() singleton followed get_hermes_home(), so a first touch
inside the multiplex cron ticker's per-profile override window permanently
bound the default backend to another profile's state.db (#102526).
close_all_under returned after the last release dropped the generation
and before the physical close finished, so rmtree still saw the open
handle. Wait directory-matching teardown barriers the same way close_all
does.
delete_profile already force-closes holographic memory_store.db in this
process, but the shared SessionDB registry kept state.db open. Recreate
then failed with a replaced/locked database. Close every shared handle
under the doomed directory, same contract as MemoryStore.release_all_under.
Co-authored-by: Cursor <cursoragent@cursor.com>
Since 0.21.0 reads go through mode=ro pooled connections. A read-only OPEN already
rides out the millisecond WAL transition window (checkpoint / WAL reset / frame flush
by a sibling process; the ro reader cannot rewrite the -shm index) with a bounded retry
(#100436), but a WARM pooled reader hitting the same window while its SELECT executes
propagated `disk I/O error` straight out of get_session(): 37 identical tracebacks on a
multi-process WSL2 ext4-on-vhdx install, each followed by "compression session recovery
failed", with quick_check=ok (#100871). The reporter's A/B shows the operator
workaround (journal_mode=delete) collapses read throughput ~30000x, so the flake has to
be absorbed on the read path.
_read_one/_read_all now replay the idempotent statement within the existing read-only
IOERR budget (3 x 50 ms) on the SAME connection -- close+reopen would cancel this
process's POSIX locks for every sibling connection -- and a persistent IOERR still
propagates. No quarantine: EIO on a read is busy, not broken. Every SELECT in the
SessionDB siblings (63 call sites) reaches the pool through these two helpers, so the
class is covered without a wrapper type.
Same-connection retry per #100882's analysis (@fangliquanflq); #100883
(@Sahilvishnaliya) diagnosed the missing recovery in the 0.21.0 read pool.
Fixes#100871.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Sahilvishnaliya <222165401+Sahilvishnaliya@users.noreply.github.com>
apply_wal_with_fallback() reports "wal" in two indeterminate cases -- the vulnerable-
SQLite gate (_apply_delete_for_wal_reset_bug) and the non-vulnerable probe-unknown path
(a7f2a593d1) -- meaning "touched nothing, the connection inherits the header's mode".
SessionDB turned that assumption into `_wal_active=True`, which enables the mode=ro read
pool that skips `self._lock`. On a file that is really in rollback-journal mode those
readers race the writer with a 5s busy timeout and no retry: random SQLITE_BUSY read
failures for the instance's lifetime (#86515).
Confirm the header on the freshly opened connection before enabling the pool. When the
probe is still blocked, reads queue on the writer connection under the lock -- slower,
never wrong. Every other apply_wal_with_fallback caller ignores the return value, so the
consumer is the right place to gate; changing the return contract to Optional across
15 call sites (#87044's shape) is not needed.
Live repro: DELETE-mode file, sibling holding BEGIN EXCLUSIVE during open ->
before: _wal_active=True and _checkout_read_conn() hands out a pooled mode=ro conn;
after: _wal_active=False, reads take the locked writer path.
Fixes#86515. Based on the analysis in #87044.
Co-authored-by: QDung210 <dqdung205@gmail.com>
_send_session_db_warning_notifications() broadcast the error recorded at startup
without asking whether it was still true. A startup `database is locked` routinely
clears while the adapters are still connecting (another profile's open, a `hermes
sessions` one-shot, a slow SMB lock release), so the home channels were told the
store was unavailable when it had already healed — and a warning that is wrong once
is ignored the next time it is right.
The RecoverableHandleCache opener already clears `_session_db_init_error` on
recovery; the broadcast now drives one open attempt through it first and only warns
when the error is still standing. A store that is genuinely still down warns exactly
as before.
Fixes#108031.
main replaced the file-local _make_db/_require_wal/_unlink_sidecars helpers with
tests/hermes_state/_wal_generation_harness (make_db/require_wal/lose_sidecars) after
PR #107411 branched; the cherry-picked tests referenced the old names and failed at
collection with NameError. Fixture-only change; the assertions are unchanged.
Classify deleted WAL generations separately from main-file replacement and point operators at the captured-generation manifest and mode-aware recovery path.
Co-authored-by: crazyief <8566250+crazyief@users.noreply.github.com>
A lock-free _raise_if_db_replaced() probe at the top of the _execute_write
retry loop raced a concurrent close(). close() runs under the same _lock and
ends the WAL generation: it checkpoints, closes the connection (SQLite
unlinks the -wal/-shm sidecars), nulls _conn and clears
_db_sidecar_identity. The probe could observe the mid-teardown state —
sidecars already unlinked while _db_sidecar_identity was not yet cleared —
and misclassify this process's OWN clean close as an externally deleted WAL
generation, raising a sticky DeletedWalGenerationError that permanently
refused every later write on that handle (#105567).
Move the live probe inside the lock, ahead of the close-race reopen
decision, so it only ever observes the stable post-close state (identity
cleared -> the existing adopt/reopen path). The corrupt flag check stays on
the lock-free fast path; external file/generation replacement detection is
unchanged, just serialized with teardown.
Synthetic repro (100 rounds x 40 writes, direct SessionDB handles): before
~9 failing rounds / ~360 DeletedWalGenerationError; after 0 failures,
4000/4000 writes persisted across repeated runs. tests/state (181) plus the
generation/replaced/corrupt guard suites (55) pass.
Fixes#105567
st_nlink == 0 alone cannot distinguish a genuine orphan from one that
still has a surviving hard link (e.g. a backup) after the watched
sidecar path itself was removed or replaced — that left st_nlink >= 1
on a truly orphaned generation, letting a new opener through while a
live writer still owned the old one. Compare (st_dev, st_ino) between
the fd and the current watched sidecar path instead: only an exact
match means they're the same live file, so any mismatch or unstattable
watched path still fails closed.
iter_deleted_sqlite_sidecar_holders() and SessionDB._wal_generation_was_lost()
both treated a `` (deleted)`` suffix on a /proc/<pid>/fd/* target as proof that
state.db-wal or state.db-shm was unlinked. On OpenZFS that suffix is not proof:
a live, still-linked file whose dentry was unhashed is reported the same way,
with st_nlink still 1 and the same (dev, ino) as the path. The guard then fires
permanently and the gateway falls back to JSONL forever, because the WAL was
never actually deleted.
Add _fd_is_truly_unlinked(), which confirms via os.stat(fd_path).st_nlink == 0
before a target counts as an orphaned generation. An unstattable descriptor
still counts as deleted, so the guard keeps failing closed. _iter_proc_fd_targets()
and _proc_fd_targets() now also yield the /proc fd path itself so both call
sites (open-path and the sticky write-path probe) can run the check.
Drop two PrimaryProfilePin cases that only restate the constructor
defaults and blank-string normalisation, and the wiring-routing test that
froze POOL_LIMITS_SETTINGS_ROUTE to a literal string — a snapshot of the
constant, not a behaviour contract. The two kept pin tests cover the bug
(a live primary keeps answering for its booted profile after the stored
preference moves; teardown releases the pin), and the notifications tests
cover the toast action end-to-end.
`primaryProfileKey()` re-read active-profile.json on every call. The rail's
live workspace switch rewrites that file via `hermes:profile:remember`
WITHOUT re-homing the primary, so after a switch the routing table disagreed
with the running process: a request for the profile the primary actually
booted as (e.g. "default") no longer matched `primaryProfile` in
`resolveProfileBackendRoute`, fell through to the pool, and spawned a second
backend for the same HERMES_HOME.
The duplicate was keepalive-fresh so LRU eviction spared it, it burned a pool
slot, and with the default cap of 3 every further profile queued and failed
with `Local backend start for "<profile>" timed out while waiting for a free
slot` (repro in desktop.log: "default" spawned as a pool backend while the
primary "default" was still running; coder/qwen then timed out for 20+ min).
Snapshot the launch profile in `startHermes()` (PrimaryProfilePin.pin) and
release it in `resetHermesConnection()` so the next start follows the stored
preference again. The pin is a tiny pure module with tests; main.ts only owns
the file read and the two call sites.
Switching the pooled dispatch probe to /api/health (salvaged from #97914)
would 404 on every dispatch against a remote older than 0.19, retire the
tunnel and reconnect forever - the same storm #107997 describes, moved to
old backends. Fall back to /api/status on an explicit 404 exactly the way
the boot readiness probe already does (backend-health.ts). The legacy
fallback idea and its test are taken from #101976 (@edosulai); the rest of
that PR (timeout-tolerance streak, ServerAlive SSH options) is not adopted.
Co-authored-by: Edo Sulaiman <edosulai@icloud.com>
Cold /api/status through a Windows no-mux SSH forward routinely exceeds
the 2.5s dispatch budget, so Desktop retires a live tunnel and respawns.
Use the cheap /api/health route (5s, same as DEFAULT_HEALTH_PROBE_TIMEOUT_MS).
Background liveness still probes /api/status at 10s.