Commit Graph

33651 Commits

Author SHA1 Message Date
KoNit-K 0b8daf30aa fix(bootstrap-installer): stamp setup app version from release semver
Fixes #107177

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-11 16:10:47 +02:00
Teknium 04dd80a977 test(web): count only the dashboard's own warning in the corrupt-store poll test
hermes_state emits a once-per-process SQLite-version advisory on CI's linked
3.50.4, which caplog captured as a second WARNING. Scope the assertion to the
hermes_cli.web_server logger the router actually writes to.
2026-09-11 06:37:27 -07:00
Teknium fdd32f81d4 fix(web): report a corrupt state.db as a throttled 503 status instead of a traceback per poll
The dashboard polls /api/analytics/usage and /api/analytics/models every
few seconds. When state.db is malformed the read raised straight through
the handler, so uvicorn logged a full traceback at ERROR on every poll —
one fleet host wrote ~520K identical journal entries in 24 h.

Wrap both analytics handlers in `corrupt_store_as_status`: a corrupt-image
sqlite3.DatabaseError (is_malformed_db_error) becomes a 503 with an
explicit `state_db_corrupt` payload pointing at `hermes doctor`, and the
warning is gated per store path via `{path: monotonic}` (>=300 s), then
debug. Busy/locked and every other error propagate unchanged, and the file
is never renamed or quarantined from the dashboard — repair stays with
`hermes doctor` / `hermes sessions repair`. `_session_db_path_for_profile`
is split out of `_open_session_db_for_profile` so the router can name the
store without opening it.

Refs #96591
Reported-by: #96591
2026-09-11 06:37:27 -07:00
jango 5a2f512390 fix(cli): exit non-zero from sessions repair --check-only on an unhealthy store
`hermes sessions repair --check-only` printed the corruption reason and
exited 0, so scripts and the console wrapper gating on the status read a
broken state.db as healthy. Return 1 from the CLI handler (main.py already
sys.exits a truthy return) and from the console handler, where
`_capture_output` turns the status into a ConsoleCommandError carrying the
printed reason.

Salvaged from PR #103321 (the check-only reporting part only; the probe
rewrite and connection-tracking changes were not taken). Refs #63386.
2026-09-11 06:37:27 -07:00
Teknium d8cf5da7ac fix(state): make the FTS write-health probe flush segments and catch IntegrityError
`_db_opens_cleanly` drove one probe row through the messages_fts* triggers
and rolled back. FTS5 only buffers that row in an in-memory segment until
commit, so the probe never wrote to `<fts>_idx`/`_data` and could not hit a
stale `messages_fts_trigram_idx` row waiting at the next segid — the class
where PRAGMA integrity_check, the FTS5 integrity-check command and MATCH all
report clean while every committed append fails with
`IntegrityError: constraint failed`. The probe also caught only
OperationalError; IntegrityError is a DatabaseError sibling, so even a
colliding probe would have escaped and been reported as healthy.

Now the probe issues `INSERT INTO <fts>(<fts>) VALUES('flush')` for every
FTS family inside the rolled-back transaction (capability / not-built errors
stay benign), catches sqlite3.DatabaseError, and always rolls back in a
finally. `hermes doctor` and `hermes sessions repair --check-only` surface
the corruption and `repair_state_db_schema` heals it via the FTS rebuild
strategy (verified with a real stale-segid fixture).

Refs #100227
Reported-by: #100227
2026-09-11 06:37:27 -07:00
Teknium 754ecff466 fix(state): pin corrupt-session recovery guidance to the failing profile
The recovery commands rendered on structural corruption — the turn explainer's
`session_persistence_failed`/corrupt body, the gateway's home-channel state.db
warning, and hermes_state_repair._persistent_repair_exhausted_error — already
interpolate the active profile's state.db path, but every `hermes ...` verb in them
was bare. A bare `hermes` follows the sticky `active_profile` file, so an operator
running the pasted `hermes doctor --fix` (or `hermes sessions recover` with a
relative source) from a named-profile incident could inspect or repair a different
profile's database (#105887).

hermes_constants.profile_cli_selector() renders `-p <name> ` for a named profile
home (default home and custom roots outside the profile tree render nothing: the
default is what a bare `hermes` already means, and a custom root is only reachable
via HERMES_HOME). Every command in the three guidance sites now carries it, and
the new `fts_index` guidance inherits the same interpolation.

Live check with HERMES_HOME=<root>/profiles/research and active_profile=other:
before `1. Run \`hermes doctor --fix\`` (targets "other"); after
`1. Run \`hermes -p research doctor --fix\`` and `hermes -p research sessions
recover --source <root>/profiles/research/state.db --inspect-only`.

Refs #105887
Reported-by: Cuttingwater
2026-09-11 06:37:27 -07:00
liuhao1024 3d167a8271 fix(doctor): name structural state.db corruption honestly and route it to sessions recover
`hermes doctor` reported every write-health-probe failure as "state.db FTS write
corruption" and `--fix` ran the FTS repair ladder — rebuild, REINDEX, sqlite_master
surgery + VACUUM — on the damaged file. When the damage is structural (canonical
tables/indexes), none of those rungs can fix it, each one writes to the torn file in
place, and the operator is then told to "restore from the backup copy beside
state.db": a `.malformed-backup` that is a snapshot of the same corrupt image.
`hermes sessions recover`, the tool that actually rebuilds canonical rows into a
fresh file, was never mentioned (#88587; the 1.7 GB field incident lost days to it).

Discriminate before mutating. hermes_state_repair.integrity_damage_is_structural
maps `PRAGMA integrity_check` output onto the file: a `Tree N` id resolved through
sqlite_master.rootpage, an index named in `row N missing from index X`, or a
`Freelist:` line is structural unless the object is a Hermes-owned messages_fts*
table/shadow (full-matched, so a user lookalike such as archive_fts_data is never
swept into the rebuildable set). state_db_has_structural_damage runs it read-only on
a fresh connection; an integrity_check that RAISES under the walk (torn root page)
is structural too — no FTS-only fixture does that while sessions/messages read
cleanly. doctor's state check consults it first: structural damage becomes a
manual issue naming `hermes [-p <profile>] sessions recover --source <this db>
--inspect-only` (profile pinned, #105887) and explicitly warning off the
.malformed-backup; nothing is mutated and no backup is written. FTS-only damage
keeps the existing in-place repair path.

Verified against real fixtures: a torn `sessions` root page (before: "FTS write
corruption", --fix wrote a 1:1 malformed-backup and failed; after: structural,
recover guidance, no writes) and the 16-byte DEADBEEF messages_fts_data stomp
(still repaired in place via rebuild_fts).

Salvaged from PR #88604 (liuhao1024) onto the split doctor_state.py; the
classifier lives beside the repair ladder in hermes_state_repair so the ladder
itself can consult it next.

Fixes #88587
2026-09-11 06:37:27 -07:00
Sulthan Zahran 38adfe90a4 fix(state): classify FTS-scoped corruption as fts_index, never whole-file damage
classify_persistence_error bucketed every _DB_CORRUPTION_MARKERS hit as "corrupt",
so an error SQLite itself scoped to the FTS5 index layer (SQLITE_CORRUPT_VTAB, or an
`fts5: corrupt structure record for table "messages_fts"` report) that escaped the
write path — the detach in _enter_fts_fail_open refused (generation/lock check),
or a read/search path with no fail-open at all — reached the turn boundary and the
gateway startup notice as structural corruption: the turn ended with `.recover` /
restore-backup advice on a file whose canonical tables were provably healthy.

One provenance rule, hermes_state_errors.is_fts_scoped_corruption_error, now feeds
both the write-repair gate (SessionDB._is_fts_write_corruption_error delegates to it,
so the gateway transcript retry inherits it) and the classifier: a known result code
outranks prose (only SQLITE_CORRUPT_VTAB is FTS-scoped; bare SQLITE_CORRUPT/NOTADB
and any contradictory code fail closed), and without a code the text must both carry
a corruption marker and name a messages_fts* object. The new "fts_index" cause
renders index-scoped guidance (doctor --fix / restart, do not run recovery) in the
turn explainer and the home-channel notice. The structural fail-close is untouched:
bare malformed / not-a-database still quarantine and still classify "corrupt".

Salvaged from PR #97843 (SulthanZahran1), trimmed: the quick_check-backed
"corrupt_unconfirmed" tier is dropped — on a live handle that just observed an
unscoped SQLITE_CORRUPT, PRAGMA quick_check on a damaged shadow b-tree raises rather
than reports on 3.53.1, so the probe could never downgrade the exact shape it was
built for, and a verdict that softens quarantine guidance on prose alone weakens the
fail-close. #97841 (Finn763) reached the same fts_index cause via text markers
only; its LIKE-degradation intent already lives in _search_messages_impl (_fts_stale).

Fixes #97794
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
2026-09-11 06:37:27 -07:00
nftpoetrist 17b2768b86 fix(state): quarantine a corrupt/replaced handle out of rebuild_fts() too
optimize_fts() and vacuum() already refuse to run against a quarantined
handle (_db_corrupt / _db_replaced / _db_wal_generation_lost): both would
rewrite index/file pages in place, turning contained, diagnosable
corruption into an amplified one. rebuild_fts() never got the same guard,
despite being the more destructive of the two ("discards and recreates the
index data entirely", per its own docstring, vs. optimize_fts's segment
merge).

It's also independently reachable outside _execute_write's own quarantine
check: gateway/session_transcript.py's _rebuild_fts_once() calls
db.rebuild_fts() directly from the FTS-corruption transcript-retry path,
with no quarantine check of its own (only a WAL split-brain / foreign-holder
check, a different concern). A quarantined handle hitting that retry path
would run a full FTS rebuild — and commit it — on a corrupt, replaced, or
split-WAL-generation file.

Add the same self._raise_if_db_corrupt()/self._raise_if_db_replaced() pair
optimize_fts() already has, at the top of rebuild_fts(), before it enters
the cross-process rebuild admission.
2026-09-11 06:37:27 -07:00
TaoMasterCoder 10777823fe fix(sessions): salvage a page-1-header-damaged state.db in the lost_and_found lane
A SIGKILL mid-write (OOM killer, #106667) can leave state.db with a garbage
page-1 header. SQLite refuses the file outright ("file is not a database",
SQLITE_NOTADB) and the sqlite3 shell's .recover opens the file like any other
client, so the lost_and_found lane failed with the same rc=26 twice although
every data page after the header survived. The #106587 quarantine now
preserves such a file as state.db.notadb-<ts>-<pid>.bak; this makes that
preserved file recoverable with `hermes sessions recover --source <bak>
--allow-partial`.

When both .recover attempts fail with "not a database", zero the 100-byte
header of the lane's private snapshot copy and rerun them. .recover trips
only on the magic check and infers page size and layout from the pages
themselves, so a zeroed header is enough; a spliced donor header (the PR's
original mechanism) instead advertises a database size / freelist that
contradicts the file and yields "database disk image is malformed" on a
direct open — verified live on a 139-page fixture, which also showed the
zeroed header recovers 60/60 sessions and 300/300 messages whether the
damage covers 100 bytes or the whole first page. The user's file is never
written; the report carries `sqlite3_cli.header_zeroed` and a warning about
the WAL boundary.

Live repro (sqlite3 shell 3.53.1 on PATH, header overwritten with random
bytes): BEFORE "page-level .recover salvage failed: ... file is not a
database (26)"; AFTER "Recovered 60 sessions and 300 messages", source md5
unchanged.

Salvaged from PR #102808 (intent; trimmed from 302 to ~40 source LOC by
dropping the donor-header/page-size sweep and the redundant preopen probe —
the shell's own refusal is the detector). Independent review on the PR by
@strzhao.

Refs #106667
Refs #106587
Reported-by: TaoMasterCoder
Cross-referenced-by: kshitijk4poor
2026-09-11 06:37:27 -07:00
liuhao1024 c816bfab20 fix(state): escape underscores in the cjk family exclusion LIKE pattern
Review feedback on #103657: the sibling clauses in the same statement
declare ESCAPE, and the trash enumeration two blocks up escapes its
underscores via replace. Use the same escaped ESCAPE form so every
underscore in the statement is a literal match instead of a
single-character wildcard. Behavior on the fixed schema is unchanged
(same enumeration split); this is consistency plus defense against
future lookalike table names.

Also reword the exclusion comment to the verified failure mechanism:
fts5 xRename renames the whole shadow family in one step, so sweeping
the cjk vtable aborts the loop on the next shadow entry and drags the
_config table (read by the vtable constructor) into the trash family.
2026-09-11 06:37:27 -07:00
liuhao1024 b811dfd0ee fix(state): exclude the cjk index family from the legacy FTS demote rename
The demote enumeration (name LIKE 'messages_fts_%') sweeps the
messages_fts_cjk vtable and its shadow tables into the fts_v22_trash_*
renames. Renaming the cjk vtable cascades to its shadow tables and breaks
the vtable constructor chain, so 'hermes sessions optimize-storage'
aborts with 'vtable constructor failed: messages_fts_cjk' on every DB
that carries both a legacy inline FTS layout and an established cjk
index (#103647). The cjk family is an independent v23+ index, not part
of the demoted legacy layout: skip it in the enumeration.
2026-09-11 06:37:27 -07:00
Teknium 8a9f9eca2e fix(recovery): seed a damaged rowid edge from min()/max() before the INT64 domain
When the leftmost (or rightmost) leaf of a table b-tree is damaged, the edge
probe `SELECT rowid ... ORDER BY rowid ASC LIMIT 1` walks the table tree and
raises, and _salvage_rowid_bounds fell back to INT64_MIN. The gallop from the
surviving edge cannot cap that side either (every probe crosses the damaged
leaf), so bisection burned the entire 10,000-query budget moving the bound
inward by a few thousand rowids out of 9.2e18 and the table was lost — a
4-row gateway_routing table in #98050, sessions + session_model_usage in
#100313.

`SELECT min(rowid), max(rowid)` is answered by the planner from any covering
index (every Hermes table has at least the PRIMARY KEY autoindex) without
touching the damaged leaf, which is exactly what the reporter verified by
hand. Ask it for the missing edge(s) first; only when it fails too does the
domain fallback + gallop run as before. Reported under `aggregate_edges` so
recovery.json still shows how the bound was obtained.

Live repro (real fixture: leftmost `sessions` leaf cell count overwritten,
400 rows): BEFORE bounds low=-9223372036854775808 copied=0 range_queries=10000
query_limit_reached=True status=failed; AFTER low=1 high=400 copied=391
range_queries=40 status=partial (only the damaged leaf's rows are lost).

Refs #98050
Refs #100313
Reported-by: Ace-Kelly
Corroborated-by: Proff506
2026-09-11 06:37:27 -07:00
Gunwoo Lee 1dcc0b06e2 fix(recovery): skip rows the destination schema rejects during partial salvage
`hermes sessions recover --allow-partial` walks damaged tables by rowid range
and falls back to an exact-rowid lookup for the last cell of a broken range.
A phantom row produced by page damage (a `sessions` row with a NULL
`started_at`, an `async_delegations` row with a NULL `state`) reads fine from
the source but violates the destination's NOT NULL constraint, and the
resulting sqlite3.IntegrityError escaped the exact-lookup boundary and aborted
the whole recovery — losing every healthy row behind it.

Catch IntegrityError at that boundary, count it under
`destination_rejected_rows` and record the rowid as a skipped singleton
("destination constraint rejected row: ...") so the run completes, verifies,
and the report shows exactly which rows were dropped.

Live repro (real fixture: source schema's NOT NULL relaxed via
writable_schema, phantom `sessions` row inserted): BEFORE
sqlite3.IntegrityError "NOT NULL constraint failed: sessions.started_at" at
session_recovery.py recover_exact_rowid; AFTER status=partial copied=3
destination_rejected_rows=1, verified=True. Reporter @i8ei confirmed the same
patch recovers the field database (22 sessions / 2,248 messages) in #102240.

Salvaged from PR #91413 (rebased onto the _RowidRangeSalvage refactor; test
reduced to one invariant).

Refs #102240
Reported-by: i8ei
2026-09-11 06:37:27 -07:00
liuhao1024 bd2c7e2457 fix(state): drop orphan FTS5 shadow tables per family on SessionDB open
`sqlite3 state.db .recover` re-emits the FTS5 shadow tables (messages_fts_data,
_idx, _docsize, _config, _content) as ordinary tables but cannot re-emit the
CREATE VIRTUAL TABLE row. The next SessionDB open ran the FTS DDL in
_ensure_fts_schema and died with "fts5: error creating shadow table
messages_fts_data: table 'messages_fts_data' already exists", so a recovered
database was unusable until someone hand-dropped the shadows.

_init_fts now runs _drop_orphan_fts_shadow_tables before any FTS DDL. It is
per-family and exact-name scoped: a family's shadows are dropped only when its
own vtable row is absent from sqlite_master (type='table' AND sql LIKE
'CREATE VIRTUAL TABLE%'), so a healthy messages_fts_trigram survives a
base-family repair untouched. A repaired base/trigram family is then treated
like a missing-trigger repair and rebuilt from the canonical messages table
under the cross-process rebuild admission; the shadows are derived index
state, nothing is lost.

Live repro: real `sqlite3 x.db .recover | sqlite3 y.db` on sqlite 3.50.4
keeps the vtable rows (the shell emits CREATE VIRTUAL TABLE), so the
deterministic fixture removes the vtable row via writable_schema leaving the
shadows behind: BEFORE OperationalError on open; AFTER opens, fts_enabled,
MATCH returns every message, trigram sqlite_master rowids unchanged.

Salvaged from PR #56824 (intent applied onto the current hermes_state_fts /
hermes_state_schema siblings). The ownership-safety point (never touch a live
family's shadows) was raised by @ggoldani in #103868 / #103840.

Refs #103840
Refs #56815
Co-authored-by: ggoldani <ggoldani@users.noreply.github.com>
2026-09-11 06:37:27 -07:00
Teknium a2413495ad fix(state): fts_trigram_session_sql qualifies model_config under the json_valid guard
The _sql_json_extract wrapper removed the `COALESCE(model_config` text the
alias string-replace keyed on, so the deferred-backfill SELECT joined an
unqualified `model_config`. Build the predicate from the alias directly and
derive the unaliased constant from it, so the two can never disagree.
2026-09-11 06:24:54 -07:00
teknium1 46afbfec10 fix(state): route the remaining model_config marker reads through _sql_json_extract
Sibling widening of the #101726 salvage: FTS_TRIGRAM_SESSION_SQL (trigram view/triggers/backfill),
the v16 delegate-tagging data migration and reopen_session's legacy reset-child stamp still called
json_extract() on the raw model_config cell, so one malformed JSON row could still abort FTS
maintenance, a schema migration or /resume of a reset child. Zero raw model_config json_extract
reads remain in hermes_state_*.py.
2026-09-11 06:24:54 -07:00
teknium1 c4a8572d68 test(state): pin corrupt-row robustness — one bad timestamp/oversized IN-list never kills the command
Three invariants over real SQLite fixtures (TEXT and 8.4e252 timestamps written straight into the
REAL columns; 1200 ids under a 999-variable ceiling via setlimit): list/export/insights complete
and name the corrupt session in a WARNING; writers never persist an out-of-window timestamp;
prune/delete_sessions succeed with zero orphaned messages. All three fail on origin/main.
2026-09-11 06:24:54 -07:00
Mi55ed 25670cd9e3 fix(state): batch large session cleanup queries below SQLite's variable limit
prune_sessions(), delete_empty_sessions() and prune_empty_ghost_sessions()
built one IN (?, ..., ?) clause containing every selected session id (the
parent-orphaning UPDATE), so cleaning more than SQLITE_MAX_VARIABLE_NUMBER
sessions failed atomically with "too many SQL variables". The per-row
DELETE loops that followed are folded into the same 900-id batches.

Same single _execute_write() transaction; only the binding is split.

Hand-ported from PR #100658 (targeted the pre-decomposition hermes_state.py
god file; the methods now live in hermes_state_maintenance.py /
hermes_state_sessions.py). The one-pass transcript-directory sweep from
that PR is not ported (out of scope for the variable-limit bug). Authored
by @Mi55ed; ported under --author.
2026-09-11 06:24:54 -07:00
Michael Steuer acfda37598 fix(sessions): chunk IN-list ids in bulk delete to avoid 'too many SQL variables'
`hermes sessions prune --source cron --older-than 14` on a store with ~60K
cron sessions (~50K matches) died with sqlite3.OperationalError: too many
SQL variables. SessionDB.delete_sessions, _collect_delegate_child_ids and
_delete_delegate_children each bound the full id list into a single
IN (?,?,...). SQLite caps bound parameters at SQLITE_MAX_VARIABLE_NUMBER
(999 on < 3.32, 32766 after), so any bulk delete above that failed outright.

Chunk every IN list (`_id_chunks` / `_SQL_IN_CHUNK` in hermes_state_common,
900 ids; the delegate walk binds each id twice so it chunks at half). Same
transaction, same cascade/orphan contract; only the parameter binding is
split.

Hand-ported from PR #102679 (targeted the pre-decomposition hermes_state.py
god file; the functions now live in hermes_state_sessions.py). Authored by
@mssteuer; ported under --author.
2026-09-11 06:24:54 -07:00
teknium1 496eb13bd7 fix(state): one corrupt timestamp row no longer kills sessions list, export or insights
SQLite dynamic typing lets a TEXT cell ('not-a-timestamp'), inf/nan or a
garbage double (8.4e252 salvaged from a damaged page) sit in a REAL
timestamp column. Every reader called datetime.fromtimestamp()/float
arithmetic on the raw cell, so ONE bad row raised TypeError/OverflowError
out of the row loop and took down the whole `hermes sessions list`/browse
table (#102399), all three exporters — JSONL/MD, QMD, HTML (#102352) —
and `hermes insights` (#99959).

Fix the class with ONE helper, hermes_cli.timefmt.coerce_epoch(): a
stored cell becomes float epoch seconds inside a sane 1970..2103 window
or None after a WARNING that names the session id. Every reader routes
through it — relative_time (list/browse/resume picker), format_epoch
(prune/candidates tables), the three exporters' timestamp formatters,
insights' _get_sessions/_day/period range — so a bad row renders as
'?'/'N/A'/raw text for that one cell and the command completes.

Write side: hermes_state_messages._coerce_timestamp (append_message,
append_messages_batch, import) and the import path's started_at now use
the same window, so a new out-of-range timestamp falls back to now()
instead of being persisted — new bad rows cannot be written by Hermes.

Reported-by: #102399, #102352, #99959 reporters; kokhlo's insights
analysis pointed at every reporting site, not just line 860.
2026-09-11 06:24:54 -07:00
liuhao1024 df1388e77b test(tools): extend the delegation fd-leak injection to the schema replay
The schema initializer now replays SCHEMA_SQL through executescript (the
single-authority path from #94701's follow-up), which bypasses the
execute()-level DDL failure injection — the regression stopped raising.
Bind the same simulated failure onto the executescript path so the
connect-close-on-init-failure contract stays pinned for both replay
mechanisms.
2026-09-11 06:24:54 -07:00
liuhao1024 a6d65cdd09 fix(state): single durable-shape authority for async_delegations
Review follow-up on #94701: the delegation tool's _initialize_schema
still carried its own CREATE TABLE + ALTER column list for
async_delegations, leaving a second durable-shape authority even with
the column declared in SCHEMA_SQL. Its legacy ALTER added
origin_session_id as bare TEXT (nullable, no default); reconciliation
repairs missing column names only, so a database first opened through
the tool kept a non-canonical shape forever (#94691).

Remove the private DDL entirely. The tool's initializer now calls a new
reconcile_state_schema() in hermes_state_schema, which replays the
canonical SCHEMA_SQL (idempotent CREATE IF NOT EXISTS for every table,
canonical indexes included) and reuses SessionDB's declarative
_reconcile_columns for missing-column backfill — one reconciliation
implementation, one authority. Because _parse_schema_columns
reconstructs each column's full constraint expression (type, NOT NULL,
DEFAULT), the tool-first legacy path now adds origin_session_id as
TEXT NOT NULL DEFAULT '' — the canonical shape — and SQLite backfills
existing rows with the '' default.

Opening-order regressions compare FULL PRAGMA table_info metadata
(type, notnull, dflt_value, pk) plus the canonical index set across
fresh SessionDB→tool, legacy→SessionDB, and legacy→tool→SessionDB,
each preserving a pre-existing legacy delegation row.
2026-09-11 06:24:54 -07:00
liuhao1024 0bb27d37cf fix(state): declare origin_session_id in the canonical async_delegations schema
The delegation tool carries its own CREATE TABLE for async_delegations
(tools/async_delegation.py _initialize_schema) plus a lazy ALTER TABLE
ADD COLUMN for the tables it finds already existing. Its column list
had drifted ahead of the canonical SCHEMA_SQL: origin_session_id
(raw api_server session id of the originating request, the wake
self-post target) existed only through the tool's lazy path, so two
databases at the same schema_version had different
async_delegations shapes depending solely on whether the delegation
tool had ever run. Rebuild/replay pipelines that reconstruct state.db
from the canonical schema then hit the column with no version gate to
explain it (#94691).

Declare the column in SCHEMA_SQL with the same TEXT NOT NULL DEFAULT ''
shape the tool uses. Fresh installs now carry it canonically; the
declarative _reconcile_columns backfills it into legacy databases on
the next writable open (same pattern as earlier additive columns); the
tool's lazy ALTER keeps serving pre-reconciliation databases. The two
schema authorities now agree, pinned by a test that runs the tool's
initializer over a canonical database and asserts the shape is
unchanged.

Fixes #94691
2026-09-11 06:24:54 -07:00
Efe Büken a239c4f811 fix(state): tolerate malformed session marker JSON 2026-09-11 06:24:54 -07:00
Xipong 78f85112d9 perf(state): batch message hydration in export_all 2026-09-11 06:24:54 -07:00
teknium1 16e4496d90 fix(agent): corrupt-session recovery guidance names the session's own state.db
The corruption explainer filled `{db_path}` from `_default_db_path()`, the
process default. A Desktop `serve` backend launched on the root home hosts
named-profile sessions whose SessionDB is `profiles/<name>/state.db`, so the
operator was told to inspect/repair a different profile's database. Pass the
agent's own `_session_db.db_path` from the turn finalizer; the process
default remains the fallback for agents without a bound store.

Reported in #105887.
2026-09-11 06:24:11 -07:00
teknium1 df0eed4f6b fix(session_search): a bare session id never reads another profile's state.db
Reading a session by id that missed the caller's store fell through to
_locate_session_db(), which opened every profile's state.db read-only and
returned the first owner's full transcript — no opt-in, no profile named, and
the miss path even fired after an explicit non-matching profile= read. Any
caller holding an id (ids appear in logs and tool output) could read a
foreign profile's conversation. Profiles are isolated islands by design.

A miss now stays a miss, with a hint to name the owning profile
(profile=<name> / @session:<profile>/<id>), which remains the sanctioned,
explicit cross-profile read. The schema eval runner no longer needs to fake
the scan.

Reported by the #106761 filer; reproduced by @kokhlo. Refs #87779.
2026-09-11 06:24:11 -07:00
teknium1 5a720bbb1b test(tui): launch-handle pin test opens a real registry SessionDB
The salvaged #102534 test patched hermes_state.get_shared_session_db, a seam
main dropped (server._get_db now calls hermes_state_registry.acquire), so the
fixture errored at setup and the file reported 0 passed. Assert on the real
handle's db_path and on the foreign home staying untouched instead of on a
fake factory; red with the pin reverted, green with it.
2026-09-11 06:24:11 -07:00
HexLab98 342705135a test(tui): cover launch SessionDB binding under profile override windows
Add a #102526 regression and adjust resume ownership leak filtering now that
the shared launch handle carries an explicit state.db path.
2026-09-11 06:24:11 -07:00
HexLab98 3c3ca061a6 fix(desktop): pin launch SessionDB handle to the launch home
The lazy _get_db() singleton followed get_hermes_home(), so a first touch
inside the multiplex cron ticker's per-profile override window permanently
bound the default backend to another profile's state.db (#102526).
2026-09-11 06:24:11 -07:00
Sora-bluesky 3a350508f7 fix(cli): wait for SessionDB teardown when deleting a profile
close_all_under returned after the last release dropped the generation
and before the physical close finished, so rmtree still saw the open
handle. Wait directory-matching teardown barriers the same way close_all
does.
2026-09-11 06:24:11 -07:00
Sora-bluesky f1ee8e0f39 fix(cli): release SessionDB handles when deleting a profile
delete_profile already force-closes holographic memory_store.db in this
process, but the shared SessionDB registry kept state.db open. Recreate
then failed with a replaced/locked database. Close every shared handle
under the doomed directory, same contract as MemoryStore.release_all_under.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-11 06:24:11 -07:00
teknium1 d956e05e5e chore: contributor mappings for the statedb-wal salvage (gaoanze888 work email, nikkoxgonzales noreply) 2026-09-11 06:23:23 -07:00
teknium1 e00a21cd39 fix(state): retry a transient SQLITE_IOERR on the pooled read path before surfacing it
Since 0.21.0 reads go through mode=ro pooled connections. A read-only OPEN already
rides out the millisecond WAL transition window (checkpoint / WAL reset / frame flush
by a sibling process; the ro reader cannot rewrite the -shm index) with a bounded retry
(#100436), but a WARM pooled reader hitting the same window while its SELECT executes
propagated `disk I/O error` straight out of get_session(): 37 identical tracebacks on a
multi-process WSL2 ext4-on-vhdx install, each followed by "compression session recovery
failed", with quick_check=ok (#100871). The reporter's A/B shows the operator
workaround (journal_mode=delete) collapses read throughput ~30000x, so the flake has to
be absorbed on the read path.

_read_one/_read_all now replay the idempotent statement within the existing read-only
IOERR budget (3 x 50 ms) on the SAME connection -- close+reopen would cancel this
process's POSIX locks for every sibling connection -- and a persistent IOERR still
propagates. No quarantine: EIO on a read is busy, not broken. Every SELECT in the
SessionDB siblings (63 call sites) reaches the pool through these two helpers, so the
class is covered without a wrapper type.

Same-connection retry per #100882's analysis (@fangliquanflq); #100883
(@Sahilvishnaliya) diagnosed the missing recovery in the 0.21.0 read pool.

Fixes #100871.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Sahilvishnaliya <222165401+Sahilvishnaliya@users.noreply.github.com>
2026-09-11 06:23:23 -07:00
teknium1 295edf2557 fix(state): gate the lock-free read pool on a confirmed WAL header, not the assumed mode
apply_wal_with_fallback() reports "wal" in two indeterminate cases -- the vulnerable-
SQLite gate (_apply_delete_for_wal_reset_bug) and the non-vulnerable probe-unknown path
(a7f2a593d1) -- meaning "touched nothing, the connection inherits the header's mode".
SessionDB turned that assumption into `_wal_active=True`, which enables the mode=ro read
pool that skips `self._lock`. On a file that is really in rollback-journal mode those
readers race the writer with a 5s busy timeout and no retry: random SQLITE_BUSY read
failures for the instance's lifetime (#86515).

Confirm the header on the freshly opened connection before enabling the pool. When the
probe is still blocked, reads queue on the writer connection under the lock -- slower,
never wrong. Every other apply_wal_with_fallback caller ignores the return value, so the
consumer is the right place to gate; changing the return contract to Optional across
15 call sites (#87044's shape) is not needed.

Live repro: DELETE-mode file, sibling holding BEGIN EXCLUSIVE during open ->
before: _wal_active=True and _checkout_read_conn() hands out a pooled mode=ro conn;
after: _wal_active=False, reads take the locked writer path.

Fixes #86515. Based on the analysis in #87044.
Co-authored-by: QDung210 <dqdung205@gmail.com>
2026-09-11 06:23:23 -07:00
teknium1 9f9ec647f5 fix(gateway): re-check the live store before broadcasting the state.db warning
_send_session_db_warning_notifications() broadcast the error recorded at startup
without asking whether it was still true. A startup `database is locked` routinely
clears while the adapters are still connecting (another profile's open, a `hermes
sessions` one-shot, a slow SMB lock release), so the home channels were told the
store was unavailable when it had already healed — and a warning that is wrong once
is ignored the next time it is right.

The RecoverableHandleCache opener already clears `_session_db_init_error` on
recovery; the broadcast now drives one open attempt through it first and only warns
when the error is still standing. A store that is genuinely still down warns exactly
as before.

Fixes #108031.
2026-09-11 06:23:23 -07:00
teknium1 045704ce00 test(state): port the #107411 dentry tests to the shared WAL-generation harness
main replaced the file-local _make_db/_require_wal/_unlink_sidecars helpers with
tests/hermes_state/_wal_generation_harness (make_db/require_wal/lose_sidecars) after
PR #107411 branched; the cherry-picked tests referenced the old names and failed at
collection with NameError. Fixture-only change; the assertions are unchanged.
2026-09-11 06:23:23 -07:00
nikkoxgonzales 4d9ec9be50 fix(db): survive DELETE fallback failure in WAL setup 2026-09-11 06:23:23 -07:00
gaoanze 031d3760d0 fix(state): distinguish retired WAL recovery guidance
Classify deleted WAL generations separately from main-file replacement and point operators at the captured-generation manifest and mode-aware recovery path.

Co-authored-by: crazyief <8566250+crazyief@users.noreply.github.com>
2026-09-11 06:23:23 -07:00
ca-shrimp 3a3ec6eeaf fix(state): serialize replaced/generation probe with close() in _execute_write
A lock-free _raise_if_db_replaced() probe at the top of the _execute_write
retry loop raced a concurrent close(). close() runs under the same _lock and
ends the WAL generation: it checkpoints, closes the connection (SQLite
unlinks the -wal/-shm sidecars), nulls _conn and clears
_db_sidecar_identity. The probe could observe the mid-teardown state —
sidecars already unlinked while _db_sidecar_identity was not yet cleared —
and misclassify this process's OWN clean close as an externally deleted WAL
generation, raising a sticky DeletedWalGenerationError that permanently
refused every later write on that handle (#105567).

Move the live probe inside the lock, ahead of the close-race reopen
decision, so it only ever observes the stable post-close state (identity
cleared -> the existing adopt/reopen path). The corrupt flag check stays on
the lock-free fast path; external file/generation replacement detection is
unchanged, just serialized with teardown.

Synthetic repro (100 rounds x 40 writes, direct SessionDB handles): before
~9 failing rounds / ~360 DeletedWalGenerationError; after 0 failures,
4000/4000 writes persisted across repeated runs. tests/state (181) plus the
generation/replaced/corrupt guard suites (55) pass.

Fixes #105567
2026-09-11 06:23:23 -07:00
chelsealong d6b036b726 fix(state): compare fd identity against the watched path, not st_nlink
st_nlink == 0 alone cannot distinguish a genuine orphan from one that
still has a surviving hard link (e.g. a backup) after the watched
sidecar path itself was removed or replaced — that left st_nlink >= 1
on a truly orphaned generation, letting a new opener through while a
live writer still owned the old one. Compare (st_dev, st_ino) between
the fd and the current watched sidecar path instead: only an exact
match means they're the same live file, so any mismatch or unstattable
watched path still fails closed.
2026-09-11 06:23:23 -07:00
chelsealong 84a3c4de74 fix(state): require nlink==0 before treating a /proc fd as an unlinked WAL sidecar
iter_deleted_sqlite_sidecar_holders() and SessionDB._wal_generation_was_lost()
both treated a `` (deleted)`` suffix on a /proc/<pid>/fd/* target as proof that
state.db-wal or state.db-shm was unlinked. On OpenZFS that suffix is not proof:
a live, still-linked file whose dentry was unhashed is reported the same way,
with st_nlink still 1 and the same (dev, ino) as the path. The guard then fires
permanently and the gateway falls back to JSONL forever, because the WAL was
never actually deleted.

Add _fd_is_truly_unlinked(), which confirms via os.stat(fd_path).st_nlink == 0
before a target counts as an orphaned generation. An unstattable descriptor
still counts as deleted, so the guard keeps failing closed. _iter_proc_fd_targets()
and _proc_fd_targets() now also yield the /proc fd path itself so both call
sites (open-path and the sticky write-path probe) can run the check.
2026-09-11 06:23:23 -07:00
Teknium 7dec81568e test(desktop): trim salvaged pool tests to invariants
Drop two PrimaryProfilePin cases that only restate the constructor
defaults and blank-string normalisation, and the wiring-routing test that
froze POOL_LIMITS_SETTINGS_ROUTE to a literal string — a snapshot of the
constant, not a behaviour contract. The two kept pin tests cover the bug
(a live primary keeps answering for its booted profile after the stored
preference moves; teardown releases the pin), and the notifications tests
cover the toast action end-to-end.
2026-09-11 06:23:18 -07:00
Mabolla 19cff347ad fix(desktop): wait for an evicted backend to release its slot
Await LRU teardown in each pooled backend creation path so replacement wakes do not race an exiting child for the hard spawn slot.
2026-09-11 06:23:18 -07:00
Mabolla 919dea9020 fix(desktop): await LRU teardown before queued wake
Wait for selected stale backend processes to exit before a replacement profile wake enters the bounded spawn queue.
2026-09-11 06:23:18 -07:00
joaomarcos a863bbcc28 fix(desktop): make pool slot timeouts actionable 2026-09-11 06:23:18 -07:00
Alexandru Ionescu d27180ba7f fix(desktop): pin primary backend routing to its booted profile
`primaryProfileKey()` re-read active-profile.json on every call. The rail's
live workspace switch rewrites that file via `hermes:profile:remember`
WITHOUT re-homing the primary, so after a switch the routing table disagreed
with the running process: a request for the profile the primary actually
booted as (e.g. "default") no longer matched `primaryProfile` in
`resolveProfileBackendRoute`, fell through to the pool, and spawned a second
backend for the same HERMES_HOME.

The duplicate was keepalive-fresh so LRU eviction spared it, it burned a pool
slot, and with the default cap of 3 every further profile queued and failed
with `Local backend start for "<profile>" timed out while waiting for a free
slot` (repro in desktop.log: "default" spawned as a pool backend while the
primary "default" was still running; coder/qwen then timed out for 20+ min).

Snapshot the launch profile in `startHermes()` (PrimaryProfilePin.pin) and
release it in `resetHermesConnection()` so the next start follows the stored
preference again. The pin is a tiny pure module with tests; main.ts only owns
the file read and the two call sites.
2026-09-11 06:23:18 -07:00
Teknium d1bd7b00df fix(desktop): dispatch probe falls back to /api/status on pre-/api/health remotes
Switching the pooled dispatch probe to /api/health (salvaged from #97914)
would 404 on every dispatch against a remote older than 0.19, retire the
tunnel and reconnect forever - the same storm #107997 describes, moved to
old backends. Fall back to /api/status on an explicit 404 exactly the way
the boot readiness probe already does (backend-health.ts). The legacy
fallback idea and its test are taken from #101976 (@edosulai); the rest of
that PR (timeout-tolerance streak, ServerAlive SSH options) is not adopted.

Co-authored-by: Edo Sulaiman <edosulai@icloud.com>
2026-09-11 06:22:35 -07:00
bennybuoy ee8257d601 fix(desktop): probe /api/health on pooled SSH dispatch, not /api/status
Cold /api/status through a Windows no-mux SSH forward routinely exceeds
the 2.5s dispatch budget, so Desktop retires a live tunnel and respawns.
Use the cheap /api/health route (5s, same as DEFAULT_HEALTH_PROBE_TIMEOUT_MS).
Background liveness still probes /api/status at 10s.
2026-09-11 06:22:35 -07:00