On origin/main only the TUI copied durable `_row_id`s onto the installed warm
prefix (its clients address follow-ups by row); folding the loop into
`rewind_user_turn` made CLI /undo grow `_row_id` on a resumed history that never
had it. Live CLI rows already carry ids from the flush, so for them it was a
no-op, but a resumed transcript changed shape. The loop now runs only with
`adopt_row_ids=True`, which the TUI passes; the CLI history shape is unchanged.
Review follow-up on #109610.
The 3-surface parity test only rewound plain text rows. Add one parametrized
case where every turn is a multi-part (text + image) user ask answered through
an assistant tool_call, its tool result and a final reply, asserting all three
surfaces archive the whole exchange identically (no orphaned tool row) and the
surviving image rows stay byte-identical to what was stored.
Review follow-up on #109610.
CLI `_rewind_persisted_user_turn`, TUI `_rewind_active_session_history` and gateway
`rewind_session` each re-ran get_active_message_ids -> get_messages_as_conversation ->
split_user_originated_turn -> rewind_to_message with their own warm/durable comparison
helpers and three different out-of-range contracts (RuntimeError / ValueError / None).
The durable transcript is the authority for a rewind, so the implementation now lives
with the data: `SessionDB.rewind_user_turn` (hermes_state_rewind.py) with one typed
out-of-range error (`RewindTargetUnavailableError`). Surfaces keep only lock, eviction
and rendering glue and map that error to their own message.
optimize_fts() and vacuum() already refuse to run against a quarantined
handle (_db_corrupt / _db_replaced / _db_wal_generation_lost): both would
rewrite index/file pages in place, turning contained, diagnosable
corruption into an amplified one. rebuild_fts() never got the same guard,
despite being the more destructive of the two ("discards and recreates the
index data entirely", per its own docstring, vs. optimize_fts's segment
merge).
It's also independently reachable outside _execute_write's own quarantine
check: gateway/session_transcript.py's _rebuild_fts_once() calls
db.rebuild_fts() directly from the FTS-corruption transcript-retry path,
with no quarantine check of its own (only a WAL split-brain / foreign-holder
check, a different concern). A quarantined handle hitting that retry path
would run a full FTS rebuild — and commit it — on a corrupt, replaced, or
split-WAL-generation file.
Add the same self._raise_if_db_corrupt()/self._raise_if_db_replaced() pair
optimize_fts() already has, at the top of rebuild_fts(), before it enters
the cross-process rebuild admission.
Three invariants over real SQLite fixtures (TEXT and 8.4e252 timestamps written straight into the
REAL columns; 1200 ids under a 999-variable ceiling via setlimit): list/export/insights complete
and name the corrupt session in a WARNING; writers never persist an out-of-window timestamp;
prune/delete_sessions succeed with zero orphaned messages. All three fail on origin/main.
close_all_under returned after the last release dropped the generation
and before the physical close finished, so rmtree still saw the open
handle. Wait directory-matching teardown barriers the same way close_all
does.
delete_profile already force-closes holographic memory_store.db in this
process, but the shared SessionDB registry kept state.db open. Recreate
then failed with a replaced/locked database. Close every shared handle
under the doomed directory, same contract as MemoryStore.release_all_under.
Co-authored-by: Cursor <cursoragent@cursor.com>
Since 0.21.0 reads go through mode=ro pooled connections. A read-only OPEN already
rides out the millisecond WAL transition window (checkpoint / WAL reset / frame flush
by a sibling process; the ro reader cannot rewrite the -shm index) with a bounded retry
(#100436), but a WARM pooled reader hitting the same window while its SELECT executes
propagated `disk I/O error` straight out of get_session(): 37 identical tracebacks on a
multi-process WSL2 ext4-on-vhdx install, each followed by "compression session recovery
failed", with quick_check=ok (#100871). The reporter's A/B shows the operator
workaround (journal_mode=delete) collapses read throughput ~30000x, so the flake has to
be absorbed on the read path.
_read_one/_read_all now replay the idempotent statement within the existing read-only
IOERR budget (3 x 50 ms) on the SAME connection -- close+reopen would cancel this
process's POSIX locks for every sibling connection -- and a persistent IOERR still
propagates. No quarantine: EIO on a read is busy, not broken. Every SELECT in the
SessionDB siblings (63 call sites) reaches the pool through these two helpers, so the
class is covered without a wrapper type.
Same-connection retry per #100882's analysis (@fangliquanflq); #100883
(@Sahilvishnaliya) diagnosed the missing recovery in the 0.21.0 read pool.
Fixes#100871.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Sahilvishnaliya <222165401+Sahilvishnaliya@users.noreply.github.com>
apply_wal_with_fallback() reports "wal" in two indeterminate cases -- the vulnerable-
SQLite gate (_apply_delete_for_wal_reset_bug) and the non-vulnerable probe-unknown path
(a7f2a593d1) -- meaning "touched nothing, the connection inherits the header's mode".
SessionDB turned that assumption into `_wal_active=True`, which enables the mode=ro read
pool that skips `self._lock`. On a file that is really in rollback-journal mode those
readers race the writer with a 5s busy timeout and no retry: random SQLITE_BUSY read
failures for the instance's lifetime (#86515).
Confirm the header on the freshly opened connection before enabling the pool. When the
probe is still blocked, reads queue on the writer connection under the lock -- slower,
never wrong. Every other apply_wal_with_fallback caller ignores the return value, so the
consumer is the right place to gate; changing the return contract to Optional across
15 call sites (#87044's shape) is not needed.
Live repro: DELETE-mode file, sibling holding BEGIN EXCLUSIVE during open ->
before: _wal_active=True and _checkout_read_conn() hands out a pooled mode=ro conn;
after: _wal_active=False, reads take the locked writer path.
Fixes#86515. Based on the analysis in #87044.
Co-authored-by: QDung210 <dqdung205@gmail.com>
main replaced the file-local _make_db/_require_wal/_unlink_sidecars helpers with
tests/hermes_state/_wal_generation_harness (make_db/require_wal/lose_sidecars) after
PR #107411 branched; the cherry-picked tests referenced the old names and failed at
collection with NameError. Fixture-only change; the assertions are unchanged.
Classify deleted WAL generations separately from main-file replacement and point operators at the captured-generation manifest and mode-aware recovery path.
Co-authored-by: crazyief <8566250+crazyief@users.noreply.github.com>
st_nlink == 0 alone cannot distinguish a genuine orphan from one that
still has a surviving hard link (e.g. a backup) after the watched
sidecar path itself was removed or replaced — that left st_nlink >= 1
on a truly orphaned generation, letting a new opener through while a
live writer still owned the old one. Compare (st_dev, st_ino) between
the fd and the current watched sidecar path instead: only an exact
match means they're the same live file, so any mismatch or unstattable
watched path still fails closed.
iter_deleted_sqlite_sidecar_holders() and SessionDB._wal_generation_was_lost()
both treated a `` (deleted)`` suffix on a /proc/<pid>/fd/* target as proof that
state.db-wal or state.db-shm was unlinked. On OpenZFS that suffix is not proof:
a live, still-linked file whose dentry was unhashed is reported the same way,
with st_nlink still 1 and the same (dev, ino) as the path. The guard then fires
permanently and the gateway falls back to JSONL forever, because the WAL was
never actually deleted.
Add _fd_is_truly_unlinked(), which confirms via os.stat(fd_path).st_nlink == 0
before a target counts as an orphaned generation. An unstattable descriptor
still counts as deleted, so the guard keeps failing closed. _iter_proc_fd_targets()
and _proc_fd_targets() now also yield the /proc fd path itself so both call
sites (open-path and the sticky write-path probe) can run the check.
The two test files each carried their own force_wal/_make_db/_require_wal, a byte-identical
_write_second_generation, and a ~60-line Popen/reader-thread/next_event driver around a
near-identical gateway child script. They now share tests/hermes_state/_wal_generation_harness.py
(gateway_writer() context manager, one child script). The mock-only 'no wal_checkpoint SQL string'
test is dropped: its own comment calls it necessary but not sufficient, and the real-file
close/exit integrity tests cover the same contract.
Every production close of a SessionDB goes through hermes_state_registry.release_or_close,
whose teardown swallows any exception from close() at DEBUG. With the capture in place that
meant a RetiredGenerationCaptureError (disk full, permissions) left the handle open silently
and, on Python 3.11, skipped the retention pin: the raise happened before the pin, so the
interpreter's exit still ran sqlite3_close's implicit checkpoint and wrote the stale frames
over the newer generation (probe: 302 rows -> 4 after exit, worse than main).
- close() now decides retention first and takes the pin BEFORE attempting the capture on
runtimes without setconfig; the capture failure is logged at ERROR and re-raised, so a
later close() retries it while the pin already protects the newer generation.
- The registry logs RetiredGenerationCaptureError at ERROR instead of DEBUG (other teardown
errors stay quiet). Same treatment on the release_or_close fallback path.
- The lost-generation settlement moves out of close() into _settle_lost_generation_locked();
the sticky _close_checkpoint_disabled attribute and the dead "setconfig exists but failed"
late-bind block are gone (_disable_close_time_checkpoint returns the call's outcome).
Regression test drives release_or_close with a failing capture: no raise, error text in the
log, pin taken exactly once (3.11), retry closes cleanly. Red on the salvaged head on both
3.11 and 3.12, green here.
With the retired generation captured durably, closing a lost-generation
handle is safe wherever SQLITE_DBCONFIG_NO_CKPT_ON_CLOSE took effect: the
frames are preserved and sqlite3_close no longer checkpoints them into the
newer generation. On Python 3.11, where sqlite3 has no setconfig, closing
still runs SQLite's internal checkpoint over the newer main file, so only
there the exact quarantined connection is retained instead of closed: one
public Py_IncRef reference via ctypes.pythonapi (no struct-layout access),
bound before the writer opens, taken in close() after the capture. Runtimes
with setconfig never touch ctypes; a writable SessionDB requires CPython
with ctypes only where retention is the sole guard.
Read-only handles and every other quarantine reason close as before.
Regressions cover both branches: where retention applies, the retired inode
stays readable after close() and GC and an independent deleted WAL under
the same pathname is untouched; elsewhere the capture is the surviving copy.
A separate process's newer generation survives close, GC and normal exit
in both cases. The mock-based close-time regression originally written for
Refs #105670
Co-authored-by: fangliquanflq <fangliquan@qq.com>
After DeletedWalGenerationError the retired frames exist only in an
unlinked -wal inode that this process keeps open. #106315 stops them from
being checkpointed under wrong page numbers, but the canonical remediation
("stop the gateway, dashboard and cron writers, then reopen") lets the
kernel drop the inode with them: committed transactions whose disposition
is unknown were destroyed by the prescribed recovery path itself.
Capture the exact retired generation next to the database at the first
halt, or at close() if that sees the loss first. The WAL is located by the
(st_dev, st_ino) recorded at open among this process's own descriptors,
never by pathname, so a sibling's deleted WAL or a newer sidecar minted at
the same path cannot be mistaken for it; it is read with pread and no
descriptor is closed, moved or truncated. The artifact holds the WAL, the
-shm when still ours, the main image (or its header past a size cap) and a
manifest with identities, digests and the sidecar generation found at the
path at capture time. Nothing is merged: whether the frames belong on top
of the file now at the path stays an operator decision. close() refuses to
settle without the capture: it raises and leaves the handle open.
Regressions: capture at halt and at close, refusal to settle on capture
failure, inode-not-pathname selection with a second deleted WAL under the
same name, refusal to guess without a recorded identity, and a subprocess
control where rows committed only in the retired WAL are recovered from
the capture alone after the writer process has exited.
Refs #105670
The publish-side taxonomy (_end_stamp_class, _compression_parent_obstacle,
compression_parent_deliberately_ended) and the agent-guard delegation produced
exactly main's verdict -- every non-automatic stamp still fails closed -- so
they only changed an error string. Inline the "explicit close with no
continuation" test into reopen_if_explicitly_closed(), restore main's publish
branch and agent guard untouched, and keep two tests: the field shape rotates
after the host clears the stamp; boundary/compression/automatic stamps and a
session already claimed for teardown are never cleared.
Consolidated from PR #106543 (5 commits, final tree d2c4d908) by @Totoro-qaq.
publish_compression_child() fails closed on any non-automatic end stamp and
end_session() is first-stamp-wins, so a stale tui_close on a session the TUI
still routes turned every rotation into "compute the summary, then discard it"
(#106459). The host that still routes the session clears the stale explicit
close via SessionDB.reopen_if_explicitly_closed() before the turn starts;
publication never heals explicit closes. Review probes by @ehz0ah.
Six new tests collapsed into three that each go red without the indexed
projection: bounded VM steps for a page read and for an append (the
symptom), composite user-handoff identity + first-row position + internal
columns never leaking + signed-zero timestamp identity (the parity
contract), and a legacy store read-only then lazily migrated (the
persistence contract). The parametrized symmetry/visibility-toggle test
exercised the same key through raw SQL edits no production path performs
and passed on base, so it was a change-detector for the trigger shape.
A quarantined/replaced/split-generation handle must never run a full-file rewrite or an FTS5
'optimize': both read damaged or foreign pages and commit the result back, turning contained,
diagnosable corruption into an amplified one. Same guard _execute_write applies to every write.
Salvaged from #102092 onto current main: the _try_wal_checkpoint half landed via #106315's
_quarantine_reason(), so only the two rewrite sites remain.
Follow-up to #106315. vacuum() ran PRAGMA wal_checkpoint + VACUUM + wal_checkpoint(TRUNCATE) on
self._conn with no quarantine check; the only guard it inherited (optimize_fts raising
DeletedWalGenerationError) was swallowed by its own try/except and the rewrite proceeded on the
split-brain handle. Mutation on main: vacuum() returned 2 and rewrote pages after the write stop.
retry_deferred_fts_recovery gated only on _db_corrupt ("mirrors _try_wal_checkpoint /
close") — after this PR it no longer mirrored them: on a replaced/lost-generation handle
the periodic housekeeping tick still ran FTS DDL/DML + commit, the same split-brain write
class as the #105670 checkpoint. One SessionDB._quarantine_reason() now decides for the
periodic checkpoint, close(), and the FTS retry, with the halt path's precedence
(replaced before generation loss) and the operator wording in one place.
Test: the periodic-checkpoint case folds into the close test (same setup), which now
also proves the FTS retry returns False without touching the file; the mutation with
main's schema sibling swapped in returns True (a rebuild ran).
AGENTS.md: a bare skipif(sys.platform != linux) is never listed by
scripts/ci/list_os_marked_tests.py, so the tests would run nowhere on the
OS lanes. The marker is the contract.
- close() and _try_wal_checkpoint() now skip when _db_replaced or _db_wal_generation_lost
(previously only _db_corrupt was checked) — prevents checkpointing stale-generation frames
into the main DB, which is the shutdown-time damage reported in #105670
- _halt_if_db_generation_changed() calls _disable_close_time_checkpoint() alongside the flag
set (3.12+: disables SQLite internal last-connection checkpoint too)
- Regression tests: halted handle must not run explicit PRAGMA checkpoint on close(),
halt must call setconfig(NO_CKPT_ON_CLOSE), periodic _try_wal_checkpoint() must skip
- acquire(): the discard-and-recurse path becomes one more iteration of the existing wait loop;
the two inline 'with lifecycle_lock: _teardown(db)' copies reuse _teardown_generation, and the
type-narrowing asserts go away with the recursion. release() reads generation.path directly.
- Restore the real os.replace inode swaps in test_state_db_file_identity.py and the registry tests:
those files carry no windows_only marker so they never run on Windows, and the monkeypatched
predicate stopped exercising the stat->identity->halt path anywhere.
- Drop the auto-archive change and its 4 tests: on main the sweep gets a bare SessionDB and
db.close() already releases a registry-shared handle, so the described NameError leak only
existed on this branch's earlier head. trace_upload: acquire(None) already defaults.
- Trim the barrier tests to the invariant pair (retired drain must not lift a pending current
teardown; replacement not published before the last close settles) plus the raising-close
settlement; comments say the WHY once.
The #102827 corruption is pure zero holes -- frames lost across a WAL
generation. SessionDB.close() produces exactly that when it runs against a
file another live handle is still writing: PRAGMA wal_checkpoint(PASSIVE),
then the connection close that lets SQLite unlink -wal/-shm. The dangerous
event is a physical close overlapping any other live physical lifetime for
the same path, so both sides of it are closed here.
Late write vs. close: a cron watchdog timeout only stops waiting, and
ThreadPoolExecutor.shutdown(wait=False) cannot interrupt a worker already
inside run_conversation. The agent and its registry reference are now held
until that worker's Future completes, so its last frames land before any
checkpoint.
Close vs. open: the per-path barrier now COUNTS admitted teardowns. A path
can own several closes at once -- the current generation's final release and
a retired generation's drain are admitted independently under the registry
lock, and the per-path mutex only serializes teardowns that already entered
it. With one bare event per path, a releasing thread descheduled between
generation removal and the mutex let the next teardown to settle remove and
signal the shared event: close_all() returned over a pending close and
acquire() published a replacement writer on top of a handle still inside
checkpoint/unlink. _TeardownBarrier tracks event + pending count,
_admit_teardown_locked registers each close in the same lock section that
removes the generation, and only the last settled teardown lifts the
barrier. Physical I/O stays outside the registry lock and unrelated paths
still progress independently.
The auto-archive sweep called release_or_close in its finally while the
import was local to a different function, so every eligible sweep raised
NameError, the outer except Exception swallowed it at debug level, and the
borrowed registry reference was never returned -- a holder leak that pins a
retired generation open. The helper is now bound in the calling scope.
Remaining in-process writable SessionDB() call sites (trace upload, the
API-server profile cache, the web-server writable paths, startup schema
reconcile) go through the canonical registry acquire/release_or_close, and
gateway maintenance borrows pinned handles instead of iterating an unpinned
snapshot.
Regressions: overlapping final releases of the current and retired
generations in both orderings with the first paused before the lifecycle
mutex, teardown-error settlement, an unrelated-path control, and refcount
assertions for the auto-archive sweep on success, on failure, across
repeated sweeps and with auto-archive disabled.
Fixes#102827
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzxCWw6SuHXhXMdkiEwMa2
_dedupe_display_generations chose the right representative row per logical
message but sorted the survivors by that representative's id. A protected-tail
copy written into a newer compaction generation has a higher id than messages
emitted after the original, so include_compacted reads came back as C, A, B.
Anchor the sort on the logical message's first-ever row id instead.
Salvaged from #93869 (the tui_gateway half of that PR is superseded by #100504
and #104137); the code moved from hermes_state.py to hermes_state_messages.py
since, so the change is re-applied to its new home with the PR's regression
test verbatim.
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
Preselect indexed recent candidates before rich hydration, cap and deduplicate compression lineage traversal, and interrupt sustained SQLite work through a cooperative progress deadline. Fail closed when the bounded browse API is unavailable and cover legacy-schema reconciliation plus malformed lineage cases.
Assert the invariant — every display read of one session returns the same
transcript — plus the two things that must not grow with it: the model-fed
projection stays compressed, and Undo/Rewind rows stay hidden. Six of these
fail on the unfixed read paths.
Follow-ups on the salvaged #101081 guard:
- A clean close() lets SQLite unlink the WAL sidecars legitimately; the
guard treated that as a lost generation and permanently halted the
handle, so the #94736 late-write self-heal reopen dropped transcript
tails (4 existing tests failed). close() now clears the recorded
sidecar generation, and _wal_generation_was_lost() re-adopts the
current sidecars after a clean /proc/self probe instead of relying on
a stale snapshot.
- Healthy writes no longer walk /proc/self/fd: once a sidecar
generation is recorded, the stat-based inode check alone detects an
unlink/replace. The fd probe only runs in the empty-identity state
(fresh DB, post-close reopen).
- DeletedWalGenerationError now subclasses StateDbReplacedError, so the
gateway retry queue and run_agent flush divert transcripts to the
JSONL fallback exactly as they do for a replaced store, instead of
retrying forever against a halted handle.
- __init__ refuses once (under the startup lock) instead of twice per
open, halving the system-wide /proc scan; dropped the dead
include_self parameter and the dead _IS_WINDOWS clause.
- Test fixes: rstrip(' (deleted)') char-set bug -> removesuffix; the
non-linux test now patches sys.platform (the real gate) instead of
_IS_WINDOWS.
A live writer can keep a deleted state.db-wal inode while a second opener
mints a fresh WAL at the same path. Fail closed on writable open (before
connect) and on the write-path sidecar identity check so the second
generation is never created.
Co-authored-by: Noa <rainbowgore@users.noreply.github.com>
Follow-up to the salvaged #99560 commit: restore the original docstring
wording (user writes still always land everywhere else), name the
llm-outranks-derived hole the guard now closes, and annotate the two new
tests with what each pins.
Skipping the explicit PRAGMA wal_checkpoint(PASSIVE) in close() left
sqlite3.Connection.close() running SQLite's own last-connection PASSIVE
checkpoint, which still checkpoints the WAL and unlinks -wal/-shm on a
structurally corrupt file (E2E: the -wal vanished on close despite the
quarantine). Python 3.12+ exposes SQLITE_DBCONFIG_NO_CKPT_ON_CLOSE via
Connection.setconfig(); arm it in _halt_db_corrupt so the WAL image
survives close() for forensics/recovery. On 3.11 the switch does not
exist; the docstring and docs now say so instead of claiming sqlite3
cannot reach it at all.
Follow-up to #101095; flagged by JoaoMarcos44 on #101093.
A bare SQLITE_CORRUPT/NOTADB on a live write (not FTS-scoped, not a
replaced file) now sets a sticky per-instance flag: later writes fail
fast with StateDbCorruptError, the handle never reopens after close(),
and close() skips its explicit PASSIVE WAL checkpoint. Gateway and agent
flush paths divert pending transcripts to JSONL/spool like the replaced
case instead of retrying forever.
Field evidence: a handle that kept writing for ~50 minutes after the
first structural error checkpointed 15 pages under the wrong page
numbers on shutdown (page 1 <- messages_fts_trigram_data leaf), turning
"malformed" into "file is not a database".
Refs #90837, #90950, #97940, #89332, #45383
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CNX8rNYHqA5pT4tAGSzXtb
The _clean_registry fixture clears _generations and _retired between
tests; the new _opening map needs the same reset so a test that aborts
mid-construction cannot leave a stale opening event that stalls the
next test's cold acquire.
Follow-up on top of the salvaged #94095 commit:
- maybe_auto_prune_and_vacuum() now returns 'closed' (stale open state-owned
sessions marked ended) alongside 'pruned', so entrypoints can report the
reconciliation without parsing logs.
- Docstring explains the two-window lifecycle (close now, delete after a
further retention window).
- Regression test: cron/kanban/subagent rows with ended_at NULL are closed on
pass 1 and deleted on pass 2; a telegram row is never touched.
- website/docs sessions.md documents the automatic stale-open sweep.