The gateway's three display reads — session.resume, the ancestor lineage
prefix, and the warm-session payload — filtered active = 1, so a compacted
conversation rendered as its summary plus the carried-forward tail. The REST
transcript read has included the archived rows since #80680, so the same
session read two ways gave two different answers.
Extract the generation-dedupe policy the REST read carried inline into a
shared _dedupe_display_generations() and point every display projection at
it. The model-fed history stays active-only: compaction still compresses the
working context, it just no longer erases the user's transcript.
Flip the state.db retention defaults per Teknium's decision on #54189:
- sessions.auto_prune: false -> true. A stock install now prunes ENDED
sessions inactive for retention_days at CLI/gateway/cron startup
(at most once per min_interval_hours). Open, pinned and mid-turn
sessions are never deleted; the only open rows touched are stale
automation sessions (#100903 sweep), which are closed, not deleted,
and aged a further full window before removal.
- sessions.retention_days stays 90 (already the default; verified).
- Auto-VACUUM is now additionally gated on the reclaimable fraction of
the file: PRAGMA freelist_count / page_count must exceed 25%
(AUTO_VACUUM_MIN_FREELIST_RATIO) on top of the existing
min_vacuum_interval_days throttle. Pruning a few small sessions on a
dense multi-GB DB no longer rewrites the whole file to reclaim a few MB.
Unknown ratio (pragma read failure) falls back to the time throttle.
Existing installs that explicitly set any sessions.* key keep their
values (load_config deep-merges DEFAULT_CONFIG under user YAML); only
unset keys pick up the new defaults. No _config_version bump needed.
cli-config.yaml.example documents the section commented-out so
installers that copy it verbatim never pin these as explicit settings.
Tests: ratio gate (below/above/at-threshold/unknown/override), real-DB
freelist ratio, default assertions, fresh-config startup hook reaches
the prune call, explicit opt-out respected, template-does-not-pin-keys.
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:
* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
agent.tools from live probes with no predecessor to preserve. Persist
the session's resolved tool-name order in a new `sessions.tool_names`
JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
alongside the system prompt and re-pinned on every published refresh
(so /reload-mcp and compaction naturally reset it; /new mints a new
row). On restore-for-existing-session the fresh definitions are folded
onto the saved order via the SAME `_merge_preserving_prefix` helper —
a probe-flipped tool is carried forward from the registry schema, a
deregistered one dropped, new tools appended at the tail.
* /reload-mcp (CLI, gateway, TUI RPC) now also calls
`reprobe_tool_availability()` — drops the check_fn verdict cache and the
get_tool_definitions memo — so a user can consciously pick up a
credential/daemon that appeared mid-session. Docs updated.
The per-profile store partition (17ba992108, 5ffaed6e45, 5cc3da6827)
already keeps fresh rows apart, but legacy rows written to root state.db
before the partition still sat where the default profile's peer-tuple
fallback could adopt them: a Telegram DM's tuple (chat_id == user_id, no
thread) is identical for every bot. Three residual holes, closed with the
smallest predicate that fits main's design:
- hermes_state.find_latest_gateway_session_for_peer: the fallback query
now requires COALESCE(s.profile_name, <store owner>) = <store owner>
(owner via SessionDB._own_profile_name). Handles NULL legacy rows and
the single→multiplex migration case a key-namespace fence would break;
stores outside the profile tree (no derivable owner) are unchanged.
- gateway/session._recovered_row_allowed_for_active_profile: under
multiplexing no longer `return True` — the recovered row's agent:<ns>:
must match the REQUESTED key's namespace (the active profile is
meaningless when several profiles serve concurrently). Single-profile
behavior unchanged; keyless/unnamespaced rows stay adoptable.
- hermes_state create_session parent COALESCE: profile_name inherits only
when parent and child agree on agent:<ns>: (or either is keyless), so a
default child forked from a sibling row is not durably mislabelled.
Co-authored-by: pcaruba <31041167+pcaruba@users.noreply.github.com>
Co-authored-by: jiangtaoliu-source <308256854+jiangtaoliu-source@users.noreply.github.com>
Co-authored-by: 69k4xmdfm2-blip <275826864+69k4xmdfm2-blip@users.noreply.github.com>
Same behavior as the salvaged #76487 migration (fresh installs get the v3
shape; v1/v2 tables rebuild with profile_name leading the PK, legacy rows
into 'default' only, CASCADE FK supplied on the way), with the per-table
DDL written once instead of three times and the now-redundant v1->v2
CASCADE-only rebuild dropped (the v3 rebuild subsumes it).
Co-authored-by: Celio Monteiro <crdesign8@hotmail.com>
Issue #76423: under multiplex_profiles a shared state.db keyed topic mode
and bindings only by Telegram chat_id/thread_id, so private-chat ids
collided across bots/profiles.
- Add profile_name to telegram_dm_topic_mode and telegram_dm_topic_bindings
- Schema v2→v3 rebuild; legacy rows migrate into the "default" namespace
- Keyword-only profile_name="default" on SessionDB topic APIs (compat)
Follow-ups on the salvaged #101081 guard:
- A clean close() lets SQLite unlink the WAL sidecars legitimately; the
guard treated that as a lost generation and permanently halted the
handle, so the #94736 late-write self-heal reopen dropped transcript
tails (4 existing tests failed). close() now clears the recorded
sidecar generation, and _wal_generation_was_lost() re-adopts the
current sidecars after a clean /proc/self probe instead of relying on
a stale snapshot.
- Healthy writes no longer walk /proc/self/fd: once a sidecar
generation is recorded, the stat-based inode check alone detects an
unlink/replace. The fd probe only runs in the empty-identity state
(fresh DB, post-close reopen).
- DeletedWalGenerationError now subclasses StateDbReplacedError, so the
gateway retry queue and run_agent flush divert transcripts to the
JSONL fallback exactly as they do for a replaced store, instead of
retrying forever against a halted handle.
- __init__ refuses once (under the startup lock) instead of twice per
open, halving the system-wide /proc scan; dropped the dead
include_self parameter and the dead _IS_WINDOWS clause.
- Test fixes: rstrip(' (deleted)') char-set bug -> removesuffix; the
non-linux test now patches sys.platform (the real gate) instead of
_IS_WINDOWS.
A live writer can keep a deleted state.db-wal inode while a second opener
mints a fresh WAL at the same path. Fail closed on writable open (before
connect) and on the write-path sidecar identity check so the second
generation is never created.
Co-authored-by: Noa <rainbowgore@users.noreply.github.com>
Follow-up to the salvaged #99560 commit: restore the original docstring
wording (user writes still always land everywhere else), name the
llm-outranks-derived hole the guard now closes, and annotate the two new
tests with what each pins.
Compression rotates a conversation's tip id while tiles stay keyed by
whichever segment id they were opened with. focusOpenSession and
openSessionTile tested exact ids, so right after a rotation the same
chat read as 'not open' and opened again in a second tab — and a tile
keyed to a MIDDLE segment (the tip when it was opened) could no longer
prove it names the conversation at all, rendering as an untitled ghost.
The projected list row now carries the full chain
(SessionDB.get_compression_chain, served as _lineage_ids by
list_sessions_rich and the sidebar tree row), lineageAliases indexes
every segment, sessionMatchesStoredId accepts membership, and the tab
focus/open paths dedupe through the lineage instead of the exact id.
Older gateways omit the field and degrade to today's root/tip pairing.
Skipping the explicit PRAGMA wal_checkpoint(PASSIVE) in close() left
sqlite3.Connection.close() running SQLite's own last-connection PASSIVE
checkpoint, which still checkpoints the WAL and unlinks -wal/-shm on a
structurally corrupt file (E2E: the -wal vanished on close despite the
quarantine). Python 3.12+ exposes SQLITE_DBCONFIG_NO_CKPT_ON_CLOSE via
Connection.setconfig(); arm it in _halt_db_corrupt so the WAL image
survives close() for forensics/recovery. On 3.11 the switch does not
exist; the docstring and docs now say so instead of claiming sqlite3
cannot reach it at all.
Follow-up to #101095; flagged by JoaoMarcos44 on #101093.
A bare SQLITE_CORRUPT/NOTADB on a live write (not FTS-scoped, not a
replaced file) now sets a sticky per-instance flag: later writes fail
fast with StateDbCorruptError, the handle never reopens after close(),
and close() skips its explicit PASSIVE WAL checkpoint. Gateway and agent
flush paths divert pending transcripts to JSONL/spool like the replaced
case instead of retrying forever.
Field evidence: a handle that kept writing for ~50 minutes after the
first structural error checkpointed 15 pages under the wrong page
numbers on shutdown (page 1 <- messages_fts_trigram_data leaf), turning
"malformed" into "file is not a database".
Refs #90837, #90950, #97940, #89332, #45383
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CNX8rNYHqA5pT4tAGSzXtb
Two pieces of PR #100130 (@HexLab98) re-applied on top of the orphaned-flock
break (894fc35337) and fail-closed admission (#100895) that landed since:
* `is_advisory_lock_contention` (hermes_state_common): only EAGAIN /
EWOULDBLOCK / EACCES / EDEADLK mean "another process holds the lock".
ESTALE / ENOTSUP / ENOLCK / EIO from flock or msvcrt.locking are
environment failures that polling cannot fix — `_acquire_db_flock` and
both Windows msvcrt loops (FTS rebuild admission, state.db repair lock)
now defer immediately with the real errno instead of burning the full
120s / holder timeout and then logging a fake "held by another process".
* `retry_deferred_fts_recovery` (hermes_state_schema): a SessionDB whose
open-time `_recover_stale_fts` deferred (foreign holders or busy rebuild
lock) stayed `_fts_stale` — LIKE-only search — until the process
reopened state.db. Short-lived CLIs reopen every run; the gateway opens
once and stays up for days, so the deferral was effectively permanent
(#100108). The retry runs from the EXISTING gateway housekeeping tick
(`_start_gateway_housekeeping`, 60s) against the shared SessionDB
instances via `hermes_state_registry.live_shared_session_dbs()`:
non-blocking admission (`fts_rebuild_admission(timeout_seconds=0)`),
bounded backoff 60s -> 1h, no new thread, still fails closed on live
holders. `fts_rebuild_admission` gains the `timeout_seconds` kwarg.
* WAL-reset warning names `sys.executable` so a "linked SQLite 3.45.1"
line can be matched to the interpreter that actually linked it
(#100108 point 3).
Deliberately NOT carried from #100130: the "leftover lock file = holder"
premise (a 0-byte lock file never blocked flock; the real cause was the
fork-inherited fd, fixed in 894fc35337) and the `_rebuild_fts_once`
one-shot rework.
Co-authored-by: HexLab98 <liruixinch@outlook.com>
The #14694 recovery clock (`_anti_thrash_recovery_deadline`) was a
process-local `time.monotonic()` value zeroed in `bind_session_state()`.
The gateway rebuilds the AIAgent (and its ContextCompressor) on every
cache eviction, so each fresh compressor bound to a durably tripped
session row (#69872) re-armed a full 300s window and the half-open probe
never fired — a long messaging conversation above the threshold stayed
blocked permanently.
Persist the deadline as a wall-clock epoch in a new
`sessions.compression_recovery_deadline REAL` column (declarative column
reconciliation; SCHEMA_VERSION 26 -> 27) with
`SessionDB.get/set_compression_recovery_deadline`. The compressor loads it
in `bind_session_state()` and writes it on change only via
`_set_anti_thrash_recovery_deadline()`. A fresh compressor with no stored
deadline still starts a full window blocked (#54923 restart contract); one
that loads an armed deadline resumes that window. Backward clock jumps are
bounded to one window. The 300s window is unchanged.
Minimal salvage of #100185 (the probe-lease/fencing state machine and
model_config-blob storage were not carried).
Refs #100185
Co-authored-by: Komzpa <me@komzpa.net>
Follow-up on top of the salvaged #94095 commit:
- maybe_auto_prune_and_vacuum() now returns 'closed' (stale open state-owned
sessions marked ended) alongside 'pruned', so entrypoints can report the
reconciliation without parsing logs.
- Docstring explains the two-window lifecycle (close now, delete after a
further retention window).
- Regression test: cron/kanban/subagent rows with ended_at NULL are closed on
pass 1 and deleted on pass 2; a telegram row is never touched.
- website/docs sessions.md documents the automatic stale-open sweep.
state.db has two cross-process admission authorities gating destructive work
on a file several Hermes processes share: fts_rebuild_admission for full
structural FTS rebuilds, and _cross_process_repair_lock for writable_schema
surgery / VACUUM. Both document themselves as fail-closed, and both honoured
that only for a timed-out acquire. When the lock file could not be open()ed at
all they yielded True and proceeded "with in-process serialisation only" —
which is no cross-process authority whatsoever.
That inversion is reachable exactly when it does the most damage. Creating the
lock file needs a directory entry and an inode, so on a full disk open() raises
ENOSPC — while a sibling that opened ITS handle before the disk filled is still
mid-rebuild or mid-surgery. Every process then ran concurrent destructive work
on the same live DB: precisely the interleaving PR #93200 added these locks to
prevent, and the shape reported in #100368 (disk-full trigger, then a fresh
corruption on every boot with other writers alive, and no re-corruption on a
boot with zero other writers).
Both helpers now yield False on OSError. This routes the error into the
outcome the locks already define and every caller already handles: rebuild_fts
returns 0, _recover_stale_fts leaves canonical writes plus LIKE search
available behind the retryable stale breadcrumb, the startup path detaches FTS
triggers, and repair_state_db_schema re-probes and reports. Nothing reachable
is lost — on a read-only directory the rebuild's and the repair's own writes
could not have committed either. The repair report's error string now names
both ways the authority can be missing, since operators read it directly.
Desktop's cold resume (defer_history + omit_messages, transcript paged over
REST) only ever holds the live tip segment in memory, but session.resume
bounded it against the FULL compression lineage (sessions.max_resume_messages,
default 20000). A Bot Chat with 85 compaction segments / ~29k lineage rows
behind a ~700-row tip was refused at 20001, sent zero model prompts, and sat on
"Waking up default…" forever — the healthiest possible session shape, rejected
by a guard sized for in-memory materialization.
- hermes_state: one `_resume_lineage_ids` definition shared by the resume
readers (get_resume_conversations, get_ancestor_display_prefix) and the
guard (assert_resume_safe / get_resume_message_count). Guard grows
`tip_only=` and names the scope it counted; the branch-aware lineage the
readers already used is now what the guard counts too (a /branch copy was
being counted against its parent's rows).
- tui_gateway session.resume: deferred, omit_messages and lazy resumes are
bounded by the tip; only the full in-memory lineage resume keeps the
lineage-wide bound. Deferred hydration falls back to tip-only history when
the lineage exceeds the limit instead of loading the rows the guard refused.
- CLI mid-setup tip-only path routes through the same guard instead of
borrowing assert_export_safe.
- docs: sessions.max_resume_messages / max_export_messages documented with the
per-surface scope.
Live repro (real SessionDB fixture, 85 segments / 29,226 lineage rows / 666 tip
rows, real tui_gateway.server.handle_request): before — deferred resume ->
4130; after — ok, hydrated history=666 prefix=0; the non-deferred full resume
still returns 4130 on the same fixture.
A concurrent WAL checkpoint / reset / frame-flush can surface SQLITE_IOERR
to a reader on a perfectly healthy database: a mode=ro connection cannot
perform the WAL recovery the read needs, because recovery writes the -shm
index and read-only mode refuses. The window is millisecond-scale.
Today that one-shot error escapes the SessionDB read-only constructor, and
GET /api/sessions turns it into a 500 the desktop reads as an authoritative
empty list.
Retry it, bounded, in the constructor so every read-only opener is covered —
the sidebar poll, cross-profile aggregation, recall, browse — rather than at
one route. A persistent IOERR still exhausts the budget and propagates.
Remaining transient failures answer 503, so the client keeps the list it has.
On the write path, BEGIN IMMEDIATE can hit the same transient IOERR before
the callback runs. That one is safe to retry on the same connection because
nothing has been mutated; once the callback starts, settlement is unknown and
the error propagates. Never close()+reopen to heal it — close() cancels this
process's POSIX advisory locks on the file for every sibling connection, and
a list poll's reader must stay disposable so a replaced state.db is observed
and the pre-repair forensic backup stays reachable.
Fixes#100436
Co-authored-by: rkfshakti <rkfshakti@users.noreply.github.com>
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
@teknium1's maintainer-side review found no blocking defect on 09004753c9 and
listed five cleanups. All five are here.
1. scratch/repro_96811.py is deleted. It would have landed on main as a
tracked file: scratch/ is not gitignored and has never existed on main, so
this PR was creating the directory. Nothing referenced the probe, and
TestConversationGenerationRotates / TestGenerationSurvivesPruning /
TestPeerIdentityIsSourceQualified already carry all four of its stages, so
it is dropped rather than parked under tests/.
2. Upgrade notes are written into this commit body (below) and the PR body.
There is no committed changelog to add them to: scripts/release.py
generates .release_notes.md from commit SUBJECTS at release time, and
.gitignore keeps that file out of the tree.
3. declared_conversation_scope() now reads the sessions row ONCE. The fork
verdict and the source the peer queries match on both live on that row, and
asking for them separately read it twice per resolution. The new
SessionDB.declared_scope_identity() returns the pair and keeps the marker
rules beside is_explicit_fork_child() instead of re-implementing them in the
caller. A SessionDB that does not expose the combined view keeps the
original two-call path, so nothing that predates it changes behaviour --
including the three doubles that certify the fail-closed contract, which are
untouched. TestOneIdentityReadPerResolution pins the single read, the
two-call fallback, the fail-closed degrade and the fork refusal; removing
the fold turns the first of those red.
The third read stays: the generation lives in conversation_generations, a
different table, and cannot be folded into a sessions lookup.
4. _declared_conversation_session() documents the concurrent first-turn race.
Two simultaneous first requests on one declared key can each miss the
lookup, mint a row and both bind, because each row is unkeyed at bind time
and the mismatch guard does not fire. That converges rather than crossing:
both rows carry the same key under the same source, so the lookup returns
the later one for every subsequent reply and the earlier row is an abandoned
transcript, never another conversation's identity.
The same docstring still claimed the generation was durable in
sessions.end_reason and that "nothing here needs a counter". That stopped
being true in 09004753c9, which moved the generation into
conversation_generations precisely because deriving it from prunable session
rows was ABA. Corrected, along with the same stale sentence on
TestConversationBoundariesRotate.
5. conversation_generations rows are now documented as deliberately never
collected, rather than merely uncollected. Dropping one resets that peer to
"no generation", so its next boundary writes 1 again and re-issues a gwk_
scope a retired conversation already used -- the exact ABA the table exists
to close. Worth stating because the repo already carries both patterns a
maintainer would extend: delete_session() cascades to messages, and
gateway_hygiene_state is already swept by session_key.
Upgrade notes, one-time on merge:
- One cold prompt-cache bucket per keyed conversation. Every gateway platform
declares gateway_session_key, so each keyed conversation's affinity scope
moves once from its compression-lineage root session id to the gwk_ hash.
One cache miss per live conversation, on its next turn only.
- hermes status counts more sessions. A declared API conversation is now
recorded as a keyed row and appears in "Active: N session(s)" where it was
invisible. Those sessions already existed; only their visibility changes.
- A database upgraded mid-conversation starts with no generation and takes its
first from the next boundary written, so a conversation that reset before the
upgrade shares its predecessor's scope once. One warm bucket, never a crossed
identity.
Verified on this head: 55 in test_declared_conversation_scope.py (51 + 4 new),
33 in test_prompt_cache_scope.py, 49 in test_api_server_declared_conversation.py,
25 in test_api_server_runs.py, 109 in test_api_server.py, 12 in
test_cross_process_turn_lease.py, and 526 across test_hermes_state.py +
tests/hermes_state/ + tests/state/. ruff clean.
Found in review by @teknium1.
Refs #96811
Four blockers from @andrexibiza's reviews of 28a2d7f0ee and dc7865765c. The
first two are defects I introduced in 99f2d4394f by replacing the wrong
occurrence of an identical call site.
1. _run_agent raised NameError on every opted-in declared bind. Its worker
finally evaluated `if _declared_selected:`, a local of _handle_responses /
_handle_runs that is neither a parameter nor an enclosing binding here, so
the successful declared-key paths failed at settlement after the agent run.
bind_declared_conversation already IS the gate; the inner name is gone.
2. /v1/runs never received the gate at all -- it landed on _run_agent instead.
_run_sync bound unconditionally, so an explicit body session_id that existed
with an empty session_key was adopted by the header key even though the
header lost precedence. It now carries the same gate.
3. COUNT(*) + MAX(ended_at) over session rows cannot prove non-reuse.
delete_session() deletes the selected row and bulk prune selects ended rows,
so the aggregate can return a pair it already emitted:
(1,T1) -> (2,T2) -> delete boundary B -> (1,T1), handing a new conversation
a retired affinity identity. The backwards-clock shape needs no pruning at
all. The generation now lives in a conversation_generations table keyed by
(source, session_key), advanced by _bump_conversation_generation inside the
same transaction that writes each boundary -- outside prunable session
history, wall-clock-free, and increment-only. end_session() and
promote_to_session_reset() both advance it, and only when they actually
wrote a boundary, so a repeated end cannot double-count.
4. The carrier could be memoized under the wrong source. _agent_source() fell
back to agent.platform before the row landed while persistence uses
_session_source_for_agent(), which honors HERMES_SESSION_SOURCE. Because a
declared scope is non-None immediately, resolve_prompt_cache_scope memoizes
it and never re-resolves once the authoritative row appears, so under an
override both sides of a /new read the platform domain and hashed the same
scope. The pre-row path now uses the persistence resolver itself.
Coverage answers the review's specific objection that mocked tests proved the
mock rather than the path. TestRealRunAgentSettlement stubs _create_agent and
lets the real _run_agent settle; the /v1/runs case persists an unkeyed explicit
row and waits for the worker to retire before asserting. Both were verified by
mutation: reinstating the inner name fails two of them, and removing the
/v1/runs gate fails the explicit-session one. The first version of that test
passed with the gate removed -- it asserted before settlement -- and would have
been the same empty proof the review called out.
TestGenerationSurvivesPruning covers deleting the newest boundary, deleting
every boundary, the backwards-clock-then-prune shape, compression and
accidental ends not advancing it, repeated ends not double-counting, promotion
advancing it, unkeyed rows advancing nothing, and peer scoping.
TestSourceOverrideDomain covers the override across a reset.
Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.
Refs #96811
Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
Both blockers from @andrexibiza's review of 28a2d7f0ee.
1. The generation lookup was not in the same identity domain as recovery.
latest_conversation_boundary() selected on session_key alone, while
_declared_conversation_session() is qualified by (source, session_key).
X-Hermes-Session-Key accepts any authenticated caller-supplied string, so an
API conversation may legally carry the same key as a Telegram row in one
database -- a /new over there rotated this conversation's gwk_ generation
while recovery correctly refused to cross the same line, moving the affinity
identity out from under a physical identity that had not moved.
The boundary read now takes (session_key, source), and the carrier is
'source|key|generation' rather than 'key|generation' -- keying on the string
alone would also collapse two same-key conversations from different sources
onto one routing key, since this value leaves the process verbatim as
OpenRouter's sticky session_id and xAI's x-grok-conv-id. The source comes
from the agent's own session row, falling back to the platform the row will
be created with before it lands.
2. The declared key's stated lower precedence did not survive settlement. Both
handlers let stored_session_id / an explicit body session_id win, then called
_bind_declared_conversation() unconditionally. record_gateway_session_peer()
does SET session_key = ? across compression ancestors, so a request carrying
conversation A's chain plus header key B silently rebound A to B: A could no
longer be recovered by its own key, and B recovered A's session.
Recording is now gated on the declared key having actually selected or
minted the session, on both paths. Behind that gate the bind itself refuses
to overwrite a row already bound to a different key, so a future caller
cannot reintroduce the same defect by opting in wrongly.
test_declaration_outranks_the_lineage_root asserted the pre-qualification
contract by comparing a DB-backed agent against a DB-less one; it now makes the
stronger statement it was written for -- one declared conversation reached
through two different physical ids on the same peer.
Refs #96811
Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.
Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
Self-review of the generation marker. MAX(ended_at) alone is wall-clock: an
NTP correction between two resets writes a SMALLER boundary, MAX keeps
returning the older one, and the next conversation silently reuses the
previous generation -- two conversations on one routing key, which is the
defect this PR exists to remove.
latest_conversation_boundary now returns (count, latest_ended_at) and the
marker is 'count:ended_at'. The two halves fail under different conditions --
a backwards clock defeats the timestamp, retention pruning of an old ended row
decrements the count -- so a generation repeats only if both happen at once.
The pair is deliberately biased toward changing: a spurious change costs one
cold prompt-cache bucket, a repeat would merge two conversations.
Pinned by test_a_backwards_clock_does_not_reuse_a_generation, which rewrites
the second boundary to land before the first and asserts three conversations
still resolve to three distinct scopes.
Refs #96811
The declared key is a per-CHAT identifier and outlives the conversation it
names: reset_session() mints a fresh physical id on /new but keeps the key, and
the idle/daily/suspended policy resets do the same. Hashing the key alone
therefore mapped the conversation before a reset and the one after it onto one
gwk_ scope -- the lifecycle violation @cervantesh raised on #97158 and
@kshitijk4poor reproduced on #97709.
No counter is introduced. The generation that must rotate is already durable:
every one of those boundaries closes the outgoing row with an
_RESET_END_REASONS end_reason, so SessionDB.latest_conversation_boundary reads
the most recent one and declared_conversation_scope hashes 'key|generation'.
That makes the carrier stable across a host's per-response physical ids -- a
host that never resets writes no boundary, so every reply hashes the same value
-- while rotating on every conversation replacement, /new and the policy
auto-resets alike. ended_at only moves forward, so a retired generation can
never be reused: no ABA.
It also cannot drift from the rest of the codebase's notion of a conversation
boundary, because find_latest_gateway_session_for_peer fences on the same set.
The read is on the memoized resolution path, not per API call, and both lookups
fail closed: an unqualified key would span a /new, so a DB error degrades to the
physical-id scope. A SessionDB without the lookup keeps the previous behaviour.
Refs #96811
Every conversation-affinity hint Hermes sends is derived from the PHYSICAL
session id: prompt_cache_key on both OpenAI-wire transports, OpenRouter's and
Nous Portal's sticky session_id, and xAI's x-grok-conv-id. A host that mints
one physical session per RESPONSE re-keys all four on every reply, so the
conversation never lands back on the routing bucket it just warmed (#96811).
Two hosts do exactly that. Hermes Studio's group chat mints
gc_run_<room>_<profile>_<name>_<uuid4hex> per reply and destroys it after,
and POST /v1/responses with client-managed history mints str(uuid4()) per
request — while parsing X-Hermes-Session-Key one screen earlier and handing
it to the agent.
Hermes must not infer the logical conversation from the id's syntax: that
rule merges independent client-supplied ids and Studio members truncated past
its 96-character boundary (the #79017 failure class). It does not have to.
gateway_session_key is already the "stable per-chat key" built by
gateway.session.build_session_key from that header, and branching
deliberately does not key off it. The affinity path simply never consulted it.
- agent/prompt_cache_scope.py: declared_conversation_scope() resolves the key
into gwk_<sha256[:24]> and outranks the lineage walk (it is stable across
rotation AND across per-response ids). Hashed because, unlike a session id,
the key embeds platform/chat/user identifiers and leaves the process
verbatim as a sticky id and as x-grok-conv-id.
- agent/portal_tags.py: a separate ambient scope for ROUTING, published only
when a host declared one. The providers read the attribution id when it is
unset, so delegate trees keep sharing their parent's sticky key and every
host that keeps one id per conversation is byte-identical to before.
- hermes_state.py: is_explicit_fork_child() — the public view of the marker
rules that keep /branch children, delegate subagents and tool children off
their parent's chat key. Background-review forks clone the live runtime, so
_persist_disabled excludes them for the same reason (#79161).
Refs #96570Fixes#96811
Desktop opens the registry id, then session.resume walked the legacy
unmarked-child fallback, so Open Chat still landed in a side chat after
the title lookup was already strict. Recoverable-archive resurrection
uses the same helper.
Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
_record_db_file_identity's PRAGMA fallback is a pure read — route it
through _read_ctx() instead of the writer lock (Pattern C gate).
test_codex_turn_persists_each_message_exactly_once leaked a live
SessionDB into shutil.rmtree, racing the WAL sidecars ('Directory not
empty' on CI); close the handle first and rmtree with ignore_errors.
The repo-wide conn-lock audit (tests/test_hermes_state_conn_lock_audit.py)
requires every self._conn statement outside construction to hold
self._lock — an unlocked statement races SessionDB.close() inside
pysqlite's statement cache (#99349). Wrap _ensure_db_file_generation's
statements and _record_db_file_identity's PRAGMA fallback in the lock;
both call sites run outside the lock, so no re-entry.
Detect same-inode cp via a generation stamp, halt FTS repair, and divert
unwritten transcripts to sessions/<id>.jsonl plus the gateway pending spool.
Co-authored-by: Cursor <cursoragent@cursor.com>
Addresses Enough1122's two non-blocking review notes on PR #92316:
1. Lease vs deadline arithmetic: the state.db init/migration/repair leases
(600-900s) are authoritative against the 300s default deadline by design.
Add explicit comments at all three lease sites pinning the trade-off:
single lease is deliberate (clamped to _MAX_LEASE_S=900), honest worst
case is up to the lease duration of zombie time on a wedged DB phase,
accepted over per-chunk renewal complexity in the migration loops.
2. Arm-site coverage: add a structural contract test asserting every
documented entry point (hermes_cli/main.py argv fast-path,
hermes_cli/gateway.py config-bridge re-arm, gateway/run.py backstop,
cli.py legacy --gateway) actually calls arm_startup_watchdog (or its
aliased import), so a future entry point can't silently ship unwatched.
Addresses the two class-level review blockers on PR #89750:
1. Bounded hard-exit seam (escort thread). The forensic fire path
(logger.critical, dump record, faulthandler, lifecycle ledger) can
itself wedge — the parked main thread may hold the logging handler
lock, or the disk may be full/hung. _fire() now starts an exit-escort
daemon thread BEFORE any forensics; it is free of log handlers,
filesystem access, module loads and application locks, and hard-exits
with the restart code after _FIRE_EXIT_BOUND_S unless the normal fire
path signals completion. Adversarial tests hold the logging handler
lock / hang the dump write at fire time and assert the exit seam is
still reached.
2. Phase-owned progress leases (report_startup_progress). Process CPU
time proves process activity, not startup progress: an unrelated busy
thread could extend forever while startup sits parked (false
negative), and I/O-bound repair/backup accrues ~zero CPU and would be
killed (false positive). Long synchronous startup phases now declare
authoritative, clamped (_MAX_LEASE_S), renewable progress leases:
state.db _init_schema + the version-gated data-migration chain
(hermes_state_schema) and repair_state_db_schema (hermes_state) are
wired. CPU progress remains only as a bounded fallback, capped at
_MAX_CPU_EXTENSIONS, with leases outranking the cap. Adversarial
tests cover both directions (lease saves zero-CPU legitimate work;
capped CPU noise no longer hides a parked deadlock).
Fire-path dump record now includes lease_count/last_lease_phase for
forensics. gateway/startup_watchdog.py shim re-exports
report_startup_progress.
OOF-298
teknium flagged it as scope creep on PR #98691 review: no callers in the diff.
_is_explicit_fork_child_row (the row-based helper actually used) is unchanged.