Commit Graph

501 Commits

Author SHA1 Message Date
Teknium eb8d628c97 refactor(hermes_state): restore WHY comments dropped by round-2 sub-branches
Comment/docstring-only (AST-identical): surrogate-scrub rationale, persisted
marker stripping invariant, generation counter upgrade semantics, CJK marker
empty-vs-populated rule, WAL 0-page ordering precondition, repair backup
live-connection case, telegram topic delete precondition, mixed-mode
corruption definition, and similar.
2026-09-02 16:48:30 -07:00
Teknium b32dbc98f5 refactor(state): final docstring/comment tightening (IDENTICAL-AST) 2026-09-02 16:27:31 -07:00
Teknium 2ae3e4d6bd refactor(state): share /proc fd walker, collapse loops in generation/holder probes 2026-09-02 16:22:10 -07:00
Teknium f4e6ccc5f5 refactor(state): unify db-generation halt guard, hoist _is_no_more_rows, collapse flag inits 2026-09-02 16:17:27 -07:00
Teknium 40fe17f83b refactor(state): fold adjacent string fragments that fit one line (AST-identical) 2026-09-02 16:14:54 -07:00
Teknium 0fe9d5f5ae refactor(state): AST-neutral formatting compaction (blank lines, signatures, arg lists) 2026-09-02 16:13:30 -07:00
Teknium c023f00c9c refactor(state): compact session-CRUD comments/docstrings (keep every WHY) 2026-09-02 16:09:13 -07:00
Teknium e4947b6871 refactor(state): compact SessionDB core comments/docstrings (keep every WHY) 2026-09-02 16:00:21 -07:00
Teknium 0679a0ca15 refactor(state): compact module-level comments/docstrings (keep every WHY) 2026-09-02 15:54:56 -07:00
Teknium 2242f55e9f refactor(state): unify session CRUD helpers, list filter builder, compact re-exports 2026-09-02 15:49:46 -07:00
Teknium 62e8c8ffbc refactor(state): lift SessionDB.__init__ nested closures into methods 2026-09-02 15:29:40 -07:00
Teknium d15c61b5dc refactor(state): split SessionDB into domain mixins and free-function modules; unify SQL boilerplate
hermes_state.py 17,220 -> 6,442 LOC. Behavior-neutral: every moved body is
AST-identical to the original, verified per extraction.

SessionDB core
- _write_sql / _write_rowcount / _read_one / _read_all replace ~120 copies of
  the `def _do(conn): conn.execute(...)` + `_execute_write(_do)` and
  `with self._read_ctx() as conn: row = conn.execute(...).fetchone()` shapes.
- _set_lineage_column replaces four copies of the recursive compression-lineage
  UPDATE (archived / pinned / hidden / last_read_at).
- _read_session_number unifies the three compression counter readers.
- Dead (zero refs repo-wide): restore_rewound, delete_gateway_routing_entries,
  _is_duplicate_replayed_user_message, SessionPortabilityMixin.get_first_assistant_text.

New mixins bound onto SessionDB via the MRO (logger name stays "hermes_state"):
  hermes_state_messages    SessionMessagesMixin       48 methods
  hermes_state_compression SessionCompressionMixin    30
  hermes_state_gateway     SessionGatewayMixin        26
  hermes_state_maintenance SessionMaintenanceMixin    13
  hermes_state_usage       SessionUsageMixin          12
  hermes_state_titles      SessionTitlesMixin         13
  hermes_state_telegram    SessionTelegramTopicsMixin 11
Origin-internal symbols resolve through a lazy `from hermes_state import ...`
inside the few methods that need them (no import cycle).

New free-function modules, every name re-imported into hermes_state so
`hermes_state.<name>` (and test monkeypatches on it) keep working; intra-module
calls to patched helpers go through the lazy origin import:
  hermes_state_repair   repair/backup/preflight (43 defs)
  hermes_state_wal      journal-mode / PRAGMA policy (33 defs)
  hermes_state_dbfile   header probes, zeroed-db quarantine, stats, holders (21 defs)

Existing mixins: search — shared FTS MATCH/LIKE builders, unified rebuild
status/step/finish engines, state_meta helpers; schema — one legacy/v23 FTS init
branch, shared _live_pk_columns, Row/tuple dual access dropped; portability —
shared _PREVIEW_RAW_SUBQUERY_SQL and _rich_row; common — single
stat_db_file_identity (was 3 copies), AUTO_VACUUM_MIN_FREELIST_RATIO.

Docstrings/comments hand-compacted (AST-identical) keeping every invariant,
ordering rule, failure mode and WHY. Schema SQL, migration order and PRAGMAs
untouched. test_repair_path_has_no_bare_connects repointed to hermes_state_repair.
2026-09-02 13:32:13 -07:00
Brooklyn Nicholson ab281990b8 fix(sessions): include compaction-archived rows in every display projection
The gateway's three display reads — session.resume, the ancestor lineage
prefix, and the warm-session payload — filtered active = 1, so a compacted
conversation rendered as its summary plus the carried-forward tail. The REST
transcript read has included the archived rows since #80680, so the same
session read two ways gave two different answers.

Extract the generation-dedupe policy the REST read carried inline into a
shared _dedupe_display_generations() and point every display projection at
it. The model-fed history stays active-only: compaction still compresses the
working context, it just no longer erases the user's transcript.
2026-09-02 12:11:15 -05:00
Teknium 73f68362b3 fix(sessions): auto-prune state.db by default (90d) and gate VACUUM on freelist ratio (#54189)
Flip the state.db retention defaults per Teknium's decision on #54189:

- sessions.auto_prune: false -> true. A stock install now prunes ENDED
  sessions inactive for retention_days at CLI/gateway/cron startup
  (at most once per min_interval_hours). Open, pinned and mid-turn
  sessions are never deleted; the only open rows touched are stale
  automation sessions (#100903 sweep), which are closed, not deleted,
  and aged a further full window before removal.
- sessions.retention_days stays 90 (already the default; verified).
- Auto-VACUUM is now additionally gated on the reclaimable fraction of
  the file: PRAGMA freelist_count / page_count must exceed 25%
  (AUTO_VACUUM_MIN_FREELIST_RATIO) on top of the existing
  min_vacuum_interval_days throttle. Pruning a few small sessions on a
  dense multi-GB DB no longer rewrites the whole file to reclaim a few MB.
  Unknown ratio (pragma read failure) falls back to the time throttle.

Existing installs that explicitly set any sessions.* key keep their
values (load_config deep-merges DEFAULT_CONFIG under user YAML); only
unset keys pick up the new defaults. No _config_version bump needed.
cli-config.yaml.example documents the section commented-out so
installers that copy it verbatim never pin these as explicit settings.

Tests: ratio gate (below/above/at-threshold/unknown/override), real-DB
freelist ratio, default assertions, fresh-config startup hook reaches
the prune call, explicit opt-out respected, template-does-not-pin-keys.
2026-09-02 07:26:52 -07:00
Teknium 8e4366d358 fix(tools): freeze tools[] across agent-cache eviction; make /reload-mcp the re-probe hatch
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:

* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
  rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
  agent.tools from live probes with no predecessor to preserve. Persist
  the session's resolved tool-name order in a new `sessions.tool_names`
  JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
  alongside the system prompt and re-pinned on every published refresh
  (so /reload-mcp and compaction naturally reset it; /new mints a new
  row). On restore-for-existing-session the fresh definitions are folded
  onto the saved order via the SAME `_merge_preserving_prefix` helper —
  a probe-flipped tool is carried forward from the registry schema, a
  deregistered one dropped, new tools appended at the tail.

* /reload-mcp (CLI, gateway, TUI RPC) now also calls
  `reprobe_tool_availability()` — drops the check_fn verdict cache and the
  get_tool_definitions memo — so a user can consciously pick up a
  credential/daemon that appeared mid-session. Docs updated.
2026-09-02 07:22:59 -07:00
Teknium 7463cd1202 fix(session): fence multiplex peer-fallback recovery and profile inheritance by owner (#74285, #88381)
The per-profile store partition (17ba992108, 5ffaed6e45, 5cc3da6827)
already keeps fresh rows apart, but legacy rows written to root state.db
before the partition still sat where the default profile's peer-tuple
fallback could adopt them: a Telegram DM's tuple (chat_id == user_id, no
thread) is identical for every bot. Three residual holes, closed with the
smallest predicate that fits main's design:

- hermes_state.find_latest_gateway_session_for_peer: the fallback query
  now requires COALESCE(s.profile_name, <store owner>) = <store owner>
  (owner via SessionDB._own_profile_name). Handles NULL legacy rows and
  the single→multiplex migration case a key-namespace fence would break;
  stores outside the profile tree (no derivable owner) are unchanged.
- gateway/session._recovered_row_allowed_for_active_profile: under
  multiplexing no longer `return True` — the recovered row's agent:<ns>:
  must match the REQUESTED key's namespace (the active profile is
  meaningless when several profiles serve concurrently). Single-profile
  behavior unchanged; keyless/unnamespaced rows stay adoptable.
- hermes_state create_session parent COALESCE: profile_name inherits only
  when parent and child agree on agent:<ns>: (or either is keyless), so a
  default child forked from a sibling row is not durably mislabelled.

Co-authored-by: pcaruba <31041167+pcaruba@users.noreply.github.com>
Co-authored-by: jiangtaoliu-source <308256854+jiangtaoliu-source@users.noreply.github.com>
Co-authored-by: 69k4xmdfm2-blip <275826864+69k4xmdfm2-blip@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
Teknium 458e2ef1d2 refactor(state): collapse telegram topic v3 migration into one table-driven rebuild
Same behavior as the salvaged #76487 migration (fresh installs get the v3
shape; v1/v2 tables rebuild with profile_name leading the PK, legacy rows
into 'default' only, CASCADE FK supplied on the way), with the per-table
DDL written once instead of three times and the now-redundant v1->v2
CASCADE-only rebuild dropped (the v3 rebuild subsumes it).

Co-authored-by: Celio Monteiro <crdesign8@hotmail.com>
2026-09-02 05:59:24 -07:00
Celio Monteiro 351e4c0067 fix(state): namespace telegram topic tables by profile_name
Issue #76423: under multiplex_profiles a shared state.db keyed topic mode
and bindings only by Telegram chat_id/thread_id, so private-chat ids
collided across bots/profiles.

- Add profile_name to telegram_dm_topic_mode and telegram_dm_topic_bindings
- Schema v2→v3 rebuild; legacy rows migrate into the "default" namespace
- Keyword-only profile_name="default" on SessionDB topic APIs (compat)
2026-09-02 05:59:24 -07:00
kshitijk4poor e9fa7bc05e fix(state): keep the #94736 teardown self-heal alive under the deleted-WAL guard
Follow-ups on the salvaged #101081 guard:

- A clean close() lets SQLite unlink the WAL sidecars legitimately; the
  guard treated that as a lost generation and permanently halted the
  handle, so the #94736 late-write self-heal reopen dropped transcript
  tails (4 existing tests failed). close() now clears the recorded
  sidecar generation, and _wal_generation_was_lost() re-adopts the
  current sidecars after a clean /proc/self probe instead of relying on
  a stale snapshot.
- Healthy writes no longer walk /proc/self/fd: once a sidecar
  generation is recorded, the stat-based inode check alone detects an
  unlink/replace. The fd probe only runs in the empty-identity state
  (fresh DB, post-close reopen).
- DeletedWalGenerationError now subclasses StateDbReplacedError, so the
  gateway retry queue and run_agent flush divert transcripts to the
  JSONL fallback exactly as they do for a replaced store, instead of
  retrying forever against a halted handle.
- __init__ refuses once (under the startup lock) instead of twice per
  open, halving the system-wide /proc scan; dropped the dead
  include_self parameter and the dead _IS_WINDOWS clause.
- Test fixes: rstrip(' (deleted)') char-set bug -> removesuffix; the
  non-linux test now patches sys.platform (the real gate) instead of
  _IS_WINDOWS.
2026-09-02 18:24:54 +05:30
Cursor Agent 7f7df1ce44 fix(state): refuse SessionDB open and writes on a deleted WAL generation
A live writer can keep a deleted state.db-wal inode while a second opener
mints a fresh WAL at the same path. Fail closed on writable open (before
connect) and on the write-path sidecar identity check so the second
generation is never created.

Co-authored-by: Noa <rainbowgore@users.noreply.github.com>
2026-09-02 18:24:54 +05:30
Teknium 89e2e4f572 docs(sessions): explain why the canonical Bot Chat guard is provenance-blind (#99517)
Follow-up to the salvaged #99560 commit: restore the original docstring
wording (user writes still always land everywhere else), name the
llm-outranks-derived hole the guard now closes, and annotate the two new
tests with what each pins.
2026-09-02 05:39:24 -07:00
fangliquanflq fb9a294794 fix(sessions): protect canonical Bot Chat from auto-title 2026-09-02 05:39:24 -07:00
Alonso 202997b51d fix(desktop): one conversation never opens as two tabs after compaction
Compression rotates a conversation's tip id while tiles stay keyed by
whichever segment id they were opened with. focusOpenSession and
openSessionTile tested exact ids, so right after a rotation the same
chat read as 'not open' and opened again in a second tab — and a tile
keyed to a MIDDLE segment (the tip when it was opened) could no longer
prove it names the conversation at all, rendering as an untitled ghost.

The projected list row now carries the full chain
(SessionDB.get_compression_chain, served as _lineage_ids by
list_sessions_rich and the sidebar tree row), lineageAliases indexes
every segment, sessionMatchesStoredId accepts membership, and the tab
focus/open paths dedupe through the lineage instead of the exact id.
Older gateways omit the field and degrade to today's root/tip pairing.
2026-09-02 05:37:05 -07:00
kshitijk4poor e9160625dc fix(state): also disable SQLite's internal close-time checkpoint on quarantine (py3.12+)
Skipping the explicit PRAGMA wal_checkpoint(PASSIVE) in close() left
sqlite3.Connection.close() running SQLite's own last-connection PASSIVE
checkpoint, which still checkpoints the WAL and unlinks -wal/-shm on a
structurally corrupt file (E2E: the -wal vanished on close despite the
quarantine). Python 3.12+ exposes SQLITE_DBCONFIG_NO_CKPT_ON_CLOSE via
Connection.setconfig(); arm it in _halt_db_corrupt so the WAL image
survives close() for forensics/recovery. On 3.11 the switch does not
exist; the docstring and docs now say so instead of claiming sqlite3
cannot reach it at all.

Follow-up to #101095; flagged by JoaoMarcos44 on #101093.
2026-09-02 16:57:21 +05:30
leomcamilo bcc2e65818 fix(state): quarantine SessionDB handle after structural corruption
A bare SQLITE_CORRUPT/NOTADB on a live write (not FTS-scoped, not a
replaced file) now sets a sticky per-instance flag: later writes fail
fast with StateDbCorruptError, the handle never reopens after close(),
and close() skips its explicit PASSIVE WAL checkpoint. Gateway and agent
flush paths divert pending transcripts to JSONL/spool like the replaced
case instead of retrying forever.

Field evidence: a handle that kept writing for ~50 minutes after the
first structural error checkpointed 15 pages under the wrong page
numbers on shutdown (page 1 <- messages_fts_trigram_data leaf), turning
"malformed" into "file is not a database".

Refs #90837, #90950, #97940, #89332, #45383

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CNX8rNYHqA5pT4tAGSzXtb
2026-09-02 16:57:21 +05:30
teknium1 fd05029430 fix(state): fail fast on non-contention flock errors and retry deferred FTS rebuilds in-process (salvage #100130)
Two pieces of PR #100130 (@HexLab98) re-applied on top of the orphaned-flock
break (894fc35337) and fail-closed admission (#100895) that landed since:

* `is_advisory_lock_contention` (hermes_state_common): only EAGAIN /
  EWOULDBLOCK / EACCES / EDEADLK mean "another process holds the lock".
  ESTALE / ENOTSUP / ENOLCK / EIO from flock or msvcrt.locking are
  environment failures that polling cannot fix — `_acquire_db_flock` and
  both Windows msvcrt loops (FTS rebuild admission, state.db repair lock)
  now defer immediately with the real errno instead of burning the full
  120s / holder timeout and then logging a fake "held by another process".

* `retry_deferred_fts_recovery` (hermes_state_schema): a SessionDB whose
  open-time `_recover_stale_fts` deferred (foreign holders or busy rebuild
  lock) stayed `_fts_stale` — LIKE-only search — until the process
  reopened state.db. Short-lived CLIs reopen every run; the gateway opens
  once and stays up for days, so the deferral was effectively permanent
  (#100108). The retry runs from the EXISTING gateway housekeeping tick
  (`_start_gateway_housekeeping`, 60s) against the shared SessionDB
  instances via `hermes_state_registry.live_shared_session_dbs()`:
  non-blocking admission (`fts_rebuild_admission(timeout_seconds=0)`),
  bounded backoff 60s -> 1h, no new thread, still fails closed on live
  holders. `fts_rebuild_admission` gains the `timeout_seconds` kwarg.

* WAL-reset warning names `sys.executable` so a "linked SQLite 3.45.1"
  line can be matched to the interpreter that actually linked it
  (#100108 point 3).

Deliberately NOT carried from #100130: the "leftover lock file = holder"
premise (a 0-byte lock file never blocked flock; the real cause was the
fork-inherited fd, fixed in 894fc35337) and the `_rebuild_fts_once`
one-shot rework.

Co-authored-by: HexLab98 <liruixinch@outlook.com>
2026-09-02 04:15:02 -07:00
Teknium 238b6c1ab9 fix(compression): persist the anti-thrash recovery deadline so gateway agent rebuilds cannot block a session forever
The #14694 recovery clock (`_anti_thrash_recovery_deadline`) was a
process-local `time.monotonic()` value zeroed in `bind_session_state()`.
The gateway rebuilds the AIAgent (and its ContextCompressor) on every
cache eviction, so each fresh compressor bound to a durably tripped
session row (#69872) re-armed a full 300s window and the half-open probe
never fired — a long messaging conversation above the threshold stayed
blocked permanently.

Persist the deadline as a wall-clock epoch in a new
`sessions.compression_recovery_deadline REAL` column (declarative column
reconciliation; SCHEMA_VERSION 26 -> 27) with
`SessionDB.get/set_compression_recovery_deadline`. The compressor loads it
in `bind_session_state()` and writes it on change only via
`_set_anti_thrash_recovery_deadline()`. A fresh compressor with no stored
deadline still starts a full window blocked (#54923 restart contract); one
that loads an armed deadline resumes that window. Backward clock jumps are
bounded to one window. The 300s window is unchanged.

Minimal salvage of #100185 (the probe-lease/fencing state machine and
model_config-blob storage were not carried).

Refs #100185
Co-authored-by: Komzpa <me@komzpa.net>
2026-09-02 04:14:10 -07:00
Teknium 9bc249c7e5 fix(state): report closed stale-open count from auto-maintenance, document the sweep (#54189)
Follow-up on top of the salvaged #94095 commit:
- maybe_auto_prune_and_vacuum() now returns 'closed' (stale open state-owned
  sessions marked ended) alongside 'pruned', so entrypoints can report the
  reconciliation without parsing logs.
- Docstring explains the two-window lifecycle (close now, delete after a
  further retention window).
- Regression test: cron/kanban/subagent rows with ended_at NULL are closed on
  pass 1 and deleted on pass 2; a telegram row is never touched.
- website/docs sessions.md documents the automatic stale-open sweep.
2026-09-02 01:20:47 -07:00
Alex 0aa84bb3fd fix(state): reap stale state-owned sessions safely 2026-09-02 01:20:47 -07:00
HexLab98 f5f4cefc3d fix(state): fail closed when a state.db admission lock file cannot be opened
state.db has two cross-process admission authorities gating destructive work
on a file several Hermes processes share: fts_rebuild_admission for full
structural FTS rebuilds, and _cross_process_repair_lock for writable_schema
surgery / VACUUM. Both document themselves as fail-closed, and both honoured
that only for a timed-out acquire. When the lock file could not be open()ed at
all they yielded True and proceeded "with in-process serialisation only" —
which is no cross-process authority whatsoever.

That inversion is reachable exactly when it does the most damage. Creating the
lock file needs a directory entry and an inode, so on a full disk open() raises
ENOSPC — while a sibling that opened ITS handle before the disk filled is still
mid-rebuild or mid-surgery. Every process then ran concurrent destructive work
on the same live DB: precisely the interleaving PR #93200 added these locks to
prevent, and the shape reported in #100368 (disk-full trigger, then a fresh
corruption on every boot with other writers alive, and no re-corruption on a
boot with zero other writers).

Both helpers now yield False on OSError. This routes the error into the
outcome the locks already define and every caller already handles: rebuild_fts
returns 0, _recover_stale_fts leaves canonical writes plus LIKE search
available behind the retryable stale breadcrumb, the startup path detaches FTS
triggers, and repair_state_db_schema re-probes and reports. Nothing reachable
is lost — on a read-only directory the rebuild's and the repair's own writes
could not have committed either. The repair report's error string now names
both ways the authority can be missing, since operators read it directly.
2026-09-01 23:58:51 -07:00
Teknium 52f359a011 fix(sessions): Desktop resume of a heavily-compacted chat no longer fails with 4130
Desktop's cold resume (defer_history + omit_messages, transcript paged over
REST) only ever holds the live tip segment in memory, but session.resume
bounded it against the FULL compression lineage (sessions.max_resume_messages,
default 20000). A Bot Chat with 85 compaction segments / ~29k lineage rows
behind a ~700-row tip was refused at 20001, sent zero model prompts, and sat on
"Waking up default…" forever — the healthiest possible session shape, rejected
by a guard sized for in-memory materialization.

- hermes_state: one `_resume_lineage_ids` definition shared by the resume
  readers (get_resume_conversations, get_ancestor_display_prefix) and the
  guard (assert_resume_safe / get_resume_message_count). Guard grows
  `tip_only=` and names the scope it counted; the branch-aware lineage the
  readers already used is now what the guard counts too (a /branch copy was
  being counted against its parent's rows).
- tui_gateway session.resume: deferred, omit_messages and lazy resumes are
  bounded by the tip; only the full in-memory lineage resume keeps the
  lineage-wide bound. Deferred hydration falls back to tip-only history when
  the lineage exceeds the limit instead of loading the rows the guard refused.
- CLI mid-setup tip-only path routes through the same guard instead of
  borrowing assert_export_safe.
- docs: sessions.max_resume_messages / max_export_messages documented with the
  per-surface scope.

Live repro (real SessionDB fixture, 85 segments / 29,226 lineage rows / 666 tip
rows, real tui_gateway.server.handle_request): before — deferred resume ->
4130; after — ok, hydrated history=666 prefix=0; the non-deferred full resume
still returns 4130 on the same fixture.
2026-09-01 22:14:33 -07:00
Brooklyn Nicholson 3a0e7df799 fix(state): a busy session store reads as busy, not as damaged or empty
A concurrent WAL checkpoint / reset / frame-flush can surface SQLITE_IOERR
to a reader on a perfectly healthy database: a mode=ro connection cannot
perform the WAL recovery the read needs, because recovery writes the -shm
index and read-only mode refuses. The window is millisecond-scale.

Today that one-shot error escapes the SessionDB read-only constructor, and
GET /api/sessions turns it into a 500 the desktop reads as an authoritative
empty list.

Retry it, bounded, in the constructor so every read-only opener is covered —
the sidebar poll, cross-profile aggregation, recall, browse — rather than at
one route. A persistent IOERR still exhausts the budget and propagates.
Remaining transient failures answer 503, so the client keeps the list it has.

On the write path, BEGIN IMMEDIATE can hit the same transient IOERR before
the callback runs. That one is safe to retry on the same connection because
nothing has been mutated; once the callback starts, settlement is unknown and
the error propagates. Never close()+reopen to heal it — close() cancels this
process's POSIX advisory locks on the file for every sibling connection, and
a list poll's reader must stay disposable so a replaced state.db is observed
and the pre-repair forensic backup stays reachable.

Fixes #100436

Co-authored-by: rkfshakti <rkfshakti@users.noreply.github.com>
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
2026-09-01 22:42:19 -05:00
Teknium 894fc35337 fix(state): break provably-orphaned repair/FTS-rebuild locks left by dead holders (#100108) 2026-09-01 10:52:06 -07:00
Teknium 09b88bab88 fix(state): stop the on-write identity probe cancelling our own POSIX locks (#100368) 2026-09-01 10:51:52 -07:00
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
joaomarcos 6b9b3e0145 chore(cache): take the pre-merge cleanups on the declared conversation scope
@teknium1's maintainer-side review found no blocking defect on 09004753c9 and
listed five cleanups. All five are here.

1. scratch/repro_96811.py is deleted. It would have landed on main as a
   tracked file: scratch/ is not gitignored and has never existed on main, so
   this PR was creating the directory. Nothing referenced the probe, and
   TestConversationGenerationRotates / TestGenerationSurvivesPruning /
   TestPeerIdentityIsSourceQualified already carry all four of its stages, so
   it is dropped rather than parked under tests/.

2. Upgrade notes are written into this commit body (below) and the PR body.
   There is no committed changelog to add them to: scripts/release.py
   generates .release_notes.md from commit SUBJECTS at release time, and
   .gitignore keeps that file out of the tree.

3. declared_conversation_scope() now reads the sessions row ONCE. The fork
   verdict and the source the peer queries match on both live on that row, and
   asking for them separately read it twice per resolution. The new
   SessionDB.declared_scope_identity() returns the pair and keeps the marker
   rules beside is_explicit_fork_child() instead of re-implementing them in the
   caller. A SessionDB that does not expose the combined view keeps the
   original two-call path, so nothing that predates it changes behaviour --
   including the three doubles that certify the fail-closed contract, which are
   untouched. TestOneIdentityReadPerResolution pins the single read, the
   two-call fallback, the fail-closed degrade and the fork refusal; removing
   the fold turns the first of those red.

   The third read stays: the generation lives in conversation_generations, a
   different table, and cannot be folded into a sessions lookup.

4. _declared_conversation_session() documents the concurrent first-turn race.
   Two simultaneous first requests on one declared key can each miss the
   lookup, mint a row and both bind, because each row is unkeyed at bind time
   and the mismatch guard does not fire. That converges rather than crossing:
   both rows carry the same key under the same source, so the lookup returns
   the later one for every subsequent reply and the earlier row is an abandoned
   transcript, never another conversation's identity.

   The same docstring still claimed the generation was durable in
   sessions.end_reason and that "nothing here needs a counter". That stopped
   being true in 09004753c9, which moved the generation into
   conversation_generations precisely because deriving it from prunable session
   rows was ABA. Corrected, along with the same stale sentence on
   TestConversationBoundariesRotate.

5. conversation_generations rows are now documented as deliberately never
   collected, rather than merely uncollected. Dropping one resets that peer to
   "no generation", so its next boundary writes 1 again and re-issues a gwk_
   scope a retired conversation already used -- the exact ABA the table exists
   to close. Worth stating because the repo already carries both patterns a
   maintainer would extend: delete_session() cascades to messages, and
   gateway_hygiene_state is already swept by session_key.

Upgrade notes, one-time on merge:

- One cold prompt-cache bucket per keyed conversation. Every gateway platform
  declares gateway_session_key, so each keyed conversation's affinity scope
  moves once from its compression-lineage root session id to the gwk_ hash.
  One cache miss per live conversation, on its next turn only.
- hermes status counts more sessions. A declared API conversation is now
  recorded as a keyed row and appears in "Active: N session(s)" where it was
  invisible. Those sessions already existed; only their visibility changes.
- A database upgraded mid-conversation starts with no generation and takes its
  first from the next boundary written, so a conversation that reset before the
  upgrade shares its predecessor's scope once. One warm bucket, never a crossed
  identity.

Verified on this head: 55 in test_declared_conversation_scope.py (51 + 4 new),
33 in test_prompt_cache_scope.py, 49 in test_api_server_declared_conversation.py,
25 in test_api_server_runs.py, 109 in test_api_server.py, 12 in
test_cross_process_turn_lease.py, and 526 across test_hermes_state.py +
tests/hermes_state/ + tests/state/. ruff clean.

Found in review by @teknium1.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 832d68aba4 fix(cache): repair settlement, and make the generation unprunable
Four blockers from @andrexibiza's reviews of 28a2d7f0ee and dc7865765c. The
first two are defects I introduced in 99f2d4394f by replacing the wrong
occurrence of an identical call site.

1. _run_agent raised NameError on every opted-in declared bind. Its worker
   finally evaluated `if _declared_selected:`, a local of _handle_responses /
   _handle_runs that is neither a parameter nor an enclosing binding here, so
   the successful declared-key paths failed at settlement after the agent run.
   bind_declared_conversation already IS the gate; the inner name is gone.

2. /v1/runs never received the gate at all -- it landed on _run_agent instead.
   _run_sync bound unconditionally, so an explicit body session_id that existed
   with an empty session_key was adopted by the header key even though the
   header lost precedence. It now carries the same gate.

3. COUNT(*) + MAX(ended_at) over session rows cannot prove non-reuse.
   delete_session() deletes the selected row and bulk prune selects ended rows,
   so the aggregate can return a pair it already emitted:
   (1,T1) -> (2,T2) -> delete boundary B -> (1,T1), handing a new conversation
   a retired affinity identity. The backwards-clock shape needs no pruning at
   all. The generation now lives in a conversation_generations table keyed by
   (source, session_key), advanced by _bump_conversation_generation inside the
   same transaction that writes each boundary -- outside prunable session
   history, wall-clock-free, and increment-only. end_session() and
   promote_to_session_reset() both advance it, and only when they actually
   wrote a boundary, so a repeated end cannot double-count.

4. The carrier could be memoized under the wrong source. _agent_source() fell
   back to agent.platform before the row landed while persistence uses
   _session_source_for_agent(), which honors HERMES_SESSION_SOURCE. Because a
   declared scope is non-None immediately, resolve_prompt_cache_scope memoizes
   it and never re-resolves once the authoritative row appears, so under an
   override both sides of a /new read the platform domain and hashed the same
   scope. The pre-row path now uses the persistence resolver itself.

Coverage answers the review's specific objection that mocked tests proved the
mock rather than the path. TestRealRunAgentSettlement stubs _create_agent and
lets the real _run_agent settle; the /v1/runs case persists an unkeyed explicit
row and waits for the worker to retire before asserting. Both were verified by
mutation: reinstating the inner name fails two of them, and removing the
/v1/runs gate fails the explicit-session one. The first version of that test
passed with the gate removed -- it asserted before settlement -- and would have
been the same empty proof the review called out.

TestGenerationSurvivesPruning covers deleting the newest boundary, deleting
every boundary, the backwards-clock-then-prune shape, compression and
accidental ends not advancing it, repeated ends not double-counting, promotion
advancing it, unkeyed rows advancing nothing, and peer scoping.
TestSourceOverrideDomain covers the override across a reset.

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Refs #96811

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
joaomarcos d63e5d8a10 fix(cache): source-qualify the peer identity and gate the declared bind
Both blockers from @andrexibiza's review of 28a2d7f0ee.

1. The generation lookup was not in the same identity domain as recovery.
   latest_conversation_boundary() selected on session_key alone, while
   _declared_conversation_session() is qualified by (source, session_key).
   X-Hermes-Session-Key accepts any authenticated caller-supplied string, so an
   API conversation may legally carry the same key as a Telegram row in one
   database -- a /new over there rotated this conversation's gwk_ generation
   while recovery correctly refused to cross the same line, moving the affinity
   identity out from under a physical identity that had not moved.

   The boundary read now takes (session_key, source), and the carrier is
   'source|key|generation' rather than 'key|generation' -- keying on the string
   alone would also collapse two same-key conversations from different sources
   onto one routing key, since this value leaves the process verbatim as
   OpenRouter's sticky session_id and xAI's x-grok-conv-id. The source comes
   from the agent's own session row, falling back to the platform the row will
   be created with before it lands.

2. The declared key's stated lower precedence did not survive settlement. Both
   handlers let stored_session_id / an explicit body session_id win, then called
   _bind_declared_conversation() unconditionally. record_gateway_session_peer()
   does SET session_key = ? across compression ancestors, so a request carrying
   conversation A's chain plus header key B silently rebound A to B: A could no
   longer be recovered by its own key, and B recovered A's session.

   Recording is now gated on the declared key having actually selected or
   minted the session, on both paths. Behind that gate the bind itself refuses
   to overwrite a row already bound to a different key, so a future caller
   cannot reintroduce the same defect by opting in wrongly.

test_declaration_outranks_the_lineage_root asserted the pre-qualification
contract by comparing a DB-backed agent against a DB-less one; it now makes the
stronger statement it was written for -- one declared conversation reached
through two different physical ids on the same peer.

Refs #96811

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
joaomarcos e5bce4df4b fix(cache): make the conversation generation survive a backwards clock
Self-review of the generation marker. MAX(ended_at) alone is wall-clock: an
NTP correction between two resets writes a SMALLER boundary, MAX keeps
returning the older one, and the next conversation silently reuses the
previous generation -- two conversations on one routing key, which is the
defect this PR exists to remove.

latest_conversation_boundary now returns (count, latest_ended_at) and the
marker is 'count:ended_at'. The two halves fail under different conditions --
a backwards clock defeats the timestamp, retention pruning of an old ended row
decrements the count -- so a generation repeats only if both happen at once.
The pair is deliberately biased toward changing: a spurious change costs one
cold prompt-cache bucket, a repeat would merge two conversations.

Pinned by test_a_backwards_clock_does_not_reuse_a_generation, which rewrites
the second boundary to land before the first and asserts three conversations
still resolve to three distinct scopes.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos d7995bffaf fix(cache): qualify the declared key with the conversation generation
The declared key is a per-CHAT identifier and outlives the conversation it
names: reset_session() mints a fresh physical id on /new but keeps the key, and
the idle/daily/suspended policy resets do the same. Hashing the key alone
therefore mapped the conversation before a reset and the one after it onto one
gwk_ scope -- the lifecycle violation @cervantesh raised on #97158 and
@kshitijk4poor reproduced on #97709.

No counter is introduced. The generation that must rotate is already durable:
every one of those boundaries closes the outgoing row with an
_RESET_END_REASONS end_reason, so SessionDB.latest_conversation_boundary reads
the most recent one and declared_conversation_scope hashes 'key|generation'.

That makes the carrier stable across a host's per-response physical ids -- a
host that never resets writes no boundary, so every reply hashes the same value
-- while rotating on every conversation replacement, /new and the policy
auto-resets alike. ended_at only moves forward, so a retired generation can
never be reused: no ABA.

It also cannot drift from the rest of the codebase's notion of a conversation
boundary, because find_latest_gateway_session_for_peer fences on the same set.

The read is on the memoized resolution path, not per API call, and both lookups
fail closed: an unqualified key would span a /new, so a DB error degrades to the
physical-id scope. A SessionDB without the lookup keeps the previous behaviour.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 65672e3a93 fix(cache): honor the host-declared conversation key on the affinity-key path
Every conversation-affinity hint Hermes sends is derived from the PHYSICAL
session id: prompt_cache_key on both OpenAI-wire transports, OpenRouter's and
Nous Portal's sticky session_id, and xAI's x-grok-conv-id. A host that mints
one physical session per RESPONSE re-keys all four on every reply, so the
conversation never lands back on the routing bucket it just warmed (#96811).

Two hosts do exactly that. Hermes Studio's group chat mints
gc_run_<room>_<profile>_<name>_<uuid4hex> per reply and destroys it after,
and POST /v1/responses with client-managed history mints str(uuid4()) per
request — while parsing X-Hermes-Session-Key one screen earlier and handing
it to the agent.

Hermes must not infer the logical conversation from the id's syntax: that
rule merges independent client-supplied ids and Studio members truncated past
its 96-character boundary (the #79017 failure class). It does not have to.
gateway_session_key is already the "stable per-chat key" built by
gateway.session.build_session_key from that header, and branching
deliberately does not key off it. The affinity path simply never consulted it.

- agent/prompt_cache_scope.py: declared_conversation_scope() resolves the key
  into gwk_<sha256[:24]> and outranks the lineage walk (it is stable across
  rotation AND across per-response ids). Hashed because, unlike a session id,
  the key embeds platform/chat/user identifiers and leaves the process
  verbatim as a sticky id and as x-grok-conv-id.
- agent/portal_tags.py: a separate ambient scope for ROUTING, published only
  when a host declared one. The providers read the attribution id when it is
  unset, so delegate trees keep sharing their parent's sticky key and every
  host that keeps one id per conversation is byte-identical to before.
- hermes_state.py: is_explicit_fork_child() — the public view of the marker
  rules that keep /branch children, delegate subagents and tool children off
  their parent's chat key. Background-review forks clone the live runtime, so
  _persist_disabled excludes them for the same reason (#79161).

Refs #96570
Fixes #96811
2026-09-01 02:14:35 -07:00
Brooklyn Nicholson adb23c13cb fix(bot-mode): keep Bot Chat resume on a proven compression tip
Desktop opens the registry id, then session.resume walked the legacy
unmarked-child fallback, so Open Chat still landed in a side chat after
the title lookup was already strict. Recoverable-archive resurrection
uses the same helper.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
2026-08-31 20:47:53 -05:00
Teknium 8dbf07e950 fix: satisfy the no-locked-pure-readers gate and close the DB before rmtree
_record_db_file_identity's PRAGMA fallback is a pure read — route it
through _read_ctx() instead of the writer lock (Pattern C gate).
test_codex_turn_persists_each_message_exactly_once leaked a live
SessionDB into shutil.rmtree, racing the WAL sidecars ('Directory not
empty' on CI); close the handle first and rmtree with ignore_errors.
2026-08-31 14:02:55 -07:00
Teknium a9b6b979e9 fix: take the writer lock inside the file-identity stamp helpers
The repo-wide conn-lock audit (tests/test_hermes_state_conn_lock_audit.py)
requires every self._conn statement outside construction to hold
self._lock — an unlocked statement races SessionDB.close() inside
pysqlite's statement cache (#99349). Wrap _ensure_db_file_generation's
statements and _record_db_file_identity's PRAGMA fallback in the lock;
both call sites run outside the lock, so no re-entry.
2026-08-31 14:02:55 -07:00
rainbowgits 71256dfd01 fix(state): fail loudly when state.db is replaced under a live process
Detect same-inode cp via a generation stamp, halt FTS repair, and divert
unwritten transcripts to sessions/<id>.jsonl plus the gateway pending spool.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-31 14:02:55 -07:00
Shannon Sands 5a264f9a58 docs+test: pin lease worst-case trade-off and cover all arm sites (OOF-298 review follow-ups)
Addresses Enough1122's two non-blocking review notes on PR #92316:

1. Lease vs deadline arithmetic: the state.db init/migration/repair leases
   (600-900s) are authoritative against the 300s default deadline by design.
   Add explicit comments at all three lease sites pinning the trade-off:
   single lease is deliberate (clamped to _MAX_LEASE_S=900), honest worst
   case is up to the lease duration of zombie time on a wedged DB phase,
   accepted over per-chunk renewal complexity in the migration loops.

2. Arm-site coverage: add a structural contract test asserting every
   documented entry point (hermes_cli/main.py argv fast-path,
   hermes_cli/gateway.py config-bridge re-arm, gateway/run.py backstop,
   cli.py legacy --gateway) actually calls arm_startup_watchdog (or its
   aliased import), so a future entry point can't silently ship unwatched.
2026-08-31 14:01:39 -07:00
Shannon Sands 852db61abe fix(startup-watchdog): bounded hard-exit escort + phase-owned progress leases
Addresses the two class-level review blockers on PR #89750:

1. Bounded hard-exit seam (escort thread). The forensic fire path
   (logger.critical, dump record, faulthandler, lifecycle ledger) can
   itself wedge — the parked main thread may hold the logging handler
   lock, or the disk may be full/hung. _fire() now starts an exit-escort
   daemon thread BEFORE any forensics; it is free of log handlers,
   filesystem access, module loads and application locks, and hard-exits
   with the restart code after _FIRE_EXIT_BOUND_S unless the normal fire
   path signals completion. Adversarial tests hold the logging handler
   lock / hang the dump write at fire time and assert the exit seam is
   still reached.

2. Phase-owned progress leases (report_startup_progress). Process CPU
   time proves process activity, not startup progress: an unrelated busy
   thread could extend forever while startup sits parked (false
   negative), and I/O-bound repair/backup accrues ~zero CPU and would be
   killed (false positive). Long synchronous startup phases now declare
   authoritative, clamped (_MAX_LEASE_S), renewable progress leases:
   state.db _init_schema + the version-gated data-migration chain
   (hermes_state_schema) and repair_state_db_schema (hermes_state) are
   wired. CPU progress remains only as a bounded fallback, capped at
   _MAX_CPU_EXTENSIONS, with leases outranking the cap. Adversarial
   tests cover both directions (lease saves zero-CPU legitimate work;
   capped CPU noise no longer hides a parked deadlock).

Fire-path dump record now includes lease_count/last_lease_phase for
forensics. gateway/startup_watchdog.py shim re-exports
report_startup_progress.

OOF-298
2026-08-31 14:01:39 -07:00
the3asic 18ac3c4fb6 fix(state): defer corrupt FTS rebuilds past live operations 2026-08-31 12:08:30 -07:00
Pavel Diatchenko 96739033c4 fix(state): fail closed on unscoped corruption 2026-08-31 11:42:23 -07:00
joaomarcos db88b6c97c refactor(state): drop unused is_explicit_fork_child wrapper
teknium flagged it as scope creep on PR #98691 review: no callers in the diff.
_is_explicit_fork_child_row (the row-based helper actually used) is unchanged.
2026-08-31 11:23:48 -07:00