A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
The 19:53 canary run failed the in_channel seed EVEN WITH the deterministic
session_id fix live, and the mirror's two failure paths (no-session bail,
append exception) both logged at debug — invisible in production. WARNING
both, with the explicit session id in the exception path, so the next run
names the failing branch instead of another blind retest.
Two live failures from the Alice canary (2026-08-19, jobs 28a24afebd81 /
83b93f8be379), both leaving a continuable in_channel cron with amnesia:
1. Seed mirrored via origin heuristics and silently dropped the brief.
_seed_cron_channel_session created the flat session row, then
mirror_to_session RE-DISCOVERED the target via find_session_by_origin —
whose multi-candidate bail-out returns None on a populated chat (flat
session + N per-message thread sessions sharing one chat_id, mixed
user_ids). Receipt: 'in_channel seed did NOT land on slack:D0BJTDCSR7C'.
mirror_to_session now accepts an explicit session_id and both cron seeds
pass the exact row they just created; origin-scan remains the fallback
for callers that genuinely don't know the target.
2. The brief's OWN THREAD was never seeded. in_channel delivers flat, but
a flat Slack message still invites a thread reply (the natural mobile
affordance — exactly what the user did). That reply keys to
(chat, thread=<brief ts>), which no seed touched. The delivery's
message_id now anchors a companion _seed_cron_thread_session so BOTH
reply surfaces (plain channel message AND in-thread reply) continue the
job.
Also: thread-seed failures upgraded debug→WARNING (silent seed failure IS
the user-facing bug), and the thread seed reports landed/not-landed.
Regression tests drive both against the live failure shapes: exact-session
mirror asserted via session_id kwarg; thread companion asserted via the
SendResult message_id anchor. Clean-fixture blind spot noted: the E2E
harness used a fresh store with one row, which is why heuristic rediscovery
looked fine pre-production.
Moves gateway routing metadata (display_name, origin_json, expiry_finalized)
into state.db, making SQLite the single source of truth for gateway session
discovery. Eliminates the dual-file (sessions.json + state.db) polling
dependency that caused the mcp_serve new-conversation race (#8925).
- hermes_state.py: schema v18 (3 new sessions columns + sessions.json
backfill migration), record_gateway_session_peer gains
display_name/origin_json, new set_expiry_finalized(),
list_gateway_sessions(), find_session_by_origin()
- gateway/session.py: peer recorder persists display_name + full origin
JSON; new SessionStore.set_expiry_finalized() single write-path
- gateway/run.py: expiry watcher success + give-up paths use the store
helper so the flag lands in both sessions.json and state.db
- mcp_serve.py: routing index reads state.db first (sessions.json fallback
for pre-migration DBs); _poll_once collapses to a single state.db mtime
check — the #8925 race is structurally impossible now
- gateway/mirror.py, gateway/channel_directory.py, hermes_cli/status.py:
query state.db first, sessions.json fallback
Closes#9006
Addresses review on #51077 (kxee). The continuable-cron mirror reused
gateway.mirror.mirror_to_session, which writes role=assistant — re-
introducing the exact alternation violation #2313 (37a997945)
deliberately removed: a cron brief landing as assistant after the
agent's last turn yields assistant->assistant, which breaks strict-
alternation providers (OpenAI/OpenRouter) per issue #2221. The mirror/
mirror_source metadata is also dropped at the SQLite boundary, so the
[Delivered from cron] label is lost on replay.
This is an intentional, opt-in (default OFF) reversal of #2313's
'cron output does not belong in interactive history' for the reply-to-
cron use case — gated behind cron.mirror_delivery / attach_to_session.
Fixes:
- mirror_to_session gains a role param (default 'assistant' — interactive
send_message mirror unchanged, it IS the agent speaking). Cron paths
pass role='user' with a '[Cron delivery: <task>]' prefix so the brief
collapses via repair_message_sequence's consecutive-user merge on every
provider, and stays distinguishable on replay despite the metadata drop.
- thread_seeded: defer seeding + the flag until delivery into the new
thread actually succeeds. Previously set pre-delivery, so an open-
succeeds / deliver-fails case both stranded a seeded-but-unseen brief
AND suppressed the DM-fallback mirror.
- seed mirror now passes user_id='system:cron' to resolve the exact
thread-keyed session row it just created.
- dedupe the duplicate BasePlatformAdapter import in _deliver_result.
- trim oversized docstrings to non-obvious WHY (AGENTS.md).
- docs: document cron.mirror_delivery / attach_to_session in
website/docs/user-guide/features/cron.md.
- test: assert the cron mirror writes role='user' with the label prefix.
204 cron+mirror tests pass.
Users who inspect ~/.hermes/sessions/sessions.json see only gateway entries
(e.g. agent:main:whatsapp:dm:...) and mistake it for the session index that
hermes sessions list / /sessions read — which is actually state.db. Issue
#49361 reported CLI sessions as 'invisible' on this premise.
- gateway/session.py: write a self-documenting _README sentinel at the top of
sessions.json explaining it's the gateway routing index and that ALL sessions
(CLI/TUI/gateway) live in state.db; skip _-prefixed keys on load so the
sentinel never round-trips into a SessionEntry.
- Harden every sessions.json reader against the sentinel: mcp_serve loader,
gateway/mirror.py, gateway/channel_directory.py all skip _-prefixed keys.
- docs/user-guide/sessions.md: warning callout naming the exact symptom.
- tests: assert prune ignores metadata sentinels; add round-trip coverage.
Mirror messages are persisted via _append_to_sqlite. JSONL writer was
a redundant dual-write. Updated test assertions from JSONL file checks
to SQLite mock verification.
Authored by 0xbyt4. Fixes #N/A (no linked issue).
- Sanitize user input before FTS5 MATCH to prevent OperationalError on
special characters (C++, unbalanced quotes, dangling operators, etc.)
- Close SessionDB connection in mirror._append_to_sqlite() via finally block
- Added tests for both fixes
Two bugs fixed:
1. search_messages() crashes with OperationalError when user queries
contain FTS5 special characters (+, ", (, {, dangling AND/OR, etc).
Added _sanitize_fts5_query() to strip dangerous operators and a
fallback try-except for edge cases.
2. _append_to_sqlite() in mirror.py creates a new SessionDB per call
but never closes it, leaking SQLite connections. Added finally block
to ensure db.close() is always called.
On Windows, Python's open() defaults to the system locale encoding
(e.g. cp1254 for Turkish, cp1252 for Western European) instead of
UTF-8. The gateway already uses ensure_ascii=False in json.dumps()
to preserve Unicode characters in chat messages, but the
corresponding open() calls lack encoding="utf-8". This mismatch
causes UnicodeEncodeError / UnicodeDecodeError when users send
non-ASCII messages (Turkish, Japanese, Arabic, emoji, etc.) through
Telegram, Discord, WhatsApp, or Slack on Windows.
The project already fixed this for .env files in hermes_cli/config.py
(line 624) but the gateway module was missed.
Files fixed:
- gateway/session.py: session index + JSONL transcript read/write (5 calls)
- gateway/channel_directory.py: channel directory read/write (3 calls)
- gateway/mirror.py: session index read + transcript append (2 calls)
- Introduced a new channel directory to cache reachable channels/contacts for messaging platforms, enhancing the send_message tool's ability to resolve human-friendly names to numeric IDs.
- Added functionality to mirror sent messages into the target's session transcript, providing context for cross-platform message delivery.
- Updated the send_message tool to support listing available targets and improved error handling for channel resolution.
- Enhanced the gateway to build and refresh the channel directory during startup and at regular intervals, ensuring up-to-date channel information.