On a fresh boot with no resume_pending sessions, _finish_startup_restore
opened the inbound gate almost immediately while the agent-side turn
machinery (run_agent import graph, tool schemas + check_fn probes,
context-file tier) was still cold. A message arriving in that window was
served with a skeleton system prompt (~1.7K tokens vs ~14.6K healthy):
no AGENTS.md/context tier, no tool schemas, memory provider initializing
mid-turn.
Fix: start a background turn-machinery warm-up when the startup gate
closes (overlapping the network-bound platform connects) and have
_finish_startup_restore await it — BOUNDED by
agent.gateway_startup_warmup_timeout (default 20s, 0 disables) — before
draining the queue and opening the gate. On timeout the gate opens
anyway and the warm-up finishes in the background, so a wedged init can
never make the gateway permanently unavailable.
Reported by @yhfmstr in #99373.
Fixes#99373
TUI server shutdown stamps ended_at/end_reason='tui_shutdown' on sessions
whose agent keeps running; every rotation then aborts at
publish_compression_child's liveness check forever (the #88197 wedge; the
amplification half was fixed by #88411).
Class fix: is_automatic_end_reason() in hermes_state_common owns the
"accidental infrastructure cleanup vs deliberate boundary" taxonomy.
publish_compression_child clears automatic stamps in its own transaction
and proceeds (parent re-closes with its TRUE boundary,
end_reason='compression'); the #88411 pre-flush guard no longer aborts on
stamps the publish can heal. Deliberate boundaries (compression,
session_reset, explicit close) still fail closed at both sites.
TEST REPIN (deliberate contract change):
test_ended_parent_aborts_before_the_prepublish_flush pinned
"tui_shutdown stamp => rotation aborts and parent must not grow" — the
abort it required IS the #88197 wedge. Repinned as two tests:
- test_automatic_stamp_no_longer_wedges_rotation: automatic stamp =>
rotation COMMITS (no abort loop, so no growth-by-abort is possible);
- test_deliberately_ended_parent_aborts_before_the_prepublish_flush:
session_reset (deliberate boundary) => still aborts BEFORE the #47202
flush, preserving #88411's no-growth contract where an abort remains
correct.
The class invariant "no aborted rotation grows the parent" holds
everywhere: automatic stamps no longer produce aborts, deliberate
boundaries still abort pre-flush.
Hardening on top of @Soju06's forwarding fix: v2 providers written against
the original docs example (def on_pre_compress(self, messages)) must not
TypeError when the host forwards require_checkpoint — inspect the signature
and fall back to the legacy call shape. Docs example updated to advertise
the keyword.
MemoryManager.on_pre_compress() detects checkpoint API v2 providers,
selects the normalized evidence list for them, and re-raises their
failures under require_checkpoint — but it never tells the provider
that a checkpoint is required: the call passes only the messages.
A v2 provider therefore runs in its default best-effort mode, swallows
durable-write failures, and returns normally; the host then treats the
checkpoint as succeeded and lossy compression proceeds. With
compression.checkpoint_required: true this silently defeats the
guarantee the option exists to provide.
Forward require_checkpoint only to providers advertising the requested
checkpoint API version. Legacy providers keep the strict one-argument
on_pre_compress(self, messages) contract, so bundled v1 providers
(honcho, mem0, supermemory, ...) are unaffected.
Regression tests cover required and best-effort signaling, legacy
signature compatibility, and required-mode failure propagation.
Prove a fence-cancelled helper (no abort flag) persists cooldown so the
next turn does not re-arm compression, the host does not wait out the
600s ceiling after cancel, a held lock skips the sibling agent, and
unwind cancellation records the same brake.
(cherry picked from commit d2e178cd96fbcc8b2baf68488a8f46c70bff31a2)
A /stop or /restart abort left hygiene with no cooldown, so the next turn
re-armed auto-compression and waited up to 600s behind a fence that would
refuse the commit again. Record a cooldown on fence-cancel and unwind,
stop extending that wait once the fence is cancelled, and skip a new
hygiene agent while a compression lock is already held.
Pin gateway re-tagging of idle/preflight lifecycle lines as compacting,
and assert the TUI keeps that status until compacted rather than
restoring the busy bar after 4s.
Idle and preflight compaction arrived as lifecycle status without the
"Compacting context" marker, so TUI never entered a compacting state.
Re-tag those lines and freeze the busy FaceTicker on "compacting" for
the whole pause instead of restoring "running…" after 4s.
Re-applied onto 3aee29089 after `hermes update` reset main to origin/main.
1. web_routers/sessions.py: _resolve_session_id() classifies malformed-DB
errors via the existing is_malformed_db_error() and raises 503 at all five
call sites. delete_session_endpoint was the worst — an unresolvable id
counted as idempotent success, so DELETE reported it had removed a session
that was still on disk.
2. gateway/lifecycle_ledger.py: check_state_db_integrity() runs PRAGMA
quick_check(1) on the unclean-exit path only (~2s on 500MB) and records the
verdict into gateway-exit-diag.log. The 2026-08-31 corruption sat undetected
for 3.5 days because nothing ever looked.
3. hermes_cli/gateway.py: `gateway run --replace` gave the outgoing gateway 5s
before SIGKILL; SessionDB.close() runs a PASSIVE WAL checkpoint that does
not finish in 5s on a WAL 4x past the autocheckpoint threshold, and a kill
mid-checkpoint tears b-tree pages. Grace raised to 30s via a testable
_await_gateway_exit() that also re-checks after the final sleep (a PID
exiting in the last interval must not be SIGKILLed — PID-reuse hazard).
NOT added: wal_checkpoint(TRUNCATE) at shutdown — removed upstream in #45383
because a TRUNCATE reset races the live writer and tears b-tree pages.
Adversarial review: Codex gpt-5.6-sol, 9.0/10 across three groups, no must-fix.
The salvaged #98948 change returned False for every db=None, which also fired in deliberately store-less/degraded contexts (no _db_error), regressing six prompt.submit tests. Gate the loud failure on _db_error being set — the actual #98924 symptom — and keep the pinned best-effort contract otherwise.
Follow-ups on the salvage: regular-file guard before the zeroed byte-probe (a FIFO at the state.db path would block startup forever — #98017 review P2), plus an on-main-reproducing UnicodeDecodeError fixture for #98924 (raw bytes in sqlite_master, not messages.content, are what reach pysqlite error-message decode).
Companion to #98935, which fixes _fts_table_probe itself. This covers the
surfaces that PR does not touch:
- web_server._open_session_db_at_path: the one-writable-open heal only
caught sqlite3.DatabaseError; a raw UnicodeDecodeError (pysqlite failing
to decode SQLite's own error message over corrupt file bytes) bypassed
it, so the heal documented for malformed schema never fired (#98924
Failure 1). Both catches widened; decode errors dispatch to the heal.
- SessionSchemaMixin._recover_stale_fts_locked: drop-and-recreate skipped
vtables whose probe raised UnicodeDecodeError, the same too-narrow
catch the issue identified in the probe.
- TUI gateway: _ensure_session_db_row returned silently when the store
could not open, so prompt.submit streamed the turn while persisting
nothing (#98924 Failure 2). It now returns False and prompt.submit
fails the RPC with code 5072 so desktop maps it to a toast, mirroring
the disk-full/5070 convention. session.create stays silent per its
pinned degraded-mode contract.
The tokenizer-not-loaded branch ran its sqlite_master presence check
and self-heal statements (drop stale cjk triggers) with no guard, so a
transient sqlite3.OperationalError there (e.g. a locked database)
escaped the method despite its documented "Never raises" contract.
That exception then hit _migrate_broad_fts_update_triggers's
quarantine-then-reraise handler, aborting _init_schema and the whole
SessionDB open. Wrap the presence-check/self-heal block in the same
kind of OperationalError guard the tokenizer-loaded branch already
has, degrading to "no cjk index" instead of propagating.
Invalid UTF-8 bytes in messages.content (e.g. 0x81 from hardware issues,
corrupted disk I/O, or manual DB edits) caused read-only SessionDB init
to die on a bare UnicodeDecodeError in _fts_table_probe, taking down
every read endpoint (GET /api/sessions, Desktop read-only opens of other
profiles' DBs). The probe caught only sqlite3.OperationalError and would
re-raise any other exception, including UnicodeDecodeError (a ValueError,
not an sqlite3.Error subclass).
On some Python/SQLite builds the decode failure surfaces as
UnicodeDecodeError; on others as OperationalError('Could not decode to
UTF-8 column ...'). The fix catches both and treats them the same:
the FTS index is degraded (search may return less or fail), but the store
itself stays accessible for writes and non-FTS reads. Writable init
schedules a rebuild or degrades to LIKE search until repaired.
Adds test_98924_readonly_fts_decode_error.py with a regression test that
injects invalid UTF-8 via CAST(x'...' AS TEXT) through the Python sqlite3
module, triggers an FTS rebuild, then confirms that read-only init succeeds
instead of raising.
- Guard against concurrent-opener race where newly created 0-byte state.db was falsely quarantined before first schema write
- Wrap startup in quarantine_cross_process_lock when database is uninitialized or zeroed
- Guard is_zeroed_sqlite_file and is_zeroed_state_db against active live connections in current process
- Add concurrent-opener and live-connection regression tests
Salvaged from PR #90734 by @Kyzcreig onto current main:
- hermes_state_search.py: post-commit FTS incremental merge failures
(including the bare SystemError CPython's sqlite3 layer raises under
cross-thread errmsg scrambling) are contained and logged instead of
escaping and making the caller replay an ambiguous, possibly-durable
write (exactly-once refinement by @yuzilongleif-collab).
- tests/state/test_writer_conn_thread_safety.py: live reader/writer race
hammer + AST sweep freezing the no-unlocked-writer-conn invariant.
On top: the sweep now also flags self._conn PASSED to helpers, which
caught one more live site on main — get_session_delete_targets handed
the shared writer connection to _collect_delegate_child_ids inside a
_read_ctx block, executing on it without self._lock. Routed to the
borrowed read connection.
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
Salvaged from PR #99435. Fails on the next lock-free self._conn.<method>()
call site added to hermes_state.py; reads belong in _read_ctx(), writes
under self._lock (#99349).
Co-authored-by: dimitrysuen <dimitrysuen@users.noreply.github.com>
- fresh conversation subscribes with no since floor + bounded limit; seeded
channels keep since=last_ts-1 (#78429)
- all three production CLOSED membership phrasings prune per-subscription
without reconnect (#76850 + #97502)
- restricted channels are not re-adopted by discovery sweeps
- WS periodic discovery subscribes conversations found without a membership
event (#93557) and the companion task dies with its connection
- watch-all live channel adoption vs explicit-list scoping (#75107)
All six new tests sabotage-checked: each fails when its fix is reverted.
Three sibling gaps in the WebSocket transport's conversation lifecycle:
- #78429: _send_channel_subscription defaulted a zero last_ts to
'since ~ now', so the message that CREATED a new conversation (created_at
fractionally before the subscription) was never delivered. A channel with
no high-water mark now subscribes from the beginning with
limit=_FETCH_LIMIT instead; seeded channels still resume from last_ts-1.
- #93557: relays do not guarantee a kind-44100 membership event per new
conversation, so WS-transport deployments never discovered DMs opened
mid-session until a reconnect. The WS loop now runs the same
_discover_dms sweep the poll transport uses, on the same cadence
(poll_interval * _DM_DISCOVERY_EVERY), via a companion task that is
cancelled with the connection.
- #75107: _discover_dms only ever adopted DM-shaped conversations, so a
real community channel the agent joined mid-run was never subscribed
until restart. In watch-all mode (no explicit channels list) newly
listed real channels are now adopted and seeded from their newest events
(history predating the join is not replayed). Explicit watch lists stay
authoritative.
- Widen the permanent-rejection match to the exact relay phrasings seen in
production (#97502): 'not a channel member' and 'auth-required', alongside
'restricted'.
- Close the re-adoption hole called out in review: _discover_dms() (both the
dms-list path and the channels-list fallback) now skips channels in
_restricted_channels, so a restricted channel dropped at runtime cannot be
silently re-added by the next discovery sweep and re-trigger the rejection.
- Credit: runtime CLOSED matching terms from PR #97502 by @repfigit; the
per-subscription drop + restricted set is PR #76850 by @xozai.
When a Buzz relay sends a CLOSED frame for a single subscription with a
'restricted: not a channel member' error, the adapter was raising
ConnectionError, tearing down the entire WebSocket connection, and
immediately reconnecting — causing a ~1.6 s flood in gateway.log.
Root cause: the CLOSED handler unconditionally raised ConnectionError
regardless of whether the error was permanent (restricted) or transient
(e.g. server shutdown).
Fix:
- On a 'restricted' CLOSED, drop only the offending subscription and
record the channel in a new _restricted_channels set instead of
tearing down the whole connection.
- Skip restricted channels during connect() seeding and
_subscribe_websocket() so reconnects don't re-trigger the same error.
- Non-restricted CLOSED frames still raise ConnectionError and reconnect
as before.
Adds three regression tests:
- test_websocket_loop_drops_restricted_channel_without_reconnect
- test_websocket_loop_reconnects_on_non_restricted_closed
- test_restricted_channels_skipped_during_subscribe
Tested on macOS against buzz.xozai.com: gateway.log shows zero
'restricted' errors and stable 'watching N channel(s) via websocket'
after the fix.
`connect()` calls `_seed_channel()` unconditionally, and seeding marks every
event currently in the channel as seen so a start never replays history at the
agent. A message that arrives after the process starts but before the seed
completes — or at any point while the gateway is down — sits in exactly that
history, so the seed swallows it permanently even though the Buzz relay still
has it. The `seen` set and `last_ts` lived only in memory, so there was nothing
to distinguish "already handled" from "never seen".
Each watched channel's cursor (`chat_type`, `last_ts`, and the bounded `seen`
id list) is now persisted under `HERMES_HOME/buzz/channel-cursors.json` and
restored at connect. Where a cursor exists the channel resumes from it and the
history fetch is skipped entirely; where none exists the old seed-from-history
behaviour is unchanged, so a first-ever run still never replays a backlog.
Details worth noting:
- The file records the identity and relay it was written for. A cursor from a
different bot or relay is ignored rather than trusted — the channel ids
would collide while the event stream behind them is a different one.
- Any read or parse failure leaves the cursors empty, which degrades to
seeding instead of failing the connect. Writes go through
`utils.atomic_json_write` (temp + fsync + replace), so a crash mid-write
cannot leave a truncated cursor behind.
- The restored `seen` list is trimmed to `_SEEN_CAP` on load, keeping the
newest ids, so a hand-edited or legacy file cannot grow the de-dupe set
without bound.
- Saves are gated on the cursor actually moving, so an idle channel does not
rewrite the file every poll interval. Both inbound transports are covered:
the poll sweep and the WebSocket event path share the same check.
Tests: six new cases in `TestChannelCursorPersistence` — the cursor is written
on seed, a restart resumes without spending a CLI call on history and then
delivers the mention that landed while the gateway was down, a foreign
identity or relay is ignored, a corrupt file falls back to seeding, the
restored `seen` set stays bounded, and an idle poll leaves the file untouched.
All six fail on main.
Tested on: Windows 11, Python 3.12. `python -m pytest
tests/gateway/test_buzz_adapter.py tests/gateway/test_buzz_websocket.py -q` —
33 passed (23 pre-existing + 6 new here, plus 4 WebSocket). Requires
`pytest-asyncio` (pinned at 1.3.0 in pyproject) — without it the async cases
in this file error out as unknown marks.
A relay-side close the transport never surfaces (observed as a CLOSE_WAIT
socket behind Cloudflare, #98097) parks the read loop forever while the
gateway keeps reporting connected: inbound stops, gateway_state.json stays
healthy, and only a restart recovers. The library keepalive should catch
this first, but as a last resort the read side now waits at most
_WS_READ_IDLE_TIMEOUT (300s) for a frame before raising into the existing
reconnect path, which re-authenticates and re-subscribes with per-channel
since filters intact.
Fixes#98097
- lightpanda_engine_status: check use_real_profile before the cloud
provider, matching browser_exec's actual resolution order (real-profile
resolution runs before backend resolution), so /browser status and
hermes doctor name the right shadowing setting when both are set.
- launch_lightpanda: drop the unreachable Windows popen_kwargs branch
(find_lightpanda_binary returns None on nt, launch errors out earlier).
- doctor: drop the over-defensive try/except around the cached
_using_lightpanda_engine() config read.
- Docstring: 'no-I/O gates' -> 'no network I/O (config reads only)'.
- New test pinning real-profile-over-cloud-provider reason precedence.
Browser Use mode never read browser.engine: _resolve_backend_cdp() went
BU_CDP_* env -> CDP override -> cloud provider -> local Chrome, so
`engine: lightpanda` was a silent no-op on the default backend, and on
the built-in path it was skipped whenever a cloud provider, Camofox or a
CDP override was active without anyone saying so.
- browser_use_cli: when the engine is lightpanda and nothing with higher
precedence claimed the session, get a session from _get_session_info()
and export its endpoint as BU_CDP_URL; the browser is private to the
session key, so the own-tab preamble is skipped. The browser_exec
description gains a Lightpanda header (text-first, new_tab once then
goto_url — lightpanda-io/browser#1962).
- browser_tool: _create_local_session() spawns `lightpanda serve
--host 127.0.0.1 --port <free>` per session key (new
tools/browser_lightpanda.py), reusing the session cache, inactivity
reaper and atexit cleanup; a dead process is respawned on the next call;
orphans from a crashed Hermes are reaped through per-process records in
$HERMES_HOME/cache/browser-use/lightpanda/. New lightpanda_engine_status()
reports whether the engine is in effect or what shadows it.
- tools_config: "Lightpanda" row in the Browser Automation picker
(cloud_provider: local + engine: lightpanda; "Local Browser" resets the
engine to auto) with a binary-check post-setup.
- /browser status and hermes doctor print the engine state and, when it
is shadowed, the reason.
Follow-up for salvaged PRs #82646 + #83414: resolution probes precede
publishes, the presentation-escape retry precedes the self-mention
downgrade, and a new test pins the escape-retry-delivers path.
Salvaged from PR #83414 (4 commits squashed to final state) and composed
with the presentation-mention escape retry from PR #82646 already on this
branch: send() now resolves @Name tokens to channel-member pubkeys
(membership-accurate via `channels members`, TTL-cached, Unicode token
boundaries, ambiguous names stay presentation-only) and passes explicit
--mention args; recovery ladder handles membership drift, unresolvable
prose @tokens (escape retry, #78797), and a final self-mention downgrade.
require_mention gated only on visible text, so Desktop thread replies
(e.g. /approve session) to the agent's own prompts were dropped with no
log. Cache event_id→(author, snippet) from seed/poll/WS/send, resolve
the direct e-tag parent, and dispatch when that parent is ours; also
populate reply_to_* on MessageEvent for gateway context injection.
Fixes#75826
The inbound path hardcoded Nostr kind 9 at both the WebSocket
subscription filter and the dispatch gate (which runs before mention
gating), so Buzz forum channels — kind 45001 thread roots and 45003
comment replies — were silently never dispatched to the agent; chat and
stream channels worked, making the gap invisible (#90309). Block's own
ACP harness documents the forum kinds explicitly.
Introduce _DISPATCH_KINDS = {9, 45001, 45003} for the subscription
filter and dispatch gate. The stream kinds (46010/40007/45002) stay out
of scope until their dispatch semantics are confirmed.
_is_direct_message_event deliberately keeps its kind-9-only check:
widening it would let a p-tagged forum post be reclassified as a DM and
bypass mention gating. The send path already works unchanged (send()
omits --kind and threads via --reply-to).
Fixes#90309
The fallback URL probe now uses 'get url' ({data:{url}}), not
'eval window.location.href' ({data:{result}}); the npx-resolution test's
mocked response predates the #81673 change.
Review follow-up for the #81673 salvage: the strip was silent, which
would confuse a user whose AGENT_BROWSER_ARGS applies to Chrome commands
but vanishes on Lightpanda ones.
Strip Chromium-only launch variables from Lightpanda commands, use a non-recursive Lightpanda URL lookup for Chrome fallback, and share sandbox argument injection with the temporary Chrome path.
Co-authored-by: forjd-hermes-bot <282037251+forjd-hermes-bot@users.noreply.github.com>
Follow-up to the salvaged manifest guard (#90859): the doctor's temporary
HERMES_HOME now enters the ExitStack before the staging copytree, so
ENOSPC / KeyboardInterrupt / any exception during staging deterministically
removes the hermes-plugin-doctor-* directory instead of relying on the
TemporaryDirectory GC finalizer (which cannot run while the traceback pins
the frame). Regression test proven via sabotage run against the old code.