On a fresh boot with no resume_pending sessions, _finish_startup_restore
opened the inbound gate almost immediately while the agent-side turn
machinery (run_agent import graph, tool schemas + check_fn probes,
context-file tier) was still cold. A message arriving in that window was
served with a skeleton system prompt (~1.7K tokens vs ~14.6K healthy):
no AGENTS.md/context tier, no tool schemas, memory provider initializing
mid-turn.
Fix: start a background turn-machinery warm-up when the startup gate
closes (overlapping the network-bound platform connects) and have
_finish_startup_restore await it — BOUNDED by
agent.gateway_startup_warmup_timeout (default 20s, 0 disables) — before
draining the queue and opening the gate. On timeout the gate opens
anyway and the warm-up finishes in the background, so a wedged init can
never make the gateway permanently unavailable.
Reported by @yhfmstr in #99373.
Fixes#99373
TUI server shutdown stamps ended_at/end_reason='tui_shutdown' on sessions
whose agent keeps running; every rotation then aborts at
publish_compression_child's liveness check forever (the #88197 wedge; the
amplification half was fixed by #88411).
Class fix: is_automatic_end_reason() in hermes_state_common owns the
"accidental infrastructure cleanup vs deliberate boundary" taxonomy.
publish_compression_child clears automatic stamps in its own transaction
and proceeds (parent re-closes with its TRUE boundary,
end_reason='compression'); the #88411 pre-flush guard no longer aborts on
stamps the publish can heal. Deliberate boundaries (compression,
session_reset, explicit close) still fail closed at both sites.
TEST REPIN (deliberate contract change):
test_ended_parent_aborts_before_the_prepublish_flush pinned
"tui_shutdown stamp => rotation aborts and parent must not grow" — the
abort it required IS the #88197 wedge. Repinned as two tests:
- test_automatic_stamp_no_longer_wedges_rotation: automatic stamp =>
rotation COMMITS (no abort loop, so no growth-by-abort is possible);
- test_deliberately_ended_parent_aborts_before_the_prepublish_flush:
session_reset (deliberate boundary) => still aborts BEFORE the #47202
flush, preserving #88411's no-growth contract where an abort remains
correct.
The class invariant "no aborted rotation grows the parent" holds
everywhere: automatic stamps no longer produce aborts, deliberate
boundaries still abort pre-flush.
Hardening on top of @Soju06's forwarding fix: v2 providers written against
the original docs example (def on_pre_compress(self, messages)) must not
TypeError when the host forwards require_checkpoint — inspect the signature
and fall back to the legacy call shape. Docs example updated to advertise
the keyword.
MemoryManager.on_pre_compress() detects checkpoint API v2 providers,
selects the normalized evidence list for them, and re-raises their
failures under require_checkpoint — but it never tells the provider
that a checkpoint is required: the call passes only the messages.
A v2 provider therefore runs in its default best-effort mode, swallows
durable-write failures, and returns normally; the host then treats the
checkpoint as succeeded and lossy compression proceeds. With
compression.checkpoint_required: true this silently defeats the
guarantee the option exists to provide.
Forward require_checkpoint only to providers advertising the requested
checkpoint API version. Legacy providers keep the strict one-argument
on_pre_compress(self, messages) contract, so bundled v1 providers
(honcho, mem0, supermemory, ...) are unaffected.
Regression tests cover required and best-effort signaling, legacy
signature compatibility, and required-mode failure propagation.
Prove a fence-cancelled helper (no abort flag) persists cooldown so the
next turn does not re-arm compression, the host does not wait out the
600s ceiling after cancel, a held lock skips the sibling agent, and
unwind cancellation records the same brake.
(cherry picked from commit d2e178cd96fbcc8b2baf68488a8f46c70bff31a2)
Pin gateway re-tagging of idle/preflight lifecycle lines as compacting,
and assert the TUI keeps that status until compacted rather than
restoring the busy bar after 4s.
Re-applied onto 3aee29089 after `hermes update` reset main to origin/main.
1. web_routers/sessions.py: _resolve_session_id() classifies malformed-DB
errors via the existing is_malformed_db_error() and raises 503 at all five
call sites. delete_session_endpoint was the worst — an unresolvable id
counted as idempotent success, so DELETE reported it had removed a session
that was still on disk.
2. gateway/lifecycle_ledger.py: check_state_db_integrity() runs PRAGMA
quick_check(1) on the unclean-exit path only (~2s on 500MB) and records the
verdict into gateway-exit-diag.log. The 2026-08-31 corruption sat undetected
for 3.5 days because nothing ever looked.
3. hermes_cli/gateway.py: `gateway run --replace` gave the outgoing gateway 5s
before SIGKILL; SessionDB.close() runs a PASSIVE WAL checkpoint that does
not finish in 5s on a WAL 4x past the autocheckpoint threshold, and a kill
mid-checkpoint tears b-tree pages. Grace raised to 30s via a testable
_await_gateway_exit() that also re-checks after the final sleep (a PID
exiting in the last interval must not be SIGKILLed — PID-reuse hazard).
NOT added: wal_checkpoint(TRUNCATE) at shutdown — removed upstream in #45383
because a TRUNCATE reset races the live writer and tears b-tree pages.
Adversarial review: Codex gpt-5.6-sol, 9.0/10 across three groups, no must-fix.
Follow-ups on the salvage: regular-file guard before the zeroed byte-probe (a FIFO at the state.db path would block startup forever — #98017 review P2), plus an on-main-reproducing UnicodeDecodeError fixture for #98924 (raw bytes in sqlite_master, not messages.content, are what reach pysqlite error-message decode).
Companion to #98935, which fixes _fts_table_probe itself. This covers the
surfaces that PR does not touch:
- web_server._open_session_db_at_path: the one-writable-open heal only
caught sqlite3.DatabaseError; a raw UnicodeDecodeError (pysqlite failing
to decode SQLite's own error message over corrupt file bytes) bypassed
it, so the heal documented for malformed schema never fired (#98924
Failure 1). Both catches widened; decode errors dispatch to the heal.
- SessionSchemaMixin._recover_stale_fts_locked: drop-and-recreate skipped
vtables whose probe raised UnicodeDecodeError, the same too-narrow
catch the issue identified in the probe.
- TUI gateway: _ensure_session_db_row returned silently when the store
could not open, so prompt.submit streamed the turn while persisting
nothing (#98924 Failure 2). It now returns False and prompt.submit
fails the RPC with code 5072 so desktop maps it to a toast, mirroring
the disk-full/5070 convention. session.create stays silent per its
pinned degraded-mode contract.
The tokenizer-not-loaded branch ran its sqlite_master presence check
and self-heal statements (drop stale cjk triggers) with no guard, so a
transient sqlite3.OperationalError there (e.g. a locked database)
escaped the method despite its documented "Never raises" contract.
That exception then hit _migrate_broad_fts_update_triggers's
quarantine-then-reraise handler, aborting _init_schema and the whole
SessionDB open. Wrap the presence-check/self-heal block in the same
kind of OperationalError guard the tokenizer-loaded branch already
has, degrading to "no cjk index" instead of propagating.
Invalid UTF-8 bytes in messages.content (e.g. 0x81 from hardware issues,
corrupted disk I/O, or manual DB edits) caused read-only SessionDB init
to die on a bare UnicodeDecodeError in _fts_table_probe, taking down
every read endpoint (GET /api/sessions, Desktop read-only opens of other
profiles' DBs). The probe caught only sqlite3.OperationalError and would
re-raise any other exception, including UnicodeDecodeError (a ValueError,
not an sqlite3.Error subclass).
On some Python/SQLite builds the decode failure surfaces as
UnicodeDecodeError; on others as OperationalError('Could not decode to
UTF-8 column ...'). The fix catches both and treats them the same:
the FTS index is degraded (search may return less or fail), but the store
itself stays accessible for writes and non-FTS reads. Writable init
schedules a rebuild or degrades to LIKE search until repaired.
Adds test_98924_readonly_fts_decode_error.py with a regression test that
injects invalid UTF-8 via CAST(x'...' AS TEXT) through the Python sqlite3
module, triggers an FTS rebuild, then confirms that read-only init succeeds
instead of raising.
- Guard against concurrent-opener race where newly created 0-byte state.db was falsely quarantined before first schema write
- Wrap startup in quarantine_cross_process_lock when database is uninitialized or zeroed
- Guard is_zeroed_sqlite_file and is_zeroed_state_db against active live connections in current process
- Add concurrent-opener and live-connection regression tests
Salvaged from PR #90734 by @Kyzcreig onto current main:
- hermes_state_search.py: post-commit FTS incremental merge failures
(including the bare SystemError CPython's sqlite3 layer raises under
cross-thread errmsg scrambling) are contained and logged instead of
escaping and making the caller replay an ambiguous, possibly-durable
write (exactly-once refinement by @yuzilongleif-collab).
- tests/state/test_writer_conn_thread_safety.py: live reader/writer race
hammer + AST sweep freezing the no-unlocked-writer-conn invariant.
On top: the sweep now also flags self._conn PASSED to helpers, which
caught one more live site on main — get_session_delete_targets handed
the shared writer connection to _collect_delegate_child_ids inside a
_read_ctx block, executing on it without self._lock. Routed to the
borrowed read connection.
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
Salvaged from PR #99435. Fails on the next lock-free self._conn.<method>()
call site added to hermes_state.py; reads belong in _read_ctx(), writes
under self._lock (#99349).
Co-authored-by: dimitrysuen <dimitrysuen@users.noreply.github.com>
- fresh conversation subscribes with no since floor + bounded limit; seeded
channels keep since=last_ts-1 (#78429)
- all three production CLOSED membership phrasings prune per-subscription
without reconnect (#76850 + #97502)
- restricted channels are not re-adopted by discovery sweeps
- WS periodic discovery subscribes conversations found without a membership
event (#93557) and the companion task dies with its connection
- watch-all live channel adoption vs explicit-list scoping (#75107)
All six new tests sabotage-checked: each fails when its fix is reverted.
When a Buzz relay sends a CLOSED frame for a single subscription with a
'restricted: not a channel member' error, the adapter was raising
ConnectionError, tearing down the entire WebSocket connection, and
immediately reconnecting — causing a ~1.6 s flood in gateway.log.
Root cause: the CLOSED handler unconditionally raised ConnectionError
regardless of whether the error was permanent (restricted) or transient
(e.g. server shutdown).
Fix:
- On a 'restricted' CLOSED, drop only the offending subscription and
record the channel in a new _restricted_channels set instead of
tearing down the whole connection.
- Skip restricted channels during connect() seeding and
_subscribe_websocket() so reconnects don't re-trigger the same error.
- Non-restricted CLOSED frames still raise ConnectionError and reconnect
as before.
Adds three regression tests:
- test_websocket_loop_drops_restricted_channel_without_reconnect
- test_websocket_loop_reconnects_on_non_restricted_closed
- test_restricted_channels_skipped_during_subscribe
Tested on macOS against buzz.xozai.com: gateway.log shows zero
'restricted' errors and stable 'watching N channel(s) via websocket'
after the fix.
`connect()` calls `_seed_channel()` unconditionally, and seeding marks every
event currently in the channel as seen so a start never replays history at the
agent. A message that arrives after the process starts but before the seed
completes — or at any point while the gateway is down — sits in exactly that
history, so the seed swallows it permanently even though the Buzz relay still
has it. The `seen` set and `last_ts` lived only in memory, so there was nothing
to distinguish "already handled" from "never seen".
Each watched channel's cursor (`chat_type`, `last_ts`, and the bounded `seen`
id list) is now persisted under `HERMES_HOME/buzz/channel-cursors.json` and
restored at connect. Where a cursor exists the channel resumes from it and the
history fetch is skipped entirely; where none exists the old seed-from-history
behaviour is unchanged, so a first-ever run still never replays a backlog.
Details worth noting:
- The file records the identity and relay it was written for. A cursor from a
different bot or relay is ignored rather than trusted — the channel ids
would collide while the event stream behind them is a different one.
- Any read or parse failure leaves the cursors empty, which degrades to
seeding instead of failing the connect. Writes go through
`utils.atomic_json_write` (temp + fsync + replace), so a crash mid-write
cannot leave a truncated cursor behind.
- The restored `seen` list is trimmed to `_SEEN_CAP` on load, keeping the
newest ids, so a hand-edited or legacy file cannot grow the de-dupe set
without bound.
- Saves are gated on the cursor actually moving, so an idle channel does not
rewrite the file every poll interval. Both inbound transports are covered:
the poll sweep and the WebSocket event path share the same check.
Tests: six new cases in `TestChannelCursorPersistence` — the cursor is written
on seed, a restart resumes without spending a CLI call on history and then
delivers the mention that landed while the gateway was down, a foreign
identity or relay is ignored, a corrupt file falls back to seeding, the
restored `seen` set stays bounded, and an idle poll leaves the file untouched.
All six fail on main.
Tested on: Windows 11, Python 3.12. `python -m pytest
tests/gateway/test_buzz_adapter.py tests/gateway/test_buzz_websocket.py -q` —
33 passed (23 pre-existing + 6 new here, plus 4 WebSocket). Requires
`pytest-asyncio` (pinned at 1.3.0 in pyproject) — without it the async cases
in this file error out as unknown marks.
A relay-side close the transport never surfaces (observed as a CLOSE_WAIT
socket behind Cloudflare, #98097) parks the read loop forever while the
gateway keeps reporting connected: inbound stops, gateway_state.json stays
healthy, and only a restart recovers. The library keepalive should catch
this first, but as a last resort the read side now waits at most
_WS_READ_IDLE_TIMEOUT (300s) for a frame before raising into the existing
reconnect path, which re-authenticates and re-subscribes with per-channel
since filters intact.
Fixes#98097
- lightpanda_engine_status: check use_real_profile before the cloud
provider, matching browser_exec's actual resolution order (real-profile
resolution runs before backend resolution), so /browser status and
hermes doctor name the right shadowing setting when both are set.
- launch_lightpanda: drop the unreachable Windows popen_kwargs branch
(find_lightpanda_binary returns None on nt, launch errors out earlier).
- doctor: drop the over-defensive try/except around the cached
_using_lightpanda_engine() config read.
- Docstring: 'no-I/O gates' -> 'no network I/O (config reads only)'.
- New test pinning real-profile-over-cloud-provider reason precedence.
Browser Use mode never read browser.engine: _resolve_backend_cdp() went
BU_CDP_* env -> CDP override -> cloud provider -> local Chrome, so
`engine: lightpanda` was a silent no-op on the default backend, and on
the built-in path it was skipped whenever a cloud provider, Camofox or a
CDP override was active without anyone saying so.
- browser_use_cli: when the engine is lightpanda and nothing with higher
precedence claimed the session, get a session from _get_session_info()
and export its endpoint as BU_CDP_URL; the browser is private to the
session key, so the own-tab preamble is skipped. The browser_exec
description gains a Lightpanda header (text-first, new_tab once then
goto_url — lightpanda-io/browser#1962).
- browser_tool: _create_local_session() spawns `lightpanda serve
--host 127.0.0.1 --port <free>` per session key (new
tools/browser_lightpanda.py), reusing the session cache, inactivity
reaper and atexit cleanup; a dead process is respawned on the next call;
orphans from a crashed Hermes are reaped through per-process records in
$HERMES_HOME/cache/browser-use/lightpanda/. New lightpanda_engine_status()
reports whether the engine is in effect or what shadows it.
- tools_config: "Lightpanda" row in the Browser Automation picker
(cloud_provider: local + engine: lightpanda; "Local Browser" resets the
engine to auto) with a binary-check post-setup.
- /browser status and hermes doctor print the engine state and, when it
is shadowed, the reason.
Follow-up for salvaged PRs #82646 + #83414: resolution probes precede
publishes, the presentation-escape retry precedes the self-mention
downgrade, and a new test pins the escape-retry-delivers path.
Salvaged from PR #83414 (4 commits squashed to final state) and composed
with the presentation-mention escape retry from PR #82646 already on this
branch: send() now resolves @Name tokens to channel-member pubkeys
(membership-accurate via `channels members`, TTL-cached, Unicode token
boundaries, ambiguous names stay presentation-only) and passes explicit
--mention args; recovery ladder handles membership drift, unresolvable
prose @tokens (escape retry, #78797), and a final self-mention downgrade.
require_mention gated only on visible text, so Desktop thread replies
(e.g. /approve session) to the agent's own prompts were dropped with no
log. Cache event_id→(author, snippet) from seed/poll/WS/send, resolve
the direct e-tag parent, and dispatch when that parent is ours; also
populate reply_to_* on MessageEvent for gateway context injection.
Fixes#75826
The inbound path hardcoded Nostr kind 9 at both the WebSocket
subscription filter and the dispatch gate (which runs before mention
gating), so Buzz forum channels — kind 45001 thread roots and 45003
comment replies — were silently never dispatched to the agent; chat and
stream channels worked, making the gap invisible (#90309). Block's own
ACP harness documents the forum kinds explicitly.
Introduce _DISPATCH_KINDS = {9, 45001, 45003} for the subscription
filter and dispatch gate. The stream kinds (46010/40007/45002) stay out
of scope until their dispatch semantics are confirmed.
_is_direct_message_event deliberately keeps its kind-9-only check:
widening it would let a p-tagged forum post be reclassified as a DM and
bypass mention gating. The send path already works unchanged (send()
omits --kind and threads via --reply-to).
Fixes#90309
The fallback URL probe now uses 'get url' ({data:{url}}), not
'eval window.location.href' ({data:{result}}); the npx-resolution test's
mocked response predates the #81673 change.
Strip Chromium-only launch variables from Lightpanda commands, use a non-recursive Lightpanda URL lookup for Chrome fallback, and share sandbox argument injection with the temporary Chrome path.
Co-authored-by: forjd-hermes-bot <282037251+forjd-hermes-bot@users.noreply.github.com>
Follow-up to the salvaged manifest guard (#90859): the doctor's temporary
HERMES_HOME now enters the ExitStack before the staging copytree, so
ENOSPC / KeyboardInterrupt / any exception during staging deterministically
removes the hermes-plugin-doctor-* directory instead of relying on the
TemporaryDirectory GC finalizer (which cannot run while the traceback pins
the frame). Regression test proven via sabotage run against the old code.
`hermes plugins doctor` defaults its target to `.`, and
resolve_plugin_path accepted any directory that existed. Doctor then
copytree'd the resolved path into a temporary HERMES_HOME *before* the
runtime got to reject it, so running the command from a directory that
is not a plugin copied that whole tree.
Run from $HOME on macOS this copies the home directory, and because
`Library/CloudStorage` is not excluded it also materializes every
cloud-only Google Drive/iCloud placeholder. Observed locally: 461 GB
written to /private/var/folders and still growing when the process was
killed, on a machine with 49 GB free.
Resolve now requires a manifest before returning a path, mirroring
PluginManager._scan_directory: `plugin.yaml`/`plugin.yml`/`plugin.json`
in the directory itself, or in one immediate subdirectory for the
category layout. Plugin-id candidates are only tried for values that can
be an id, since joining `.` onto a plugins root resolves to the root and
would hand Doctor every installed plugin at once.
Non-plugin targets now fail with a clear message and no copy.
/simplify-code efficiency reviewer (verified): cleanup_stale=False does
NOT deliver the exclusion the salvaged fix intended — get_running_pid
returns None whenever a record fails liveness VALIDATION (start-time
mismatch, argv drift, lock hiccup) regardless of the flag, which only
controls unlinking. In exactly the at-risk scenario the recorded PID
still never joined the exclusion set.
Exclusion evidence now comes from the RAW pidfile + lock records (no
validation, no unlink side effects); the validated non-destructive
probe is kept for the runtime-status fallback PID. For a KILL exclusion
list this is strictly safer: a stale recorded PID at worst spares one
process for one sweep, while a validation false-negative would
TerminateProcess a live gateway. Regression test reworked to drive the
real function semantics (raw record present, validation rejects);
mutation-checked: removing the raw-record read fails the test.
Follow-up on the salvaged #87158: the reaper-side fix (probe with
cleanup_stale=False so the sweep never deletes its own exclusion
evidence) fully protects a healthy standalone gateway, so the
web_server.py skip-sweep gate is dropped — a stale-but-present
registration must not veto the #77276 orphan reap that motivated the
sweep in the first place.
New regression test drives the exact failure: a registration that
fails liveness validation only surfaces its PID when probed
non-destructively; with the destructive default the standalone gateway
would be hard-killed (TerminateProcess, no drain). Mutation-checked:
reverting the probe to get_running_pid() fails the test.
The threading cluster (#99429) branched before the auth cluster (#99427)
added auth_tag to _exec_buzz's signature; merging both left two
fake_exec doubles in test_buzz_thread_topology.py without the kwarg,
red on main. Sibling-test blast radius, no behavior change.
Compose the two salvaged approaches (#78065 + #78511):
- Keep #78065's terminal-only scrub-path exemption (first-party prefix
predicate in _make_run_env / _sanitize_subprocess_env, plain env values
never scope-resolved, snapshot exclusion for cross-profile isolation,
every non-terminal surface sealed).
- Fold #78511's BUZZ_MANAGED_AGENT signal into a context gate instead of
an import-time blocklist discard: the blocklist is shared by every
scrub surface, so discarding there would leak BUZZ_PRIVATE_KEY into
execute_code / hermes_subprocess_env children too.
- New gate _buzz_terminal_context_active(): BUZZ_MANAGED_AGENT in the
process env (Buzz Desktop buzz-acp harness, #76243) OR the live
session's platform is buzz (HERMES_SESSION_PLATFORM ContextVar,
concurrency-safe under a multi-session gateway). A Telegram/CLI/cron
session on a host that also runs a Buzz gateway does NOT get the
signing key in its terminal children (maintainer triage note on
#76243: don't expose the key to unrelated shell commands).
- Snapshot exclusion stays prefix-only (conservative even when the gate
is inactive).
- Tests updated for the gate + new negative test (non-Buzz session
strips) and positive test (buzz session platform enables carve-out);
docs updated accordingly.
Closes#78026, closes#76243.
Buzz platform agents could not use the `buzz` CLI from the terminal tool:
the BUZZ_* vars (BUZZ_PRIVATE_KEY, BUZZ_AUTH_TAG, BUZZ_RELAY_URL, and the
other BUZZ_* names) are added to _HERMES_PROVIDER_ENV_BLOCKLIST from the
buzz plugin.yaml (messaging category), and env_passthrough refuses to
re-allow anything in the blocklist (GHSA-rhgp-j443-p4rf). In the reported
`hermes acp` scenario the Buzz adapter's register() is never invoked, so
the agent runs in-process and its terminal uses _make_run_env directly —
there was no path for the platform credentials to reach terminal children.
Fix: a terminal-only, first-party carve-out in the scrub paths themselves
(not adapter registration). BUZZ_* vars pass through to foreground
(_make_run_env) and background/PTY (_sanitize_subprocess_env) terminal
children via a new prefix predicate (_TERMINAL_FIRST_PARTY_ENV_PREFIXES).
Everything else stays sealed and unchanged: the blocklist itself, the
env_passthrough refusal, execute_code scrubbing, hermes_subprocess_env
(browser/TUI-host/copilot-executor spawns), and docker children. The
GHSA-rhgp-j443-p4rf seal is preserved because no registration path is
opened; skills/config still cannot register these names.
Follow-up hardening from review:
- First-party matches use the merged env value directly instead of
_resolve_passthrough_value: under multiplex with no profile secret scope
installed the resolver raised UnscopedSecretError (fail-closed) at call
sites like the webhook-filter script runner, a regression where the
script previously ran without the var. The vars are the process's own
env values and are never scope-resolved.
- LocalEnvironment now excludes first-party terminal env names from the
shared login-shell snapshot (_additional_profile_scoped_passthrough_names
override): BUZZ_PRIVATE_KEY can never be in the get_all_passthrough()
exclusion set (env_passthrough refuses blocklisted names), so without
this a multiplexed gateway would dump profile A's key into
hermes-snap-<id>.sh and profile B sharing the collapsed LocalEnvironment
would source it — a cross-profile nsec leak. The names are now excluded
from the dump and save/restored per command.
- Docs now name the _sanitize_subprocess_env consumers (search workers
like ddgs, computer-use driver, user-script runners) that also receive
first-party platform vars.
Fixes#78026
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
Real adapter imports with synthetic NIP-10 relay events; covers the composed
cluster behavior end to end: in-thread triggers anchor replies to the thread
root (never nesting), top-level triggers still open exactly one thread,
reply_in_thread/reply_to_mode opt-outs post flat on send/send_image/
standalone cron send, progress thread resolution honors the opt-out, and
buzz no longer inherits verbose _GLOBAL_DEFAULTS (#95841). Sabotage-checked:
the display-tier tests fail against origin/main's display_config.py.
Buzz has no native thread_id; channel threading is entirely --reply-to on
the triggering event. Interim commentary and progress bubbles only passed
the anchor via metadata.reply_to_message_id (or not at all), so most Kathy
posts landed as new top-level messages and cluttered channels.
- Honor metadata.reply_to_message_id in BuzzAdapter.send
- Pass reply_to on stream commentary sends
- Treat buzz like slack/mattermost for progress thread resolution
- Set _progress_reply_to to the trigger event for buzz
- Add unit tests for adapter metadata and progress routing
Every Buzz reply opened a fresh thread, including when the user was
already replying inside one. A threaded client fills up with an endless
ladder of one-message threads and the conversation becomes unreadable.
The adapter itself never had threading logic; the behaviour comes from
the generic gateway default. `_reply_anchor_for_event()` in
gateway/platforms/base.py returns `event.message_id`, which is right for
reply-style platforms (Telegram/Discord "reply to this message") but
wrong for a thread-style one: anchoring to the message you are answering
nests a new sub-thread under every single turn.
Buzz threads are NIP-10, so the information needed is already on the
inbound event. The adapter now records each inbound message's thread
root from its `e` tags and resolves the outbound anchor to that root, so
a reply joins the thread the user is typing in. When the trigger was
itself top-level there is no root and the anchor passes through
unchanged, preserving the existing behaviour of opening exactly one
thread from a top-level message.
Fixed in the adapter rather than in `_reply_anchor_for_event()`: Buzz is
a plugin-supplied platform, and its NIP-10 tag semantics do not belong in
core. Root extraction prefers an explicit `root` marker, falls back to a
lone `reply` marker (a message bearing only `reply` started the thread,
so that parent is the root for everything after it), and treats a legacy
unmarked `e` tag as the parent. A mention-only `p` tag is not a reply.
The root cache is an OrderedDict bounded at 512 entries with FIFO
eviction so a long-lived gateway cannot leak, and the resolver is applied
to the image send path as well as `send()`. Both helpers tolerate a
missing `_thread_roots` attribute, since the standalone/cron send path
constructs an adapter without running `__init__`.
Tests cover root extraction (top-level, thread opener, nested, legacy
unmarked tag), the top-level passthrough that guards the existing
behaviour, unknown/None anchors, cache bounding and eviction, and an
end-to-end assertion through `send()` that `--reply-to` carries the root.
Verified against the real event shapes returned by a live hosted relay.