Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.
- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
streak raises a halt (identical_call_streak_halt) at
hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
stay exempt; a changed result resets the streak; warning-only sessions are
unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
for pollers, or when results change.
Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
identical failing read_file main: 602 API calls, budget exhausted
branch: 8 calls, repeated_exact_failure_block
identical successful terminal main: 602 API calls, budget exhausted
branch: 5 calls, identical_call_streak_halt
BasePlatformAdapter._acquire_platform_lock emits `{scope}_lock` with
retryable=True on purpose (#54167): a MID-RUN reconnect must be able to
recover once the live holder exits or a stale record is cleared. The
startup router keyed solely off that flag, so a live foreign holder of the
bot token at zero-connected startup landed in `_failed_platforms` with
gateway_state=running — alive, deaf, and retry-storming the token every
backoff — instead of the exit-78 (EX_CONFIG / startup_failed) contract
that #51228 established for single-writer conflicts.
Minimal class fix, salvaged from #83183 (@alexgunsberg) against current
main:
- gateway/restart.py: `is_global_startup_conflict(error_code)` — matches
the `*_lock` / `lock_conflict` code families every adapter emits for
scoped-lock and identity conflicts. Code only, never message text.
- gateway/run.py primary startup routing: a lock-conflict failure is
routed as non-retryable (parked `fatal`, not queued). Nothing else
connected → exit 78; alongside a transient peer → NS-609 mixed mode,
gateway stays alive and only the peer retries.
- gateway/run.py `_schedule_secondary_profile_startup_reconnect`: the same
contract for multiplex secondaries — park `<profile>:<platform>` fatal
like `duplicate_credential` instead of scheduling a reconnect storm.
- Mid-run behavior is untouched: `_handle_adapter_fatal_error_impl` and
the reconnect watcher still treat `*_lock` as retryable (#54167).
Not carried over from #83183 (superseded on main or out of scope): the
`degraded` lifecycle write only fires on the all-retryable path and the
runner immediately overwrites it with `running` (so busy/drain already
see `running`); the secondary retry bridge landed separately in
96489f3c1b (#92064); Buzz/IRC/LINE lock-tuple unpack and the reconnect
ownership registry are separate class fixes.
Live repro (real GatewayRunner.start(), isolated HERMES_HOME + lock dir,
live holder subprocess owning the lock via production
acquire_scoped_lock): before — exit_code=None, gateway_state=running,
telegram `retrying`, queued in _failed_platforms; after — exit_code=78,
gateway_state=startup_failed, telegram `fatal`, _failed_platforms={}.
Co-authored-by: alexgunsberg <alex@gunsberg.fi>
Follow-up to @zoser69's #78111 cherry-pick:
- lift the redirect into _redirect_platform_display_key() and apply it
BEFORE _validate_config_key / type coercion, so the unknown-key hint and
the string-vs-bool coercion both see the canonical path
- widen to the sibling surfaces: config get resolves the canonical key
(previously echoed the dead top-level value — the misleading half of the
report) and config unset removes the canonical leaf
- regression tests: get mirrors gateway resolve_display_setting, unset
removes the redirected leaf, note printed, helper touches ONLY
OVERRIDEABLE_KEYS (connection keys / 4-segment / already-canonical
paths untouched)
- docs: configuration.md per-platform section names the canonical CLI
path and the accepted shorthand
Follow-up to the salvaged #100350 commits: replace the per-table
'if table == "delivery_obligations"' branches in session_recovery.py and
session_lost_and_found.py with a single _AUXILIARY_TABLE_SCHEMAS registry
(table -> destination DDL initializer) that both the SQL-level and the
lost_and_found lanes consume, so the next lazily-created state.db table is
one entry, not three code paths. The .recover lane now iterates
_CANONICAL_TABLES + _AUXILIARY_TABLES instead of a duplicated literal list.
Tests: the .recover direct-copy lane creates the missing ledger on the
destination; a source-vs-destination obligation count mismatch fails
verification (complete=False) instead of reporting a clean salvage.
Docs: state.db table inventory lists delivery_obligations.
Addresses #100313
The 10s hygiene_max_turn_hold_seconds budget (#92318) releases the arriving
user turn while the summary model is still streaming. For thinking summary
models (DeepSeek-V4-Flash etc.) whose reasoning prefix alone exceeds 10s,
the abandonment path ALWAYS cancelled the commit fence — 100% of the summary
attempt (including the full thinking prefix) was discarded on every turn,
permanently disabling auto-compression while paying the summary model 10s
of thinking per turn, and the flat 60s retry-after then blocked the
agent-side preflight from a fresh chance.
Structural fix (maintainer-chosen direction in #97963): decouple the turn
from the compression instead of holding the turn longer or making the hold
progress-aware (which would reintroduce the #90845 frozen-turn bug):
- CompressionCommitFence gains mark_commit_watermark_fenced() /
commit_watermark_fenced; compress_context marks the fence right after
capturing get_active_message_watermark() under the durable compression
lock (#75316/#87484) — the property that makes a LATE commit safe: rows
appended after compression start survive both commit paths verbatim as
cloned concurrent tail (archive_and_compact watermark= and
publish_compression_child watermark/watermark_ceiling).
- gateway hygiene turn-hold handler: when the fence is watermark-fenced,
the detached worker (already kept alive via
_defer_agent_cleanup_until_future_done) KEEPS its commit admission; the
user's turn proceeds on the uncompressed transcript at the same 10s
budget, and the summary is adopted at the worker's own watermark-fenced
commit boundary. Unfenced workers are cancelled exactly as before —
never worse than the status quo.
- No retry-after is armed while the kept-admission attempt runs (it would
block preflight adoption via the same-session cooldown); re-attempt
spacing is covered by the durable compression lock
(_session_has_compression_in_flight). If the worker ends WITHOUT
committing, a done-callback restores the flat non-escalating 60s
retry-after; a successful adoption resets the hygiene failure streak.
The streak never advances for a deferral either way.
- Docs: configuration.md hygiene_max_turn_hold_seconds one-liner updated
to describe deferred adoption and the thinking-model case;
config_defaults.py comment updated. Knob stays config.yaml-only.
Invariants preserved:
- 10s user-latency cap stays hard (#90845/#92318):
test_session_hygiene_turn_hold_budget_abandons_streaming_wait passes
UNMODIFIED (its worker is not watermark-fenced, so it pins the cancel
path through the public surface).
- Stale-clobber impossible: adoption only rides commits bounded by the
start watermark; the fence still gates admission and unfenced/late
results are discarded.
New regression tests (tests/gateway/test_session_hygiene_turnhold_adoption.py):
- watermark-fenced worker keeps admission, late summary is committed,
turn still released at the budget, no cooldown while running,
streak reset on adoption;
- kept-admission worker that ends without committing restores the flat
turn-hold retry-after (<=120s, names turn-hold, streak untouched);
- unfenced worker still cancelled and discarded (status quo).
Sabotage-verified: disabling the keep-admission branch fails the two new
adoption tests and leaves the unfenced-cancel test green.
Fixes#97963
The generic tool_call(name, arguments: object) bridge hides a deferred tool's
real parameter schema from provider-native validation. Before this, only
top-level required-key absence was checked, so invalid enums, wrong types,
nested required fields and forbidden extra properties reached the handler
or MCP server. Now the call is coerced (same coerce_tool_args path normal
dispatch uses) and validated with the schema's declared JSON Schema draft;
failures return the path, constraint and parameters schema so the model
repairs the call in one round-trip. Fails open on missing/malformed schemas,
external $ref, or missing jsonschema.
Fixes#73175
Salvaged from #73179 onto current main (post core-tool deferral #97979).
Co-authored-by: teknium1 <teknium@nousresearch.com>
Adds a bundled image-generation backend for the Meta Model API
(https://api.meta.ai/v1), which is OpenAI-compatible. Exposes the
muse-image-1.0 model via the standard image_generate tool. This is the
image-gen companion to the already-bundled meta-ai chat provider
(plugins/model-providers/meta-ai, PR #88565).
- plugins/image_gen/meta-ai/ — provider registered as `meta-ai`, matching
the chat provider's id. Reuses the openai SDK pointed at Meta's base URL.
- Auth mirrors the chat provider: MODEL_API_KEY (Meta's documented var),
with META_API_KEY / META_MODEL_API_KEY aliases and a META_BASE_URL
override.
- Text-to-image only for now (capabilities gated); base64 (WebP) and URL
responses both handled and saved under $HERMES_HOME/cache/images/.
- Auto-loads as `kind: backend` and appears in `hermes tools` with no
central list edits, matching the other bundled providers.
- tests/plugins/image_gen/test_meta_ai_provider.py — 27 tests (metadata,
auth-alias resolution, base-url override, model resolution, generate
paths incl. b64 save, aspect mapping, URL caching, error handling).
- docs: image-generation feature page + provider-plugin built-in list.
Desktop's cold resume (defer_history + omit_messages, transcript paged over
REST) only ever holds the live tip segment in memory, but session.resume
bounded it against the FULL compression lineage (sessions.max_resume_messages,
default 20000). A Bot Chat with 85 compaction segments / ~29k lineage rows
behind a ~700-row tip was refused at 20001, sent zero model prompts, and sat on
"Waking up default…" forever — the healthiest possible session shape, rejected
by a guard sized for in-memory materialization.
- hermes_state: one `_resume_lineage_ids` definition shared by the resume
readers (get_resume_conversations, get_ancestor_display_prefix) and the
guard (assert_resume_safe / get_resume_message_count). Guard grows
`tip_only=` and names the scope it counted; the branch-aware lineage the
readers already used is now what the guard counts too (a /branch copy was
being counted against its parent's rows).
- tui_gateway session.resume: deferred, omit_messages and lazy resumes are
bounded by the tip; only the full in-memory lineage resume keeps the
lineage-wide bound. Deferred hydration falls back to tip-only history when
the lineage exceeds the limit instead of loading the rows the guard refused.
- CLI mid-setup tip-only path routes through the same guard instead of
borrowing assert_export_safe.
- docs: sessions.max_resume_messages / max_export_messages documented with the
per-surface scope.
Live repro (real SessionDB fixture, 85 segments / 29,226 lineage rows / 666 tip
rows, real tui_gateway.server.handle_request): before — deferred resume ->
4130; after — ok, hydrated history=666 prefix=0; the non-deferred full resume
still returns 4130 on the same fixture.
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.
- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
registry; piper/kittentts loaders extracted so warm-up and synthesis
share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.
Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
should_use_direct_api_call() contexts (gateway cron turns #62151, delegate_task
children #60203) were short-circuited onto the NON-streaming wire because the
interrupt worker wedges inside their nested thread pools. That dropped every
liveness property streaming provides: edge proxies kill the silent POST
(z.ai HTTP 524 — three retries later the child dies as "max_iterations"), and
the non-stream stale watchdog cannot tell a reasoning model's thinking phase
from a hung provider, so children die at exactly stale_timeout (#100260).
Keep those contexts on interruptible_streaming_api_call. The request now runs
INLINE on the conversation thread (no worker → the deadlock class stays
closed) while the existing poll loop — 30s heartbeat, stale-stream detector,
cross-thread interrupt abort — moves onto a monitor thread that only ever
aborts sockets, never dispatches (same shape as direct_api_call's watchdog
timer). Interactive sessions are unchanged: worker + poll loop as before.
should_use_direct_api_call() itself is untouched; only what it routes to.
Live A/B (real SSE server, real AIAgent.run_conversation):
before: subagent/cron wire stream=None, request on conversation thread
after: subagent/cron wire stream=True, request on conversation thread
cli unchanged (stream=True, spawned worker)
inline stale detector kills a one-chunk-then-silence stream at budget;
AIAgent.interrupt() from another thread unwinds the inline stream in 0.6s.
Co-authored-by: Expri-commits <184641533+Expri-commits@users.noreply.github.com>
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
Curated picker lists (OPENROUTER_MODELS + _PROVIDER_MODELS['nous']) gain
claude-fable-5.1 above claude-fable-5 per newest-first ordering; manifest
regenerated via scripts/build_model_catalog.py.
Provider-agnostic metadata verified as already resolving for the 5.1 slug
(no new entries needed): DEFAULT_CONTEXT_LENGTHS fuzzy-matches the
claude-fable-5 prefix (1,000,000), reasoning stale-timeout floor fires
(600s), and both routes bill via official_models_api (live pricing, no
snapshot entry required).
The 20s ws-orphan grace (14b50f5edd) interrupts a RUNNING turn whenever
the client is absent past the grace window — killing healthy long turns
on deliberate client absence (desktop closed, PC asleep, mobile
backgrounded, Electron tab-switch throttling, desktop update/relaunch).
The reaper now interrupts a detached running turn ONLY when BOTH the
client is absent past the grace AND the turn's activity clock is stale
(seconds_since_activity >= dashboard.ws_orphan_activity_stale_s,
default 600s — matching agent.turn_liveness.timeout_s semantics from
PR #99758). A detached-but-actively-producing turn keeps running to
completion (the sentinel transport already buffers detached emits);
a detached AND activity-stale turn is interrupted/reaped as today.
Non-running orphaned sessions keep current behavior. Reuses the
existing AIAgent.get_activity_summary() clock — no parallel tracker
(rejected in PR #4864).
Fixes#98028Fixes#100325
- Translate leftover English 'toolset(s)' strings in ru.ts
- Cover ru aliases/config value in languages.test.ts (parity with ar)
- List Russian among desktop UI languages in website/docs/user-guide/desktop.md
The skill_manage tool schema description, prompt-builder docs, and the
skills docs page now derive the creation path from skills.create_dir
(display_skill_create_dir()) instead of hardcoding ~/.hermes/skills/ —
so pointing the config at e.g. /opt/brain/skills changes what the agent
is told everywhere, with no SOUL.md fights or read-only chmod tricks.
Adds config default + docs section + 16 tests (incl. a read-only
profile-skills-dir scenario).
The Skills tab now lists the entire official optional-skills catalog
(optional-skills/ shipped with the repo) below the installed skills.
Each catalog row has an Install button that routes through the standard
hub action pipeline; once the install finishes the row flips into the
installed list with the normal enabled/disabled toggle.
- backend: GET /api/skills/hub/official — OptionalSkillSource.list_local()
scan (no network) + per-profile installed flags from the hub lock
- desktop: catalog section in SkillsView with search/scope integration,
install-state spinners off $hubActions, and an OfficialSkillDetail pane
(hub preview: frontmatter + full SKILL.md + Install)
- CapRow gains an optional action slot (button instead of the Switch)
- electron: route the new endpoint with the skills family (primary backend)
- i18n: officialCatalog/officialPill keys across en/ja/zh/zh-hant
Final-diff review findings (/simplify-code on the full 3-commit stack):
- Temp-dir fallback uses a per-uid name (hermes-profile-exports-<uid>) and
get_profile_export_path refuses a pre-existing symlink or a directory
owned by another user — a fixed /tmp/hermes-profile-exports is a
predictable shared path a local attacker could pre-create to receive the
secret-bearing archive. Regression test mutation-checked.
- Fail-closed message reworded interface-neutrally (the web API surfaces it
as HTTP 400 detail where '-o' alone made no sense).
- Docs now cover the temp fallback and the fail-closed refusal.
Follow-up to the salvaged #92689:
- _profile_export_directory() now proves safety on the export dir's OWN
ancestry (_inside_git_checkout) instead of walking Path.cwd(). The old
heuristic missed the checkout whenever HERMES_HOME sat inside one but
the process ran from elsewhere (cron, service manager) — the export
landed back inside the source tree, the exact incident class.
- When every candidate is inside a checkout, warn instead of silently
violating the invariant.
- CLI/TUI export callers: move get_profile_export_path() inside the try
and catch OSError too — a bad profile name or read-only home printed a
raw traceback instead of the clean error main previously gave.
- Tests: bind module objects at call time (importlib) so sibling reload
pollution in the tests/hermes_cli sweep can't divorce monkeypatches
from the code under test; add regression tests for the cwd-independent
topology and the clean-error path.
- Docs: mention the ~/.hermes-profile-exports fallback store.
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
Off should mean the model is never told the tool exists. A switch that
only made the call fail leaves Hermes offering walkthroughs it cannot
give and promising to point at things it cannot point at, which reads as
a broken agent rather than a respected preference.
Both gate on the switch through a shared desktop_ui.user_enabled helper,
which is the reactions check_fn generalized — same config read, same
reason it has to be config rather than an env var: the switch belongs to
the session's client, and the client may be on another machine.
Closes the #95663 round-8 review blocker (false settlement before
commit veto): the pre-commit surface (`_surface_stall`) logged
"Force-aborting the turn and stopping lease renewal" and warned the
user "aborting it so the session can recover" BEFORE `_commit_abort`
could veto — so a turn that resumed during the warning window (or an
exceptional interrupt path that declines fail-closed) was reported as
force-aborted with lease stopped while it actually continued running.
- Split the surface: `_surface_stall` is now observational only ("no
progress for Ns; attempting recovery"), and the definitive
aborted/lease-stopped settlement moves to a new
`_surface_committed_abort` that runs only after `_commit_abort`
succeeds and the turn lease is deactivated.
- Rate-limit repeated pre-commit surfaces per observed generation: a
turn whose aborts keep declining no longer re-logs an ERROR and
re-warns the user every poll interval.
- Add the committed-path regression test
(`test_watchdog_publishes_definitive_settlement_only_after_commit`)
and extend the declined-path witness
(`...resumes_during_warning`) to assert no committed-abort or
definitive pre-commit claim appears when the abort is vetoed. Both
fail on the pre-fix tree (mutation-checked).
- Document the `_interrupt_turn` lease-loss asymmetry (fires
unconditionally, no generation claim — losing the lease means the
process no longer owns the session).
- Trim review-round archaeology from comments/docstrings (keep the
WHY, drop the round numbering), and drop the dead
`cancel_event` compat note from the test fence.
- Document `agent.turn_liveness` in the configuration guide.
On top of PR #95663 by Finn763 (cherry-picked with authorship
preserved).
A non-admin '/sessions all' or '/resume --all' silently downgraded to
chat-scoped listing with zero feedback, which reads as 'my session
vanished' (community Telegram report). Both surfaces now append a notice
that cross-chat listing requires a configured admin. Follows up the
salvaged current-session '(current)' marker (PR #68556, fixes#68547):
sibling tests updated to pin the new contract, new i18n key
gateway.resume.all_requires_admin added to all 17 locales, docs updated.
Builds on webtecnica's escape-aware _split_key_path (#84152, cherry-picked
with authorship preserved; earliest fix in the family was RelaxJonh's #80253
greedy-match approach — both behaviors now ship together):
- _greedy_literal_match: when navigating an EXISTING mapping, prefer an
existing literal key equal to the dot-join of the next N path segments
(longest match wins). Dotted model IDs are the norm, so the common
unescaped command (config set providers.p.models.grok-4.6.supports_vision
true) now hits the real key across set/get/unset instead of creating a
phantom sibling. Plain dotted paths with no dotted-key collision split
exactly as before.
- _phantom_sibling + ValueError in _set_nested: refuse to CREATE a new
intermediate mapping that would shadow an existing dotted literal sibling
(Soju06's fail-loudly suggestion on #84064); set_config_value surfaces it
as a clean CLI error with the escaped spelling to use.
- utils.py::atomic_roundtrip_yaml_update (the second split site, #91607 —
/model + TUI persistence) now uses the same escape-aware split + greedy
literal matching.
- CFG-04 empty-segment guard now splits escape-aware so escaped keys are
not misclassified.
- Tests for every repro shape in the family: #84064 provider model keys,
#80006 Matrix room IDs, #91095 dotted models under custom_providers list
index (incl. escaped creation-when-absent), #91607 model_overrides via
atomic_roundtrip_yaml_update, #99124 dotted leaf keys; plus
backward-compat coverage. Also fixed the carrier's one stale assertion
(structured-value coercion landed on main after #84152 branched) and
removed its dead _MCP_SECRETS_CONFIG fixture flagged in review.
- Docs: 'Dots inside key names' section in website/docs/reference/cli-commands.md.
Fixes#84064, fixes#80006, fixes#91095, fixes#91607, fixes#99124
Salvaged from PR #84199 by @RickyYii. DirectAlias gains api_key/key_env; the direct-alias override re-resolves credentials against the alias endpoint (host-gated, #28660) and reuses the pre-alias key only on an origin match; oneshot -m <alias> passes the alias key as explicit_api_key; direct-alias branch gains the OLLAMA_API_KEY host gate. Fixes#83612.
Localizing inbound relay media spends the agent's own Buzz credentials on
a URL chosen by the sender, so it must not run on the strength of the
adapter's local allow-list alone. Require the gateway's authorization
callback to return an explicit True before any `buzz media get` runs; a
denial, a missing callback, or a raising callback fails closed and leaves
the message text exactly as it arrived.
`_is_sender_authorized` previously wrapped the callback result in
`bool()`, so a truthy non-boolean (a status string, a sentinel) would
satisfy an `is True` gate's intent while bypassing its guarantee. Only
the literal booleans now propagate; anything else is "unknown", which the
existing Slack and Discord callers already treat as trust-unknown.
Reviewers asked for this boundary on the sibling inbound-media PRs
(#77734, #78051); it applies equally to the retrieval path in #75614,
which this change builds on.
Widens the salvaged cron doctor with the highest-value fleet check:
an active job whose next_run_at is parked >15min in the past is not
firing (dead ticker, downed gateway, wedged fire-claim). Also registers
doctor in the docs (cron guide + CLI reference) and resolves the salvage
onto current main alongside runs/incidents/notepad.
Hardening on top of @Soju06's forwarding fix: v2 providers written against
the original docs example (def on_pre_compress(self, messages)) must not
TypeError when the host forwards require_checkpoint — inspect the signature
and fall back to the legacy call shape. Docs example updated to advertise
the keyword.
Compose the two salvaged approaches (#78065 + #78511):
- Keep #78065's terminal-only scrub-path exemption (first-party prefix
predicate in _make_run_env / _sanitize_subprocess_env, plain env values
never scope-resolved, snapshot exclusion for cross-profile isolation,
every non-terminal surface sealed).
- Fold #78511's BUZZ_MANAGED_AGENT signal into a context gate instead of
an import-time blocklist discard: the blocklist is shared by every
scrub surface, so discarding there would leak BUZZ_PRIVATE_KEY into
execute_code / hermes_subprocess_env children too.
- New gate _buzz_terminal_context_active(): BUZZ_MANAGED_AGENT in the
process env (Buzz Desktop buzz-acp harness, #76243) OR the live
session's platform is buzz (HERMES_SESSION_PLATFORM ContextVar,
concurrency-safe under a multi-session gateway). A Telegram/CLI/cron
session on a host that also runs a Buzz gateway does NOT get the
signing key in its terminal children (maintainer triage note on
#76243: don't expose the key to unrelated shell commands).
- Snapshot exclusion stays prefix-only (conservative even when the gate
is inactive).
- Tests updated for the gate + new negative test (non-Buzz session
strips) and positive test (buzz session platform enables carve-out);
docs updated accordingly.
Closes#78026, closes#76243.
Buzz platform agents could not use the `buzz` CLI from the terminal tool:
the BUZZ_* vars (BUZZ_PRIVATE_KEY, BUZZ_AUTH_TAG, BUZZ_RELAY_URL, and the
other BUZZ_* names) are added to _HERMES_PROVIDER_ENV_BLOCKLIST from the
buzz plugin.yaml (messaging category), and env_passthrough refuses to
re-allow anything in the blocklist (GHSA-rhgp-j443-p4rf). In the reported
`hermes acp` scenario the Buzz adapter's register() is never invoked, so
the agent runs in-process and its terminal uses _make_run_env directly —
there was no path for the platform credentials to reach terminal children.
Fix: a terminal-only, first-party carve-out in the scrub paths themselves
(not adapter registration). BUZZ_* vars pass through to foreground
(_make_run_env) and background/PTY (_sanitize_subprocess_env) terminal
children via a new prefix predicate (_TERMINAL_FIRST_PARTY_ENV_PREFIXES).
Everything else stays sealed and unchanged: the blocklist itself, the
env_passthrough refusal, execute_code scrubbing, hermes_subprocess_env
(browser/TUI-host/copilot-executor spawns), and docker children. The
GHSA-rhgp-j443-p4rf seal is preserved because no registration path is
opened; skills/config still cannot register these names.
Follow-up hardening from review:
- First-party matches use the merged env value directly instead of
_resolve_passthrough_value: under multiplex with no profile secret scope
installed the resolver raised UnscopedSecretError (fail-closed) at call
sites like the webhook-filter script runner, a regression where the
script previously ran without the var. The vars are the process's own
env values and are never scope-resolved.
- LocalEnvironment now excludes first-party terminal env names from the
shared login-shell snapshot (_additional_profile_scoped_passthrough_names
override): BUZZ_PRIVATE_KEY can never be in the get_all_passthrough()
exclusion set (env_passthrough refuses blocklisted names), so without
this a multiplexed gateway would dump profile A's key into
hermes-snap-<id>.sh and profile B sharing the collapsed LocalEnvironment
would source it — a cross-profile nsec leak. The names are now excluded
from the dump and save/restored per command.
- Docs now name the _sanitize_subprocess_env consumers (search workers
like ddgs, computer-use driver, user-script runners) that also receive
first-party platform vars.
Fixes#78026
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
Compose/fix-up on top of the cherry-picked cluster commits:
- Unify the config surface: platforms.buzz.reply_to_mode: off (PlatformConfig
field, as Discord/Telegram) and extra.reply_in_thread: false (the key Slack
users know; env BUZZ_REPLY_IN_THREAD) are equivalent opt-outs, bridged
through _apply_yaml_config and honored by send(), send_image(), and
_standalone_send (cron delivery).
- Progress/status bubbles honor the opt-out too: gateway/run.py resolves
_progress_reply_in_thread from the Buzz adapter (mirroring the Slack path)
so the synthetic-thread fallback and the progress reply anchor are both
suppressed when the user asked for flat replies (#75082, #95842).
- Deduplicate NIP-10 parsing: inbound session thread_id now reuses
_extract_thread_root (marked root > reply > legacy positional e-tag)
instead of a second inline root-marker-only scan.
- display_config: add buzz to _PLATFORM_DEFAULTS at TIER_MEDIUM — with
edit_message now implemented, accumulate-style progress works, but without
the entry Buzz inherited the verbose _GLOBAL_DEFAULTS and every interim
update became a permanent channel post (#95841).
- plugin.yaml optional_env + platform docs for the new keys.
- contributors/emails mappings for the cherry-picked authors.
Two growth leaks closed:
1. Pushed-branch tier (the dominant survivor class — 24 of 33 preserved
trees, ~18GB on the reporting box): managed installs fetch with a
single-branch refspec, so pushed PR branches never get refs/remotes/*
entries and read as 'unpushed' forever. When a clean tree's branch head
EXACTLY matches origin (one lazy git ls-remote per sweep), the checkout
is redundant: reap the TREE, keep the BRANCH ref (shielded from the
orphaned-branch pass). Anything diverged/unverifiable stays preserved.
Applied to both the startup pruner and hermes worktree prune/list.
2. Cron-tick maintenance: the pruner only ran on hermes -w launches, so
gateway-driven boxes accumulated trees for days. The scheduler tick now
dispatches the same conservative pruner on a daemon thread, throttled
to once per 6h, against the install checkout + job-workdir repos that
have a .worktrees/ dir.
Remove the implicit hermes peer and the peer question from new connection setup. Preserve explicit peer settings and keep memory paths consistent with the captured client identity.
Add setup, configuration, request, recall, and session regression tests, plus upgrade guidance.
Follow-ups on the salvaged #97797 transport:
- Blocker 2 from the #97681 exact-head review claimed near-expiry refresh
silently mints against the target's CURRENT policy. On this head the
handler DOES refuse drift (_require_unchanged_execution_policy -> 403
room_reauthorization_required), but nothing pinned the handler-level
behavior: removing the drift check still passed the entire grants suite
(the check was only unit-tested in isolation). New HTTP-level regression
test drives /v1/room-members/grants/refresh with a drifted-policy grant
and requires the 403; sabotage-verified (check removed -> test fails).
- cancel() conflict resolution: keeps our race-retry routing loop from
#99099 with this layer's peer-stop acknowledgement body inside it
(peer receipt -> settle completion -> local interrupt escalation).
- docs: NAT one-way-reachability note in bot-mode.md — Desktop is a viewer,
not a relay; put room authority on the host everyone can reach (field
finding from /bin/bash on #97681).