No token and no cookie cannot self-heal. Tag that throw with
isReauthRequired so boot stops retrying and Sign in stays clickable.
Leave needsOauthLogin-only ticket 401s retryable for AT/RT rotation.
Widens the cherry-picked reasoning-callback fix to the whole leak class
(#93220):
- quiet branch also neutralizes tool_progress_callback,
tool_start_callback, tool_complete_callback (inline diff rendering via
render_edit_diff_with_delta was gated by NEITHER quiet_mode nor
tool_progress_mode) and syncs agent.tool_progress_mode='off'.
- _should_emit_quiet_tool_messages() returns False under
suppress_status_output: with callbacks neutralized, the quiet-mode
KawaiiSpinner fallback printed '[tool]'/'[done]' lines into captured
stdout. Also covers oneshot.py and background-review forks, which set
the same flag and expect strict silence.
E2E (isolated HERMES_HOME, live model, write_file turn): base leaks
'┊ review diff' + full SVG source into stdout; head emits exactly the
final response. Regression tests pin the quiet-branch statements and the
gate (sabotage-verified).
Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
The -Q quiet single-query path suppresses stream_delta_callback and
tool_gen_callback to keep stdout machine-readable, but missed
reasoning_callback. When display.show_reasoning is on (the default),
the reasoning box leaks into stdout before the final response,
corrupting output for automation wrappers and third-party integrations
using --source tool.
Before:
hermes chat -Q --source tool -q "Reply with exactly: PING_OK"
┌─ Reasoning ──────────────────────┐
The user wants me to reply...
PING_OK
After:
PING_OK
Follow-up on top of the salvaged cluster: sanitize_api_messages step 3
(duplicate tool_call_id dedup) still tracked only the coalesced
(call_id||id) value in outstanding_call_ids, so after step 2's
variant-aware matching preserved a result keyed on the OTHER id variant,
step 3 deleted it as answering no outstanding call — whole parallel
batches of real results vanished with no stub at all (#93251's total-loss
mode). Track the full variant set per call and consume all siblings when
answered, preserving #58327 duplicate protection and llama.cpp
constant-id re-arm semantics.
Also aligns the #58287 compressor test with the in-flight tool chain
protection (#79278) that landed after that PR was opened: a trailing
user turn keeps the negative-control assistant message out of the
protected trailing window.
New regression tests: divergent-id batch survival through the dedup
pass, sibling-id replay still dropped, constant-id re-arm preserved.
Sabotage-verified: tests fail with the old single-id tracking.
The tool_call id-matching pass in repair_message_sequence only read
`.get("id"/"call_id")` on plain dicts, skipping non-dict tool_calls
entirely (`if not isinstance(tc, dict): continue`). Host-fed and
pre-serialization histories can carry unserialized SDK tool_call
objects (e.g. `ChatCompletionMessageToolCall`) instead of dicts, which
left `known_tool_ids` empty for that assistant turn. The following
`tool` message — a legitimate result already produced by executing the
tool — was then misclassified as an orphan and silently dropped,
corrupting the persisted conversation history and leaving the
assistant's tool_calls unanswered (itself a trigger for HTTP 400 on
strict providers).
Fix: extract id/call_id via getattr() for non-dict entries too,
mirroring AIAgent._get_tool_call_id_static's existing dict-or-object
tolerance, instead of skipping them.
`repair_message_sequence` registers BOTH `id` and `call_id` for each
assistant tool_call, because a matching tool result may be keyed on either
depending on which path built it (#58168). The duplicate guard added for
dropped rather than replayed.
Those two behaviours don't compose: a Codex/Responses tool_call registers
two DIFFERENT ids (`fc_...` and `call_...`), but only the id the first
result referenced is discarded. Its sibling stays in `known_tool_ids`, so a
duplicate result keyed on that sibling still matches and is kept — two tool
messages replayed for one call, which is exactly the HTTP 400 on strict
providers the consume step exists to prevent.
Duplicates of this kind come from the retry / crash / session-resume glitch
the guard was written for; the id-variant split just lets them slip past it.
Track each registered id back to its tool_call's full variant set and
discard all of them on a match. Results keyed on either variant are still
accepted (no false orphaning), and two parallel Codex calls answered via
different variants both survive.
Adds regression tests for the sibling-keyed duplicate and for the
two-calls/mixed-keys case that must NOT be affected.
_sanitize_tool_pairs() matched tool_call/tool_result pairs using a
single-value call_id||id precedence per tool_call (_get_tool_call_id).
In the Codex Responses API format an assistant tool_call carries both a
distinct id (fc_...) and call_id (call_...); a tool result's
tool_call_id may be keyed on either depending on which code path built
it. Whenever a genuinely matching pair used the field the precedence
didn't pick, the sanitizer misclassified it as orphaned on BOTH sides:
it dropped the valid tool result AND stripped the tool_call from the
assistant message, even though neither was orphaned.
Live-verified before the fix: {"id": "fc_777", "call_id": "call_777"}
+ a tool result with tool_call_id="fc_777" (a valid pair) was fully
removed by current main.
Register both id and call_id as valid match keys via a new
_tool_call_id_variants() helper (a set per tool_call, not a single
value), matching #58168's fix for repair_message_sequence's known-id
set today. A tool_call now survives if ANY of its id variants has a
matching result, which is not vulnerable to precedence order at all
(unlike swapping which field is checked first, which only trades which
sub-case is broken).
Note on #56425 (open, unreviewed): that PR touches this same function
for the same underlying issue (#55626) by swapping the call_id||id
precedence to id||call_id. That fixes the specific case where a result
matches `id` but not the reverse case (a result matching `call_id`
while `id` is also present) -- the precedence-swap approach cannot fix
the class, only relocate which sub-case is broken. This fix instead
mirrors the already-merged #58168 pattern (register the superset of
both ids as valid matches), which has no such blind spot. Adds 2
regression tests: the previously-mismatched case, and a negative
control confirming genuine orphans are still stripped alongside a
valid dual-id pair in the same window.
Register every id variant (call_id AND id) of each assistant tool_call in
sanitize_api_messages so a tool result keyed on either variant is treated
as paired. Previously only the coalesced (call_id||id) value was
registered, so Responses-style tool_calls carrying divergent id (fc_...)
and call_id (call_...) had their real results dropped as orphans and
replaced with '[Result unavailable]' stubs.
Cherry-picked from PR #56148 (unrelated busy_ack_templates files dropped
per the author's own follow-up commit).
The hard block matched raw command text, but a shell resolves quote
splicing (`kick"start"`) and backslash escaping (`kick\start`) into the
literal verb before execution. So `launchctl kick"start" -k
gui/501/ai.hermes.gateway` ran exactly as the blocked `kickstart` form
while both the non-bypassable block and the approval detector missed it —
leaving an approval-bypassing gateway self-lifecycle operation reachable.
contains_gateway_lifecycle_command now runs a second pass over
shlex-tokenized command segments, where quotes and escapes are already
resolved. It stays anchored on a hermes-gateway identifier, so prose and
non-gateway hermes services are unaffected. Because this function is the
single choke point _contains_unsafe_gateway_action calls at every
recursion level, referenced-script and `sh -c` payload scanning inherit
the fix.
tools/approval.py had the same gap for quote splices: backslash escapes
are stripped by _normalize_command_for_detection, but quote splicing in an
ARGUMENT position is not touched by _deobfuscate_shell_word_for_detection
(scoped to command-position words, deliberately — widening it would let
quoted prose match the destructive patterns). It now delegates to the
fixed guard as a last check, so an ordinary pattern match still wins and
keeps its more specific reason string.
Tests: quoted, single-quoted and backslash-spliced verbs across the
launchctl/systemctl/hermes branches, the spliced gateway identifier
itself, a splice nested in an `sh -c` payload (resolves one level deeper,
asserted at the recursive entry point terminal_tool actually calls), plus
negative cases proving prose and non-gateway labels stay unblocked.
Verified on Windows: no regressions — the 10 remaining failures across
tests/tools/test_approval.py, tests/hermes_cli/test_gateway_restart_loop.py
and tests/cron are identical on the unmodified baseline (POSIX file modes,
symlink privileges, and /bin/bash script paths).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Branch B of _GATEWAY_LIFECYCLE_PATTERN enumerated launchd verbs but omitted
`bootout` - the modern replacement for the `unload` it already listed, and
the paired inverse of the `bootstrap` it already listed. `remove` (legacy
sibling of bootout) and `disable` (what makes an unload durable) were
missing for the same reason.
This matters because the two enforcement layers are not interchangeable. In
tools/terminal_tool.py under _HERMES_GATEWAY == "1":
- the cron.lifecycle_guard hard block is documented as applying
unconditionally ("force=True cannot help here")
- detect_dangerous_command below it is explicitly skipped when force=True
detect_dangerous_command already flags all three verbs, so the default path
was covered - but with force=True inside the gateway they reached execution
while stop/unload/kickstart did not. SIGTERM then propagates to the child
before the command completes and the service may never come back, which is
the state described in #74973.
The label anchor (\bhermes[.\-]?gateway) is unchanged, so unrelated services
such as `launchctl bootout gui/501/ai.hermes.update-checker` stay runnable.
Adds TestLifecycleGuardLaunchctlParity, which pins the one-directional
invariant: anything the bypassable approval layer flags, the unbypassable
hard block must also catch. Deliberately not equality - the hard block is
legitimately stricter (it also covers load/restart, which the approval layer
leaves alone). Verified failing on the parent commit for exactly bootout,
remove and disable.
Closes#80260
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The Bot Mode 'session not found' / bot-runs-on-wrong-backend bug. wiring's
requestGateway is ONE shared closure for every session-scoped RPC in the
window, but it derived the owning profile from the globally-FOCUSED tile
($focusedStoredSessionId). A bot chat is a background tile while another pane
is active, so its prompt.submit carried the bot's own session_id yet was
dispatched on the FOCUSED tile's backend — the default backend served the bot
via ?profile= from the default's state.db, or answered 4001 'session not
found' when it didn't hold the runtime session.
Route by the session the RPC TARGETS (params.session_id) instead. session_id
is a RUNTIME id while tiles/rows key on the STORED id, so translate via the
state cache then a reverse scan of the stored->runtime map (the same ladder
use-session-tile-delegate's storedSessionIdForRuntime uses); an unresolved id
is already a stored id (several RPCs pass stored ids directly). RPCs with no
session_id (ambient/config) keep the focused->selected fallback.
Pure helpers extracted to wiring-routing.ts so they're unit-testable without
importing the React controller; 6 tests cover the target-vs-focused routing,
the stored-id passthrough, and the no-session fallback. tsc 0 errors.
Diagnosis verified on a live install: the fix was present in source but the
running build still misrouted, and logs showed the bot's turn executing on the
default backend while its own per-profile backend sat idle.
Prose false positives (Branch A trailing boundary, Branch D leading
boundary), quoted-multiline data payloads, and the fail-closed
unbalanced-quote fallback.
- Add _split_logical_lines() to split on newlines outside quotes, fixing
false positives from quoted multi-line payloads (e.g. python -c "...")
being torn into fragments and scanned as referenced scripts.
- Make _iter_command_segments() use logical line splitting with fallback
to per-physical-line tokenization for unbalanced quotes.
- Add missing word boundaries to _GATEWAY_LIFECYCLE_PATTERN:
* Branch A: trailing \b after restart|stop
* Branch D: leading \b before p?kill to prevent matching "skill"
(and similar words ending in "kill")
Fixes#92372: gateway lifecycle guard false-blocks on prose inside
a referenced data file.
A legacy remote primary carries no registry connectionId, so the scoped
reconnect reset could not name the restarted owner and fell back to
preserving every owner-routed Bot tile -- leaving the restarted backend's
own Bot Chat bound to its dead runtime (the original bug, persisting for
that one connection shape).
Unknown identity now fails toward recovery instead: preserve only Bot
runtimes owned by provably-live secondary connections
(liveSecondaryConnectionIds()); everything else drops its binding and
re-resumes. A reset only costs a re-resume, so this is safe for the
preserved-set survivors and correct for the dead one.
Follow-up to @BrunoBza's #93062:
1. Set failed=True only for the new repeated_outer_errors exit reason.
Previously the error exit left failed=False, so finalize_turn reported
completed=True for a turn that actually failed — incorrect.
2. Don't append_message the assistant response at the break. A thinking-
prefill or interim assistant may already be the tail, and appending
would create assistant→assistant role-alternation violation.
finalize_turn (lines 341-353) handles this safely by checking
_tail_role != 'assistant' before appending.
3. Update test to assert failed=True and completed=False for the
repeated_outer_errors exit.
The outer conversation-loop except handler only left the loop on a
local-processing error or when api_call_count >= max_iterations - 1.
With the turn budget now unlimited by default (sys.maxsize), a
permanent failure that escaped the inner retry/fallback machinery
retried forever: ~64 retries/s, one core pegged, and the rotated
agent.log history overwritten within minutes.
Bound the loop with a small per-turn cap on total escaping exceptions
(_MAX_OUTER_LOOP_ERRORS = 8, scaled down by a tiny explicit
max_iterations so a manually bounded budget still governs). The legacy
local-processing and near-limit exits are byte-identical; a new
'repeated_outer_errors' exit reason gets a user-facing explanation.
The inner retry/fallback layer owns transient API recovery and
terminates on its own, so only exceptions that escape it reach this
cap - a successful turn is unaffected.
Fixes#92450
Brief transport blips often self-heal in 1–3 minutes. Raise the
non-blocking escalate toast from 45s to 5m so those windows stay quiet
while chat remains readable/draftable. Confirmed reauth still takes the
full-screen recovery overlay immediately.
Post-boot WebSocket ticket mint failures and prolonged reconnects were
promoting into the full-screen "Hermes couldn't start" overlay, locking
users out of reading/drafting during brief 1–3 minute remote flaps.
- Ignore non-reauth boot-progress errors after a healthy cold boot
- Escalate prolonged transport reconnects with a non-blocking toast
- Soft-reset remote liveness rebuilds (no boot UI reset)
- Retry transient ws-ticket mints; auth rejections still fail fast
requestGateway (contrib/wiring) resolved the owning backend from
$focusedStoredSessionId — the WINDOW's focused tile — for every RPC it
dispatched. A session-scoped RPC names its real target in
params.session_id; whenever that session's chat was NOT the focused
pane (any background bot chat — the normal Bot Mode case), the RPC was
dispatched on whichever backend the focused tile happened to own. A
profile bot's prompt.submit then executed on the DEFAULT backend
(sessions created in the default store, profile logs empty), or failed
with 4001 'session not found' when default didn't hold the session.
Root cause isolated by Teknium: submit for a Developer-profile bot ran
in root logs/agent.log while profiles/developer/logs sat empty, and
completed only when default happened to hold the session — proving the
misroute sits downstream of the tile-owner-route lookup, in the
routing-key choice itself.
requestGateway now routes by the RPC's own target first:
params.session_id (a RUNTIME id) is translated to the stored id via the
tile map — new storedSessionIdForRuntimeId() in session-states, where
tiles already carry both identities — and only session-less RPCs
(config reads, list refreshes, cron) fall back to the focused-tile key,
which is genuinely window-ambient. Stored-id claims win over runtime
bindings in the lookup so a stale tile's dead runtimeId can never
hijack a live tile's identity.
PR #93269 (kshitijk4poor) landed the outer-handler break for the same
symptom while this branch was in flight. Keep his guard (it covers
shutdown errors from local post-processing and does the resume-hint +
best-effort persist) and keep this branch's inner-retry-handler return
(it fires BEFORE the ⚠️ retry trace, credential rotation, and fallback
attempts that the outer handler never sees). Point his
_is_interpreter_shutdown_error at tools/interpreter_shutdown.py so the
class has exactly one text-matching site, preserving his RuntimeError
type gate and all 7 of his tests.
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.
Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
(matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
predicates now delegate to the shared home (tool_executor previously
matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
shutdown signal and abandons the turn — one log warning, no print,
no traceback, no debug dump, no retry; outer handler gets the same
guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
(set by the background-review fork) instead of bypassing it.
Refs #55924#58720 (same class in cron delivery), adjacent to #90683.
#90587 replaced the bundled Nous with a GitHub-theme fork. The earlier
glass-and-cream palette stays available under its own name; the default
does not change.
The status stack hydrates the goal indicator on every session open via
/goal status. A finished goal stays status=done in the DB permanently,
so hydration re-created the '✓ Goal done' chip on every mount — the 8s
linger only applies to the live completion event, not re-hydration. In
Bot Mode (one endless session) the completed overlay never went away.
applyGoalStatusText now takes a hydrate flag: terminal (done) goals are
treated as no-goal during hydration and any lingering done chip is
dropped, while live events keep the 8s linger behavior.
Test-helper + regression test from PR #92633 by @DavidMetcalfe:
mock _is_supervised_gateway_process instead of setting the raw env
var, and add a test proving a CLI agent session with inherited
_HERMES_GATEWAY=1 but no PID ownership is no longer blocked.
The terminal tool lifecycle guard and the gateway stop/restart CLI
guards keyed on the raw _HERMES_GATEWAY=1 env marker, which every
gateway descendant inherits (and importing gateway.run sets it too).
CLI/TUI agent sessions were falsely blocked from documented gateway
management commands. Gate on _is_supervised_gateway_process() instead,
which requires owning the live gateway PID file.
Salvaged from PR #92196 (guard half) by @nbxuhk. Fixes#92560.
Follow-up to the #91701 salvage: persistent registrations survive a routine
unload-all, but must not outlive their plugin.
- Targeted unload (plugin disable/uninstall) now gathers persistent rows
from the ownership ledger and disposes them.
- Unload-all parks live persistent handles in _persistent_carryover;
discover_and_load(force=True) evicts the ones whose plugin did not
re-register the same (kind, key) — superseded handles are dropped
without disposal so a same-object re-registration stays live.
The dashboard auth registry is process-global, but a bundled auth provider
was registered under the per-home plugin manager's scope and enrolled in that
manager's reverse-order teardown. A per-home manager is unloaded routinely
(profile-scoped dashboard activity, forced re-discovery), and that teardown
disposed the registration — emptying the auth registry for the whole process
and permanently disabling sign-in until restart.
Register dashboard-auth providers in the process-global slot as persistent
host-owned registrations kept out of per-home manager teardown, so a routine
unload can no longer disable authentication process-wide. Registration upserts,
so a forced re-discovery (e.g. a password change) still rotates the provider in
place. The test-only manager reset now clears the auth registry too, since
persistent registrations deliberately survive unload.
Fixes#91701
The protected agent-instruction gate grants one operation and persists
nothing, but only the TUI/desktop and Runs transports were taught that.
The prompt_toolkit panel, the input() fallback, and the ACP editor menu
still rendered "Allow for session", so a user editing SOUL.md tapped it,
got re-prompted on the next write, and read the gate as broken.
Thread allow_session through prompt_dangerous_approval so a caller that
re-asks every time collapses every surface to once/deny, and cover the
producer-to-transport contract end to end.
When the Python interpreter begins teardown (user closes hermes, SIGTERM,
OOM-kill), every executor-backed operation raises 'cannot schedule new
futures after interpreter shutdown'. The outer except handler in
run_conversation caught this error but did not recognize it as fatal —
it kept retrying (API calls #4, #5, #6) until max_iterations, each time
hitting the same dead executor and printing another traceback.
The fix adds an early check: if sys.is_finalizing() or the error matches
the 'cannot schedule new futures' pattern, break immediately with a clean
interpreter_shutdown exit reason instead of retrying. The codebase already
had this pattern in cron/scheduler.py and agent/tool_executor.py — the
conversation loop just wasn't using it.
Drives the real `syncBotsHomeWorkspace` against a shell whose reveal
does not hand the tab its zone's active slot — modelled on
`revealTreePane`'s hidden-pane early return — and asserts the passive
reconcile re-fronts once rather than on every pass. Against the
unfixed code the first case remounts the view 21 times for 21 passes.
The other three keep the bound from becoming a regression of its own:
giving up must leave the tab OPEN (a closed home drops the Bots tab
through to the ownerless Sessions composer); a cooperative shell still
gets its legitimate re-front, and gets another one the next time the
tab is genuinely backgrounded, so the budget is per-attempt rather than
one-shot for the life of the process; and an explicit gesture re-fronts
even on a shell where passive reconciles have already given up.
Re-fronting the Bots home tab is a close followed by a re-open, which
tears down and rebuilds the entire Bots view. `openBotsHomeWorkspace`
took that path on EVERY passive reconcile that found the tab open but
not holding its zone's active slot, with nothing bounding the retries.
That condition is not always transient. `revealTreePane` returns early
for a pane in `$hiddenTreePanes` without ever activating it,
`isPaneVisible` is false for a minimized zone, and a pane the tree
never adopted has no group to be active in. Pinned in any of those
states, every signal that reaches a surface sync — sidebar visibility
flips, focus churn, group changes — bought one more full remount, and
the view visibly strobed.
A passive reconcile now gets one attempt. The reveal has already
granted or refused the active slot by the time `openWorkspace` returns,
so the budget settles on that answer directly instead of waiting for a
visibility notification that is not coming: a computed store stays
silent when the value does not change. Retiring the tab starts a fresh
budget, and an explicit gesture is never blocked.
Giving up keeps the surface rather than closing it — a closed home
drops the Bots tab through to the ownerless Sessions composer, which is
the hole the home exists to plug.
Covers the behavior, not the markup: a row exposes an activation target
that opens THAT job; the opener never contains the switch or the delete
control (a nested interactive element would swallow the toggle and is
invalid markup anyway); detail rows carry only fields the gateway
actually sent, so a job that has never run drops those rows instead of
rendering "undefined"; a paused job reports Paused and promises no next
run; the raw schedule appears only when the humanized label dropped
something; and a failing job explains itself in failure order — the run
that never happened outranks the delivery of a run that did.
`routine-owner.test.mjs` asserted the row's owner routing by matching
`function RoutineRow({ job, owner })` against the plugin source, so it
broke on a parameter addition that changed no behavior. Replaced with
the real invariant it was reaching for: toggling the switch sends
`cron.manage` for the owner that rendered the row and evicts that
owner's cache key — which a signature change cannot fake.
In the Bots pane the Cronjobs rows were inert. The only interactive
controls were the enable switch and the hover-only delete button, so
clicking a cronjob to see what it runs, when it runs next, or why it
stopped did nothing at all — while the same job on the main Cron page
opens a full detail panel.
The gateway already ships every one of those facts with
`cron.manage list` (schedule, repeat, next/last run, last status,
delivery target, model, workdir, prompt preview, and the
fire/delivery/pause failures). None of it had a surface in Bot Mode: a
job failing every run reads exactly like a healthy paused one.
The row title becomes a real button that opens a read-only inspector
rendered from the record the pane is already holding — no extra RPC,
and no second mutation path beside the row's own switch and delete. The
switch and delete button stay siblings of the opener, so a toggle can
never be swallowed by the open. The inspector tracks the job by id
rather than by object, so the 20s poll keeps an open panel live instead
of freezing the snapshot it opened with.