Commit Graph

14240 Commits

Author SHA1 Message Date
Aniruddha Adak f168d857c3 test(agent): stop truncation-warning ContextVar leaking between test files
Running `pytest tests/agent/test_prompt_builder.py
tests/agent/test_system_prompt.py` failed
test_build_system_prompt_records_stable_prefix with AttributeError:
'...SimpleNamespace' object has no attribute '_emit_status'
(#93018). A truncation warning recorded by test_prompt_builder.py stays
in the shared thread context under plain pytest, so the later file's
build_system_prompt call drains a warning and forwards it to
agent._emit_status - which the test stub lacked.

Harden both sides:

- tests/agent/test_system_prompt.py: _make_agent() stub gains a no-op
  _emit_status, so draining a stray warning is harmless.
- tests/agent/test_prompt_builder.py: autouse fixture drains pending
  truncation warnings after every test, leaving the ContextVar clean.

The order-dependent failure no longer reproduces in either ordering.
2026-08-23 18:27:12 -07:00
aniruddhaadak80 cd6c088928 test(compression): align no-op strike tests with structural backoff (#93022)
Two suites still encoded the pre-#93093 contract that the three
structural no-op branches (insufficient_messages, no_compressible_window,
empty_post_handoff_window) increment _ineffective_compression_count:

- tests/agent/test_compaction_anti_thrash.py::
  TestMinimumMessagesBranch::test_too_few_messages_records_an_ineffective_pass
- tests/run_agent/test_infinite_compaction_loop.py::
  TestCompressNoOpRegistersIneffective::{test_no_op_increments_counter,
  test_two_no_ops_block_should_compress}

Structural no-ops are transcript-shape facts, not evidence of an
incompressible floor, so they now arm _structural_no_op_backoff_until
and leave the strike counter untouched. Update the tests to pin the new
contract (count unchanged, backoff armed via time.monotonic(),
should_compress blocked while it holds) and rename accordingly. The
outcome contract of test_two_no_ops_block_should_compress is preserved:
repeated no-ops still block further automatic compression.
2026-08-23 18:27:07 -07:00
Aniruddha Adak f778c0d941 fix(compression): structural no-ops defer retries instead of striking the breaker
Fixes #93022. A short session (protection window >= transcript) hits the
"insufficient messages" / "no compressible window" branches twice and
permanently trips the anti-thrash breaker, even though nothing was
eligible to compress - compression was never attempted, so there is
nothing "ineffective" to score. The session then rides past the
threshold with no compaction possible (recovery probes only soften,
not fix, the misclassification).

Distinguish "nothing eligible right now" from "attempted and
underperformed":

- New transient _structural_no_op_backoff_until (in-memory, 300s)
  armed by _record_structural_no_op() at the three structural no-op
  sites: insufficient_messages, no_compressible_window,
  empty_post_handoff_window. No strikes accumulate; auto-compaction
  resumes on its own once the backoff lapses or the transcript outgrows
  the protection window.
- The backoff gates should_compress via
  _automatic_compression_blocked_locally and surfaces in
  _compression_block_reason as "structural_backoff:<seconds>".
- #40803's frozen-CLI guarantee is preserved: a transcript that can
  never shrink retries at most once per backoff window instead of
  every turn.
- force=True (/compress) clears an active backoff before attempting;
  record_completed_compaction() lifts it - both prove the transcript
  is compressible/being worked.
- Genuine attempted-but-underperformed verdicts still strike the
  durable ineffective counter unchanged.

Tests: new tests/agent/test_context_compressor_structural_backoff.py;
updated the two tests that asserted the old strike-on-noop behavior.
2026-08-23 18:27:07 -07:00
Aintworth ce51f535d3 test(gemini): cover nested/list/non-pointer ref cases; document false-positive tolerance
Address review feedback:
- Add tests for deeply-nested $ref (recursion), top-level JSON array
  (already wrapped, no 400 path), and $ref without '#/' prefix (stays
  structured).
- Document the deliberate structural (false-positive-tolerant) detection and
  its O(n) cost in the helper docstring.
2026-08-23 18:27:04 -07:00
Aintworth 03477166f9 fix(gemini): wrap schema-bearing tool results as opaque text
Gemini 3 resolves JSON-Schema $ref/$defs pointers inside a
functionResponse.response payload and rejects unknown references with
HTTP 400 INVALID_ARGUMENT ('referenced name #/$defs/...' does not match
a display_name; see vercel/ai#14369).

tool_describe (and any tool whose result is itself a JSON Schema) returns
schema text that previously went back as a structured response, tripping
Gemini's pointer resolution. Detect such results with a $ref-pointer scan
and wrap them as opaque text instead.

Adds regression tests for the wrap path and the unchanged structured path.
2026-08-23 18:27:04 -07:00
Teknium cce2d9418b chore(tests): remove the never-executed kanban stress/chaos suite
tests/stress/ was dead weight: its own conftest set
collect_ignore_glob = ["*.py"], so pytest has never collected a single
file from it, the advertised --run-stress flag was a permanent no-op,
and no CI workflow ever invoked the scripts (#93135). Rather than wire
a nightly lane for scripts that were never verified end-to-end, remove
the suite. Kanban concurrency behavior remains covered by the regular
tests under tests/hermes_cli/ and tests/gateway/.

Closes #93135.
2026-08-23 18:25:49 -07:00
justcarlosm 478a09c06b fix(browser): floor browser-use CLI subprocess PATH with sane system dirs
Profile-spawned workers (kanban bots, cron jobs) can inherit a PATH of
only version-manager dirs — observed in the wild as one nvm node dir
repeated 7x. The uv-installed browser-use binary is a POSIX sh
trampoline that resolves dirname/realpath through PATH, so it died
with 'realpath: not found … exec: /python: not found' (exit 127)
before its own Python ever started.

_base_subprocess_env now floors the child PATH via browser_tool's
_merge_browser_path (the agent-browser backend already guards the same
hazard), degrading to appending FHS bin dirs if that import is ever
unavailable. Windows is a no-op (.cmd shims don't trampoline).

Verified: unit tests + real uvx browser-use --version under a
nvm-only-PATH worker env, rc 127 -> rc 0.
2026-08-23 18:25:35 -07:00
liuhao1024 8f9abc9873 fix(cron): re-anchor stale next_run_at after direct jobs.json schedule edits
get_due_jobs() fires purely off the stored next_run_at <= now, with no
check that the stored instant is still an occurrence of the schedule's
current expression. A direct jobs.json edit that narrows schedule.expr
(e.g. daily "0 7 * * *" -> weekdays "0 7 * * 1-5") keeps the stored
next_run_at computed under the old expression, so the job fires on days
the new expression excludes. The within-grace fire and the catch-up
"run once now" path both inherit the wrong instant.

Add a best-effort stale-schedule guard on the fire path: when the stored
next_run_at is not an occurrence of the current cron expression,
re-anchor it via compute_next_run() from the current expression and skip
the fire. Non-cron kinds, missing expr, croniter unavailability, and
malformed input all report a match so the fire path keeps its existing
semantics. Recomputation uses the current expression, so the re-anchor
converges and cannot defer a valid job forever.

Fixes #93049
2026-08-23 18:25:35 -07:00
liuhao1024 5d8b031514 fix(stt): surface the selection-specific error for explicit openai STT
When the managed openai-audio gateway is unavailable,
_resolve_openai_audio_client_config() raises a ValueError that names the
blocker (and, for managed-Nous users, the `hermes tools` remediation).
The boolean probe in _get_provider's explicit-openai branch flattened
that into False, so the log claimed "no API key available" and the
transcription result returned the all-provider install hint -- pointing
operators at unrelated setup instead of their managed route (#93045).

Resolve the config directly in the branch so the warning names the real
blocker, and let the dispatch's "none" fallback surface the
selection-specific error for an explicit openai choice. No fallback is
added: an unavailable selection still resolves to "none", it just
reports why.
2026-08-23 18:25:35 -07:00
王雪帆 c4871226f1 fix(cli): honor target_model when resolving custom providers
resolve_runtime_provider() documents target_model as the explicit model
override for mid-session switches and auxiliary slots, but the custom
provider path (_resolve_named_custom_runtime) never received it and
silently substituted the provider's configured default_model instead.

This made auxiliary slots such as auxiliary.background_review silently
run the provider's default model rather than the configured one — e.g. an
ocx-proxy slot configured for gemini-flash actually executed
cursor/claude-sonnet-5, hitting upstream rate limits.

Pass target_model through to the custom runtime resolver and prefer it
over the provider's default model in both the pooled and non-pooled
credential paths.
2026-08-23 18:25:35 -07:00
Finn763 74e6885f0d fix(review): fail-closed compressor detachment + warm-cache first request (#93057 review)
Adversarial-review fixes for the #93057 snapshot-compaction PR:

- Fail-closed detachment: only re-enable compression after
  bind_session_state successfully severs the engine's parent binding.
  A failed rebind keeps the historical compression_enabled=False
  behavior and warns, instead of running compaction against a
  compressor still bound to the parent's SessionDB (#38727 re-open).
- Warm-cache parity: defer both compression gates (turn-prologue
  preflight + pre-API pressure check) until the fork's first provider
  response, so the first request replays the full snapshot as the
  intended cached read and compaction applies from the second request
  on — matching the documented budget mental model.
- Tests: regression for the rebind-failure fail-closed path (red on
  pre-fix code) and the existing threshold-crossing test reworked to a
  two-request review asserting the warm first request + compacted
  second request. 116 tests green across all touched suites; ruff
  clean.
2026-08-23 18:25:19 -07:00
Finn763 4202a508fd fix(review): bound same-model background review replay
Detach the review fork's compressor from the parent SessionDB/session_id
and re-enable in-memory-only compaction for oversized snapshots, instead
of the historical compression_enabled=False guard that left the fork's
replayed transcript unbounded (350k-384k input tokens per request, 1.49M
total across one 8-request review). Add an aggregate input-token budget
(auxiliary.background_review.max_input_tokens, default 600k) so repeated
tool calls cannot recreate an unbounded transcript; the tool loop stops
before the provider call that would cross it.

Closes #93057
2026-08-23 18:25:19 -07:00
fangliquanflq 2033f4cc34 fix(agent): separate cancellation diagnostics from tool output 2026-08-23 18:25:19 -07:00
fangliquanflq c1c0efa375 fix(code-exec): preserve interrupt cancellation source 2026-08-23 18:25:19 -07:00
fangliquanflq ee8a66233f test(gateway): preserve replacement handles across close races 2026-08-23 18:25:12 -07:00
fangliquanflq 80cec2785d fix(gateway): preserve routing state across recovery 2026-08-23 18:25:12 -07:00
fangliquanflq 4b659f0e33 fix(gateway): retry failed session database opens 2026-08-23 18:25:12 -07:00
Andrex Ibiza, MBA 31a01f373b fix(state): make automatic repair non-destructive
Reproduce the schema-btree failure where the in-place writable_schema/VACUUM ladder can reduce a 3,048-page canonical state.db to 113 pages and still return repaired=False.

Move all mutating strategies behind a complete SQLite online-backup snapshot, retain one exclusive SQLite guard from staging through transactional promotion, preserve committed WAL frames and the live inode, fail closed on environmental hazards, and add adversarial regression coverage for failed-repair preservation, post-stage writer races, interrupted copies, stale scratch, disk admission, attempt-ledger semantics, and durability routing.

Fixes #93064
Supersedes the delivery mechanics of #87409 while preserving its implementation provenance.

Co-authored-by: cervantesh <11169707+cervantesh@users.noreply.github.com>
2026-08-23 18:25:12 -07:00
Teknium a2a43f7e82 fix(agent): widen composite-id alias matching to the compressor; unify variant policy owners (#63000)
Follow-up on top of the salvaged #93335:

- context_compressor._sanitize_tool_pairs now expands alias spellings on
  the RESULT side too (tool_result_id_variants), so a composite
  call|item-keyed result pairs with its split-field tool_call instead of
  being dropped and its call stripped.
- The compressor's _tool_call_id_variants staticmethod and
  agent_runtime_helpers' module-level _tool_call_id_variants are now thin
  forwarders to agent.message_sanitization.tool_call_id_variants — one
  policy owner for alias expansion, so the pre-call sanitizer, repair
  pass, dedup pass, and compression sanitizer can never drift apart.
- Preserved the #91768 SDK-object tolerance in repair pass 1 (the
  shared helper handles non-dict tool_calls via getattr; the salvaged
  commit's isinstance-dict guard was dropped in the merge resolution).

New regression tests: composite-keyed results through
sanitize_api_messages (both directions) and _sanitize_tool_pairs, with
negative controls. Sabotage-verified: compressor test fails with raw
tool_call_id tracking.
2026-08-23 18:24:43 -07:00
joaomarcos 5496d5995a fix(agent): preserve tool results across ID variants
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
2026-08-23 18:24:43 -07:00
Meng Chee 92edb861be fix(cron): close NUL-padded script bypass in lifecycle guard
The #76762 binary check treats any NUL byte in the first chunk as "compiled
binary, nothing to scan":

    if b"\x00" in data:
        return None, False

"Contains a NUL" and "is a compiled binary" are different questions, and the
gap between them is a guard bypass. `bash` executes a *text* script straight
past an embedded NUL, so one pad byte disables the entire scan while the
script still runs:

    #!/bin/bash
    # pad<NUL>
    hermes gateway restart

    scan("bash padded.sh")  -> False   (not blocked)
    bash padded.sh          -> executes the lifecycle command

This shape was blocked before #76762, so the crash fix traded a loud failure
for a silent one.

Keying the check on a leading `#!` is not sufficient: a shebang-less file with
a NUL on any line but the first also executes normally. (A NUL on line 1 of a
shebang-less file is the one shape bash rejects, exit 126 — but that same file
is still executable via `. file`.)

Fix: identify binaries by MAGIC NUMBER — ELF, Mach-O (incl. byte-swapped and
universal/fat), PE/COFF, static archive, gzip, zip — with a shebang always
winning. A NUL-bearing *text* file is scanned with its NULs stripped;
stripping can only splice tokens together, never apart, so it fails closed.
File extensions are deliberately not consulted, so a suffixless shell script
is still scanned.

The size check now runs BEFORE the strip: stripping shrinks the buffer, so
checking afterwards would let an oversized file slip under the threshold and
skip the fail-closed branch. (Caught by
test_oversized_nul_bearing_text_still_fails_closed, which failed on the first
cut of this patch.)

Return values are unchanged, so this does not conflict with the in-flight
crash-class fixes to the same function.

Tests (tests/hermes_cli/test_gateway_restart_loop.py), 3 of which fail on main:

- test_nul_padded_script_is_still_scanned
- test_nul_padded_script_without_shebang_is_scanned
- test_oversized_nul_bearing_text_still_fails_closed
- test_elf_binary_is_not_scanned_as_script       (#76762 stays fixed)
- test_macho_binary_is_not_scanned_as_script     (incl. fat binary)
- test_clean_script_without_lifecycle_command_not_blocked
2026-08-23 18:24:36 -07:00
Meng Chee da30db8e8c fix(cron): scan dot-operator sourced scripts in lifecycle guard
`_iter_referenced_shell_scripts` recognises the `source` builtin so a script
pulled in with `source ./restart.sh` gets scanned for lifecycle commands. The
POSIX dot operator is the same builtin, but it was not caught:

    if executable_name in {".", "source"}:

`executable_name` is `Path(executable).name`, and `Path(".").name` is the
**empty string** -- pathlib normalises "." to the current directory, whose name
is "". So the set membership never matched for `.`, the sourced script was
never added to the reference walk, and its contents were never scanned.

Verified against current main:

    . /tmp/restart.sh        -> not blocked   (script never scanned)
    source /tmp/restart.sh   -> blocked
    bash /tmp/restart.sh     -> blocked

where /tmp/restart.sh contains a `hermes gateway restart` line. Sourcing runs
the script in the current shell, so the dot spelling is not merely equivalent
to `source` -- it is the more common form in practice.

Fix compares the raw token as well as the basename:

    if executable in {".", "source"} or executable_name == "source":

Keeping the `executable_name == "source"` arm preserves the existing behaviour
for a path-qualified spelling, while the raw-token test catches `.` without
relying on pathlib normalisation.

Tests (tests/hermes_cli/test_gateway_restart_loop.py):

- test_dot_operator_sourced_script_is_scanned -- the regression; fails on main
- test_source_builtin_sourced_script_is_scanned -- `source` stays blocked
- test_dot_operator_clean_script_not_blocked -- widening the check must not
  false-block an innocent `. ./activate.sh`

Found while auditing the guard after #76762. Scoped deliberately to this one
defect; the NUL-padded-script bypass I found in the same audit is a separate
PR.
2026-08-23 18:24:36 -07:00
Teknium b4d4167d42 fix(gateway): lazy/unpersisted resume also rebinds transport and cancels the pending reap
Live WS E2E after the #93361 merge (real web_server + tui_gateway, isolated
HERMES_HOME, 2s grace): drop socket -> re-resume stored id on a new socket
still produced a ws_orphan_reap reclaim. The lazy/unpersisted resume branch
(no state.db row yet -- every fresh Bot Chat) returned the sentinel-parked
live record without rebinding its transport or cancelling the armed reap
Timer, so the storm survived for exactly the Bot Mode sessions the cluster
targeted. The unit-covered paths (_live_session_payload, _reuse_live_response,
_claim_or_reuse_live) were all correct; this branch bypassed them.

Regression test drives the real session.resume RPC against a sentinel-parked
unpersisted record (sabotage-verified: fails without the fix). After the fix
the full live E2E passes 10/10 scenarios including a 4-cycle drop/resume storm
loop with zero reclaim broadcasts.
2026-08-23 18:24:26 -07:00
Teknium 65c58651b0 feat: review slot appears in every aux-model picker (desktop, dashboard, CLI)
Follow-up to #93339: the auxiliary.review slot existed in config but was
missing from every model-picker surface, so users could only set the
review model by hand-editing config.yaml.

- hermes_cli/web_server.py: review in _AUX_TASK_SLOTS (REST allowlist,
  stale-aux warning sweep)
- hermes_cli/main.py: review in _AUX_TASKS (hermes model aux picker)
- apps/desktop model-settings.tsx + all 5 i18n locales (en/ja/zh/
  zh-hant/ar): review slot with label/hint
- web/src/pages/ModelsPage.tsx: review row in dashboard Models page
- tests: registry-sync test pinning review across DEFAULT_CONFIG,
  _AUX_TASKS, and _AUX_TASK_SLOTS (curator pattern)
- docs: aux-task table in fallback-providers.md (en) + zh-Hans mirrors
  of fallback-providers and the delegation /review section missed in
  #93339
2026-08-23 18:22:39 -07:00
zgqq 9aa0721b23 fix: inert heredoc bodies no longer trip the gateway lifecycle guard (#88336)
Runbook prose inside a quoted-delimiter heredoc feeding a data sink
(cat > file <<'EOF') is documentation, not a command this shell will
execute. Mask provably-inert heredoc bodies (tools/shell_heredoc's
conservative stripper, already used by terminal_tool) before scanning.
Fails open on any ambiguity: executable and unquoted-delimiter heredocs
stay scanned. Salvaged from PR #88336 by @zgqq (the heredoc half; its
Branch D boundary and dir-token halves already landed/were fixed).
2026-08-23 18:10:31 -07:00
Artur Hapantsou b34edd6b01 fix: execute_code and argv-list payloads no longer bypass the gateway lifecycle guard (#68289)
execute_code lacked the lifecycle guard entirely, and Python argv-list
forms (subprocess.run([...])) separated command words with brackets and
commas the shell-shaped pattern could not see. Mirror the terminal_tool
guard in execute_code (ownership-gated per #92560) and strip argv-list
punctuation in the token-join re-scan. Salvaged from PR #68289 by
@arcimun, adapted to the ownership gate and current guard structure.
2026-08-23 18:01:59 -07:00
KeaneYan 1c791cbfe6 fix(gateway): resolve uninstall lifecycle guard conflict 2026-08-23 18:01:59 -07:00
BotUser 679e07a074 fix(gateway): close order-dependency + missing-verb gap in launchctl lifecycle guards
The gateway-lifecycle guards in cron/lifecycle_guard.py (Branch B, the
unconditional hard-block used by cron creation and the terminal tool when
_HERMES_GATEWAY=1) and tools/approval.py's launchctl rule both matched
`launchctl <verb> ... hermes[.-]?gateway` as a single sequential regex,
requiring the hermes-gateway label to appear literally AFTER the verb.

A shell command that builds the label earlier in the string — e.g. a
for-loop reading labels from a list defined before the actual launchctl
call — defeats that ordering entirely:

    for item in 'ai.hermes.gateway-apollo:...' 'ai.hermes.gateway:...'; do
      label=${item%%:*}; plist=${item#*:}
      launchctl bootout "gui/$uid/$label"
      launchctl bootstrap "gui/$uid" "$plist"
    done

The literal text "hermes.gateway" only ever appears in the for-list,
never after "bootout" — so `[^\n]*\bhermes[.\-]?gateway` never matches at
the verb's position, even though the command unambiguously targets the
gateway's own launchd label.

cron/lifecycle_guard.py's verb list also didn't include `bootout` at all
(present in tools/approval.py's list and covered by its own test suite —
`launchctl bootout ai.hermes.gateway` is explicitly asserted as dangerous
there — so the omission in the sibling file looks like list drift between
the two guards rather than an intentional exclusion).

`bootout` is the verb that actually deregisters a launchd job (unlike
kickstart/stop, which just bounce a still-registered one), so a command
using it evades both guards, then removes the service from launchd with
no supervisor left to bring it back — worse than a simple restart-loop.

We hit this for real: a gateway self-restart (triggered from a chat
request to change the default model) used a raw terminal `launchctl
bootout`/`bootstrap` loop across 4 launchd labels instead of the normal
`hermes gateway restart` path. It slipped past both guards, self-bootout
killed the process mid-drain before its own follow-up bootstrap could
run, and all 4 gateway profiles ended up fully deregistered from launchd
with zero user approval (approvals.mode: manual was configured) until
someone manually re-bootstrapped them.

Fix: both guards now check "a launchctl lifecycle verb appears somewhere
AND a hermes-gateway label appears somewhere", independent of order, and
cron/lifecycle_guard.py's verb list gains bootout/kill/disable/remove to
match tools/approval.py's existing set. Internal recovery code
(hermes_cli/gateway.py's own `subprocess.run(["launchctl", "bootout",
...])` calls) is unaffected — these guards only scan shell-command
strings composed by the agent's terminal/cron tools, not the CLI's
trusted internal subprocess argument lists.

Adds regression tests in both test files reproducing the exact incident
command (label built in an earlier for-loop segment, referenced only via
`$label` at the point of the verb).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 17:56:36 -07:00
thenabbu acf8245607 fix(tools): pass single_query_deny_message to the ssh-config write approval gate
Commit 1596148ff made single_query_deny_message a required keyword-only
parameter of _run_approval_gate() and updated its two callers inside
tools/approval.py, but missed the third caller: the SSH-config write
guard in tools/file_tools.py (_check_approval_required_write,
pattern_key="ssh_config_write").

Any gated write to an SSH client config therefore raised
  TypeError: _run_approval_gate() missing 1 required keyword-only
  argument: single_query_deny_message
instead of routing through the human-approval flow.

- Pass the kwarg with a single-query-specific deny message that points
  operators at approvals.single_query_mode: approve.
- Add a regression test asserting the gate call passes every required
  kwarg (fails on unpatched main).

Fixes #93201
2026-08-23 17:54:20 -07:00
Teknium 20e308fea7 fix: use lookbehind anchor so binary-decoded and remote-read content still scans
The separator-class anchor broke two fail-closed tests (binary bytes
decode to U+FFFD adjacent to the CLI name; remote head-c reads). A
negative lookbehind excluding path/word chars keeps the #77173 fix
while preserving every fail-closed content-scan path.
2026-08-23 17:50:48 -07:00
eaglezzz0522-cloud 180f981125 fix: lifecycle guard Branch A anchors the CLI name at command position (#77173 path false positive)
A file path with embedded spaces (/docs/... with lifecycle words in the
filename) matched Branch A via the path tail and hard-blocked innocent
commands. Anchor the CLI name at command position (start, separator, or
substitution opener). Salvaged from PR #77536 by @eaglezzz0522-cloud,
reapplied onto the current pattern with subshell coverage and tests.
2026-08-23 17:50:48 -07:00
liuhao1024 51239e8e2a fix(vision): forward the API key to the server-type probe and cache failed verdicts
The image-routing vision path calls detect_local_server_type without
the provider's API key. Against a remote API-keyed endpoint (sglang /
vLLM with --api-key) every leg of the 5-request probe waterfall came
back 401 — and because a failed verdict was never written to the
in-memory cache (only positive verdicts were), the waterfall re-ran on
EVERY image-bearing turn (#89863: 51 detail-less busy-acks observed in
one Slack channel while the probe sprayed the user's own server).

Two changes:

- image_routing._should_probe_ollama_vision now takes the API key and
  forwards it; a new _resolve_inference_api_key mirrors
  _resolve_inference_base_url's resolution order (runtime value,
  model.api_key, providers blocks) so the key always matches the URL
  being probed.

- detect_local_server_type caches a None verdict in memory with a short
  failure TTL (5 min, vs 1h for positives) so the next turn is served
  from the negative entry instead of re-running the waterfall — while
  a transient failure (server starting, key being fixed) recovers in
  minutes. Negative verdicts are deliberately not written to the
  cross-process disk cache.
2026-08-23 17:47:50 -07:00
JinUltimate 4ca993c746 fix(image_routing): stop fingerprint-probing remote OpenAI-compatible endpoints
Fixes #89863. With a custom: provider pointing at a remote, API-keyed
endpoint (sglang/vLLM/OpenAI-compat), every image turn triggered a 5-request
probe waterfall without Authorization, spraying 401s at the backend.

Two fixes:

1. _should_probe_ollama_vision now takes api_key and forwards it to
   detect_local_server_type so keyed local servers don't 401.

2. When provider != 'ollama', remote endpoints (per is_local_endpoint) are
   rejected early — server-fingerprint probing is only valid for local
   boxes. Non-Ollama remotes expose Ollama-compat endpoints that can
   misidentify and trigger unnecessary /api/show probes.

_lookup_supports_vision resolves the runtime api_key via
_runtime_main_value and forwards it to both helpers. New test class
TestShouldProbeOllamaVision covers the contract in both directions.
2026-08-23 17:47:50 -07:00
honor2030 fe483de4d3 fix(agent): keep max-iteration warnings out of quiet stdout
Route the max-iterations diagnostic through logging when quiet_mode is active so automation wrappers keep stdout machine-readable.

Add a regression test covering quiet max-iteration summary handling.
2026-08-23 17:45:58 -07:00
Teknium 9b3f60c029 fix(gateway): resolve approval.respond by durable identity before failing 4001
Server half of #91684: the desktop can answer an approval prompt with a
stale live sid — its runtime record was re-minted after a reconnect while
the prompt stayed on screen. approval.respond now falls back, on 4001
only, to resolving the target session (1) by the unique approval
request_id across every live session's pending gateway approvals, then
(2) by treating session_id as a STORED session id mapped to its live
runtime record. Only when neither resolves does it return 4001.

Tests: request_id fallback, stored-id fallback, and 4001 when nothing
resolves.
2026-08-23 17:43:39 -07:00
Teknium fdd8d75ba0 fix(gateway): make ws keepalive and orphan-reap grace config-driven (#79635)
- New dashboard.ws_ping_interval / dashboard.ws_ping_timeout defaults
  (20.0/20.0) in DEFAULT_CONFIG; hermes_cli/web_server.py reads them for
  non-loopback binds. Loopback keeps ws_ping=None (event-loop stalls must
  never kill a healthy local connection).
- New dashboard.ws_orphan_reap_grace_s (20.0): tui_gateway/server.py's
  _WS_ORPHAN_REAP_GRACE_S now resolves from config via
  _resolve_ws_orphan_reap_grace(); the HERMES_TUI_WS_ORPHAN_REAP_GRACE_S
  env var is kept as an internal override for backward compat and wins
  when set.
- tests/test_ws_keepalive_config.py: real load_config against a temp
  HERMES_HOME yaml — defaults, propagation, deep-merge, env override,
  invalid-value fallback.
2026-08-23 17:43:39 -07:00
Teknium 4aa162b30d fix(gateway): cancel pending WS-orphan reaps on resume and supersede stale runtimes quietly
Storm killer for the reap->broadcast->auto-re-resume feedback loop:

- New _pending_ws_reaps registry (sid -> Timer): _schedule_ws_orphan_reap
  registers, _reap pops, and _cancel_ws_orphan_reap(sid) is called from
  every resume/reuse/rebind path — the session.resume fast-path reuse
  (methods_session.py), _claim_or_reuse_live winners, and the
  _live_session_payload live-transport rebind.
- When a resume mints a fresh runtime for stored session id S, any prior
  runtime for S still parked on the detached-WS sentinel is claimed under
  the resume lock, its reap Timer cancelled, and the record finalized
  quietly with end_reason superseded_by_resume — NOT in
  _RECLAIM_END_REASONS, so no session.reclaimed broadcast fires and the
  client's auto-re-resume can't storm.
- superseded_by_resume added to _RECOVERABLE_END_REASONS in
  hermes_state_common.py so canonical Bot Chat resurrection still applies.

Unit tests: resume cancels the reap timer, superseded runtimes finalize
without a reclaimed broadcast, and the normal orphan reap still fires
when nobody re-resumes.
2026-08-23 17:43:39 -07:00
Ayush Nangia f2dbd37ef9 fix(gateway): re-bind session transport to a surviving window on pop-out close
Live sessions hold one transport; a pop-out window's session.resume
rebinds it, and on pop-out close the disconnect path parked the session
on the drop sentinel — the original window never received stream events
again until a manual re-resume (#83716).

Sessions now track every transport that has shown them (viewers, stamped
in _live_session_payload). _close_sessions_for_transport re-binds to the
most recent surviving viewer instead of detaching when one exists; dead
viewers are filtered; the drop sentinel + grace reap remain the path for
the last viewer. Root cause and repro by CharlesR-sudo on #83716.
2026-08-23 17:43:39 -07:00
Kyzcreig 14b50f5edd fix(tui-gateway): interrupt turns after websocket disconnect
After the existing reconnect grace, route a still-detached running session through the same interrupt mechanism as session.interrupt. Preserve delegation deferral, sidecar teardown, partial history, and single-owner reap semantics.

Verified on upstream main: RED 4 failed/2 passed without production changes; GREEN 603 related gateway/compute-host tests. Ruff and py_compile passed. Momus pass 2: APPROVE.
2026-08-23 17:43:39 -07:00
686f6c61 30a37668d9 fix(tui): reschedule WS orphan reap while a turn is still running
The grace timer treated mid-turn detached sessions as not-orphaned and
returned without arming another timer. Eviction paths also skip
running sessions, so the in-memory agent leaked until process restart.

Fixes #85578
2026-08-23 17:43:39 -07:00
Teknium 12395e57b4 feat: /review command — independent reviewer subagent on every surface
/review takes the last 10 chat messages plus optional instructions,
spawns a full-privilege background subagent (the async delegation
rail) that investigates the referenced work (PR, code, docs), and its
complete review re-enters the spawning session as a normal
async-delegation completion the primary agent can act on.

- agent/review_engine.py: shared engine (snapshot, briefing,
  auxiliary.review credential resolution, dispatch, note formatting)
- tools/delegate_tool.py: internal credentials_cfg per-call override
  (never model-facing) resolved through the same credential system as
  delegation.provider pins
- auxiliary.review config block (provider/model/base_url/api_key/
  api_mode); provider auto + empty model = inherit the main model
- Surfaces: CLI process_command, gateway run.py dispatch +
  slash_commands handler (binds the approval session key so the
  completion routes back), TUI/Desktop live dispatch in
  tui_gateway/server.py, CommandDef registry (+Slack /hermes-only cap)
- Docs: delegation.md section + slash-commands.md (both tables)
- Tests: 15 engine tests (sabotage-verified: credentials_cfg and
  dispatch tests fail without the fix), 4 gateway handler tests
  through the real async rail
2026-08-23 17:38:38 -07:00
Teknium 637716755c fix(cli): -Q stdout carries only the final response — no tool diffs, spinner lines, or reasoning
Widens the cherry-picked reasoning-callback fix to the whole leak class
(#93220):

- quiet branch also neutralizes tool_progress_callback,
  tool_start_callback, tool_complete_callback (inline diff rendering via
  render_edit_diff_with_delta was gated by NEITHER quiet_mode nor
  tool_progress_mode) and syncs agent.tool_progress_mode='off'.
- _should_emit_quiet_tool_messages() returns False under
  suppress_status_output: with callbacks neutralized, the quiet-mode
  KawaiiSpinner fallback printed '[tool]'/'[done]' lines into captured
  stdout. Also covers oneshot.py and background-review forks, which set
  the same flag and expect strict silence.

E2E (isolated HERMES_HOME, live model, write_file turn): base leaks
'┊ review diff' + full SVG source into stdout; head emits exactly the
final response. Regression tests pin the quiet-branch statements and the
gate (sabotage-verified).

Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
2026-08-23 17:01:31 -07:00
Teknium faa2399e2b fix(agent): make the pre-call dedup pass variant-aware; widen batch regression coverage (#93251)
Follow-up on top of the salvaged cluster: sanitize_api_messages step 3
(duplicate tool_call_id dedup) still tracked only the coalesced
(call_id||id) value in outstanding_call_ids, so after step 2's
variant-aware matching preserved a result keyed on the OTHER id variant,
step 3 deleted it as answering no outstanding call — whole parallel
batches of real results vanished with no stub at all (#93251's total-loss
mode). Track the full variant set per call and consume all siblings when
answered, preserving #58327 duplicate protection and llama.cpp
constant-id re-arm semantics.

Also aligns the #58287 compressor test with the in-flight tool chain
protection (#79278) that landed after that PR was opened: a trailing
user turn keeps the negative-control assistant message out of the
protected trailing window.

New regression tests: divergent-id batch survival through the dedup
pass, sibling-id replay still dropped, constant-id re-arm preserved.
Sabotage-verified: tests fail with the old single-id tracking.
2026-08-23 17:01:20 -07:00
joaomarcos 36b4da5489 fix: repair_message_sequence drops tool results for SDK tool_call objects
The tool_call id-matching pass in repair_message_sequence only read
`.get("id"/"call_id")` on plain dicts, skipping non-dict tool_calls
entirely (`if not isinstance(tc, dict): continue`). Host-fed and
pre-serialization histories can carry unserialized SDK tool_call
objects (e.g. `ChatCompletionMessageToolCall`) instead of dicts, which
left `known_tool_ids` empty for that assistant turn. The following
`tool` message — a legitimate result already produced by executing the
tool — was then misclassified as an orphan and silently dropped,
corrupting the persisted conversation history and leaving the
assistant's tool_calls unanswered (itself a trigger for HTTP 400 on
strict providers).

Fix: extract id/call_id via getattr() for non-dict entries too,
mirroring AIAgent._get_tool_call_id_static's existing dict-or-object
tolerance, instead of skipping them.
2026-08-23 17:01:20 -07:00
Frowtek b9a62f6590 fix(agent): consume every tool_call id variant when pairing tool results
`repair_message_sequence` registers BOTH `id` and `call_id` for each
assistant tool_call, because a matching tool result may be keyed on either
depending on which path built it (#58168). The duplicate guard added for
dropped rather than replayed.

Those two behaviours don't compose: a Codex/Responses tool_call registers
two DIFFERENT ids (`fc_...` and `call_...`), but only the id the first
result referenced is discarded. Its sibling stays in `known_tool_ids`, so a
duplicate result keyed on that sibling still matches and is kept — two tool
messages replayed for one call, which is exactly the HTTP 400 on strict
providers the consume step exists to prevent.

Duplicates of this kind come from the retry / crash / session-resume glitch
the guard was written for; the id-variant split just lets them slip past it.

Track each registered id back to its tool_call's full variant set and
discard all of them on a match. Results keyed on either variant are still
accepted (no false orphaning), and two parallel Codex calls answered via
different variants both survive.

Adds regression tests for the sibling-keyed duplicate and for the
two-calls/mixed-keys case that must NOT be affected.
2026-08-23 17:01:20 -07:00
srojk34 52fb5081cc fix(compression): register both id/call_id variants in _sanitize_tool_pairs
_sanitize_tool_pairs() matched tool_call/tool_result pairs using a
single-value call_id||id precedence per tool_call (_get_tool_call_id).
In the Codex Responses API format an assistant tool_call carries both a
distinct id (fc_...) and call_id (call_...); a tool result's
tool_call_id may be keyed on either depending on which code path built
it. Whenever a genuinely matching pair used the field the precedence
didn't pick, the sanitizer misclassified it as orphaned on BOTH sides:
it dropped the valid tool result AND stripped the tool_call from the
assistant message, even though neither was orphaned.

Live-verified before the fix: {"id": "fc_777", "call_id": "call_777"}
+ a tool result with tool_call_id="fc_777" (a valid pair) was fully
removed by current main.

Register both id and call_id as valid match keys via a new
_tool_call_id_variants() helper (a set per tool_call, not a single
value), matching #58168's fix for repair_message_sequence's known-id
set today. A tool_call now survives if ANY of its id variants has a
matching result, which is not vulnerable to precedence order at all
(unlike swapping which field is checked first, which only trades which
sub-case is broken).

Note on #56425 (open, unreviewed): that PR touches this same function
for the same underlying issue (#55626) by swapping the call_id||id
precedence to id||call_id. That fixes the specific case where a result
matches `id` but not the reverse case (a result matching `call_id`
while `id` is also present) -- the precedence-swap approach cannot fix
the class, only relocate which sub-case is broken. This fix instead
mirrors the already-merged #58168 pattern (register the superset of
both ids as valid matches), which has no such blind spot. Adds 2
regression tests: the previously-mismatched case, and a negative
control confirming genuine orphans are still stripped alongside a
valid dual-id pair in the same window.
2026-08-23 17:01:20 -07:00
Bartok9 1a83b1e588 fix(agent): keep tool results keyed on a tool_call's id variant (#55626)
Register every id variant (call_id AND id) of each assistant tool_call in
sanitize_api_messages so a tool result keyed on either variant is treated
as paired. Previously only the coalesced (call_id||id) value was
registered, so Responses-style tool_calls carrying divergent id (fc_...)
and call_id (call_...) had their real results dropped as orphans and
replaced with '[Result unavailable]' stubs.

Cherry-picked from PR #56148 (unrelated busy_ack_templates files dropped
per the author's own follow-up commit).
2026-08-23 17:01:20 -07:00
Axel Vanni 6d501c2958 fix(cron): make gateway lifecycle matching shell-token aware (#80269)
The hard block matched raw command text, but a shell resolves quote
splicing (`kick"start"`) and backslash escaping (`kick\start`) into the
literal verb before execution. So `launchctl kick"start" -k
gui/501/ai.hermes.gateway` ran exactly as the blocked `kickstart` form
while both the non-bypassable block and the approval detector missed it —
leaving an approval-bypassing gateway self-lifecycle operation reachable.

contains_gateway_lifecycle_command now runs a second pass over
shlex-tokenized command segments, where quotes and escapes are already
resolved. It stays anchored on a hermes-gateway identifier, so prose and
non-gateway hermes services are unaffected. Because this function is the
single choke point _contains_unsafe_gateway_action calls at every
recursion level, referenced-script and `sh -c` payload scanning inherit
the fix.

tools/approval.py had the same gap for quote splices: backslash escapes
are stripped by _normalize_command_for_detection, but quote splicing in an
ARGUMENT position is not touched by _deobfuscate_shell_word_for_detection
(scoped to command-position words, deliberately — widening it would let
quoted prose match the destructive patterns). It now delegates to the
fixed guard as a last check, so an ordinary pattern match still wins and
keeps its more specific reason string.

Tests: quoted, single-quoted and backslash-spliced verbs across the
launchctl/systemctl/hermes branches, the spliced gateway identifier
itself, a splice nested in an `sh -c` payload (resolves one level deeper,
asserted at the recursive entry point terminal_tool actually calls), plus
negative cases proving prose and non-gateway labels stay unblocked.

Verified on Windows: no regressions — the 10 remaining failures across
tests/tools/test_approval.py, tests/hermes_cli/test_gateway_restart_loop.py
and tests/cron are identical on the unmodified baseline (POSIX file modes,
symlink privileges, and /bin/bash script paths).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 16:28:56 -07:00
Axel Vanni 320d884d88 fix(cron): cover bootout/remove/disable in the gateway lifecycle guard
Branch B of _GATEWAY_LIFECYCLE_PATTERN enumerated launchd verbs but omitted
`bootout` - the modern replacement for the `unload` it already listed, and
the paired inverse of the `bootstrap` it already listed. `remove` (legacy
sibling of bootout) and `disable` (what makes an unload durable) were
missing for the same reason.

This matters because the two enforcement layers are not interchangeable. In
tools/terminal_tool.py under _HERMES_GATEWAY == "1":

  - the cron.lifecycle_guard hard block is documented as applying
    unconditionally ("force=True cannot help here")
  - detect_dangerous_command below it is explicitly skipped when force=True

detect_dangerous_command already flags all three verbs, so the default path
was covered - but with force=True inside the gateway they reached execution
while stop/unload/kickstart did not. SIGTERM then propagates to the child
before the command completes and the service may never come back, which is
the state described in #74973.

The label anchor (\bhermes[.\-]?gateway) is unchanged, so unrelated services
such as `launchctl bootout gui/501/ai.hermes.update-checker` stay runnable.

Adds TestLifecycleGuardLaunchctlParity, which pins the one-directional
invariant: anything the bypassable approval layer flags, the unbypassable
hard block must also catch. Deliberately not equality - the hard block is
legitimately stricter (it also covers load/restart, which the approval layer
leaves alone). Verified failing on the parent commit for exactly bootout,
remove and disable.

Closes #80260

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 16:28:56 -07:00
webtecnica c595d3564a fix(cron): block profile-flag gateway restart/stop when self-targeting (#78028) 2026-08-23 16:24:03 -07:00