Review follow-up on #84203. Both points reproduce; neither was a regression
from the first pass, but both are live bypasses of the same guard.
**Privilege and namespace wrappers were missing.** The allowlist covered the
coreutils-shaped wrappers but not the privilege ones, so each of these ran a
lifecycle script straight past the walk:
pkexec bash ~/restart.sh
runuser -u root -- bash ~/restart.sh
setpriv --reuid=0 -- bash ~/restart.sh
systemd-run --scope bash ~/restart.sh
nsenter --target 1 --mount bash ~/restart.sh
unshare -r bash ~/restart.sh
Added `pkexec`, `su`, `runuser`, `setpriv`, `systemd-run`, `nsenter` and
`unshare`, each with the value-taking options that would otherwise be
mistaken for the command (`nsenter -t 1`, `systemd-run -p X=1`,
`runuser -u root`, …).
**An option can carry a command STRING, not an argv tail.** `env -S` and
`su`/`runuser` `-c` take shell source. The peel treated the operand as an
opaque value and skipped it, so `env -S 'bash ~/restart.sh'` was never
scanned — the string went unread rather than being recursed into.
`_STRING_COMMAND_OPTIONS` now names those options and their values are
re-scanned as shell source, the same treatment `sh -c` payloads already get.
They are read at the ORIGINAL command token, before the transparent-prefix
peel, because peeling past `su`/`env` would discard the very option carrying
the command. `--opt value` and `--opt=value` are both handled.
Scope, stated plainly: this is an enumerated allowlist, not a general
solution to "wrapper that execs its tail". A wrapper outside the set, or a
value-taking option outside these tables, still resolves to no reference —
that fails open, exactly as it did before this PR, and it is a miss rather
than a false block. The reviewer offered "extend the set with tests, or
document that the list is heuristic"; this does the first and states the
second.
Tests: 23 new cases (220 in the file) — every added wrapper against a script
reference including the value-operand option forms, both command-string
option spellings for env/su/runuser, and the same wrappers around ordinary
work (`pkexec systemctl status nginx`, `su -c 'ls -la'`, `env -S 'echo hi'`,
`nsenter -t 1 -m ps aux`) which must stay allowed. 15 fail on the tree
before this commit.
False positives re-checked at scale: the 9,258 command lines from this
repo's own scripts and docs give an identical verdict set before and after —
0 new false positives, 0 lost detections, 0 exceptions.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
`_mask_data_sink_arguments` exempts lifecycle text living in the arguments of
executables that cannot run them (`grep`, `rg`, `journalctl`, `sqlite3`, …),
so hunting for a restart string in logs is diagnostics rather than a command.
The exemption is dropped when an argument looks like an escape back into
execution — including anything starting with a dot, because sqlite3 spells
its escapes as dot-commands (`.shell`, `.system`).
But `.`, `./x` and `../x` are ordinary path operands, and
grep -r 'systemctl restart hermes-gateway' .
is the most ordinary recursive search there is. The leading-dot test treated
its `.` operand as a sqlite3 escape, disabled masking for the whole segment,
and blocked the command outright — the exact false-positive class the
exemption exists to prevent, on the shape most likely to hit it. Searching a
relative subdirectory (`./logs`, `../archive`) fails the same way, as does a
relative sqlite3 database path (`sqlite3 ./stats.db "SELECT ..."`).
Require a dot followed by a NAME character (`^\.[A-Za-z]`) so a dot-command
still defeats the exemption while a relative path stays a path. A dotfile
operand (`.env`) still reads as a dot-command — conservative, and unchanged
from today's behavior.
This narrows a security guard in the permissive direction, so the escape
hatches are pinned explicitly: with a relative-path operand present,
`.shell`/`.system`, psql's `\!`, a pipe into `sh`/`bash`/`sudo sh`/`xargs`,
command substitution, and a `;`/`&&` continuation all still block. Only the
segment's own data arguments are masked, and only when nothing in it can
reach execution.
Tests: 18 new cases in tests/hermes_cli/test_gateway_restart_loop.py — the
relative-path shapes that must now be allowed, plus the ten escape-hatch
shapes that must still block. The allow cases fail on the unfixed tree.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
`sudo`, `env`, `nohup`, `timeout` and friends exec their argument tail, so
the command that actually runs sits further right. Three guards read only the
first token of a segment, saw the wrapper, and never inspected what it runs:
bash ~/restart.sh → blocked
sudo bash ~/restart.sh → allowed
launchctl submit -l com.x -- helper → blocked
sudo launchctl submit -l com.x -- helper → allowed
Same foot-gun, one word of prefix. That reaches both enforcement points —
`cron.jobs.create_job` and `tools/terminal_tool.py` under `_HERMES_GATEWAY=1`
— and defeats the label-independent submit block that #62891 added precisely
because a persistent helper is the indirect route to a restart loop.
`_peel_transparent_prefixes()` walks past a bounded chain of these wrappers,
skipping their own options, their value-taking options (`sudo -u deploy`,
`stdbuf -o0`), `VAR=value` assignments, a `--` end-of-options separator, and
`timeout`'s duration operand, then returns the index of the real command. It
is applied to the referenced-script walk, the `sh -c` payload walk, and the
`launchctl submit`/`bootstrap` block.
In the referenced-script walk the peel is ADDITIVE — the segment is read at
the original token and again at the peeled one — because peeling must never
remove a reference the un-peeled read would have found. A local script named
`./timeout` is a script, not the coreutils wrapper, and consuming it as a
prefix would have silently stopped scanning it. (The other two call sites
need no such care: no wrapper name is also a shell name or `launchctl`, so
peeling there can only add.) That split is why the per-index logic now lives
in `_references_at()`.
This is not a new reading of shell syntax for this module — `_PIPE_TO_INTERPRETER`
already treats `sudo ` as transparent for the pipe case (`... | sudo sh`).
This generalises the same reading to the command position.
Deliberately NOT applied to the data-sink masking in
`_mask_data_sink_arguments`: peeling there would widen an exemption, and the
conservative reading is the safe one.
No false positives: peeling only changes which token is treated as the
command, so a wrapper around ordinary work resolves to a non-shell executable
and yields nothing, exactly as before (`sudo apt-get update`,
`timeout 60 curl ...`, `nice -n 10 make -j4`, a bare `env`).
Tests: 40 new cases in tests/hermes_cli/test_gateway_restart_loop.py — every
wrapper form against a script reference, a dot-source, a nested `sh -c`
payload and `launchctl submit`, plus the benign wrapped commands, a wrapped
clean script, and the `./timeout`-style lookalike names that pin the additive
reading.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
Harden the id-keyed-map flatten with an id-preserving merge:
{**value, "id": value.get("id") or key} — an inline "id" wins,
otherwise the map key is adopted (external tools often key by id and
omit the inline copy; plain list(values) would emit id-less records
that collide or get dropped downstream). Non-dict junk values are
skipped with a warning instead of crashing the load. The self-heal
rewrite persists the id-merged, junk-free records.
Tests: key adopted when no inline id (and inline id wins over a
differing key), non-dict junk skipped with warning + list_jobs
survives + self-heal persists only valid records, all-junk map
flattens to [].
Layer on the load-boundary flatten: when load_jobs() encounters an
ID-keyed jobs map ({"jobs": {"<job_id>": {...}, ...}} — written by
external tools or hand edits, never by save_jobs()), it now not only
flattens to the list contract but persists the canonical
{"jobs": [...]} form back to disk via the existing auto-repair path
(save_jobs), so the store self-heals and subsequent reads are
idempotent.
Note: _peek_jobs_unlocked() intentionally does NOT tolerate the dict
shape — it returns None so the save path never shrink-merges against
an unrepaired baseline. The flatten + repair live only at the
load_jobs() boundary.
Regression tests cover the flatten, the reported list_jobs() traceback
path, idempotent on-disk repair, and the empty-map edge case.
Salvaged from PR #92994.
Co-authored-by: a-yeyang <88581400+a-yeyang@users.noreply.github.com>
The ordinary-long-path early return now records why no resolver ran, so
the ResolvedPathReport stays truthful instead of silently inheriting a
stale value from an earlier call in the same session. Diagnostics hunk
taken from PR #93100.
Co-authored-by: aniruddhaadak80 <aniruddhaadak80@users.noreply.github.com>
ConvertTo-LongPath short-circuits for ordinary long paths (no ~\d alias),
so $script:LastResolver is only assigned when a short path actually needs
expansion. The ResolvedPathReport block read it unconditionally, which is
fatal under Set-StrictMode before any install stage runs (#93017: fresh
installs died at line 367 through three different invocation styles).
Initialize it to 'none' — the resolver's own value for "nothing ran" — at
script scope before Set-LongProfileEnvVars can invoke a resolver.
The CUA spawn-env tests froze PATH == '/usr/bin:/bin' verbatim.
_sanitize_subprocess_env now (intentionally) prepends the hermes
console-script dir for all sanitized children (#92998), so these
assertions flip to the contract: original entries preserved as
suffix, hermes bin dir first when prepended.
Exercises the real build_subprocess_env()/_resolve_hermes_bin_dir chain (no
helper mocks) under a simulated systemd/cron minimal PATH, the exact call
path cron/scheduler._run_job_script uses. Companion to #93082.
Follow-up to the ContextVar-leak fix: the autouse fixture now drains on
both sides (drain(); yield; drain()) so earlier files can't pollute this
file's assertions either.
Running `pytest tests/agent/test_prompt_builder.py
tests/agent/test_system_prompt.py` failed
test_build_system_prompt_records_stable_prefix with AttributeError:
'...SimpleNamespace' object has no attribute '_emit_status'
(#93018). A truncation warning recorded by test_prompt_builder.py stays
in the shared thread context under plain pytest, so the later file's
build_system_prompt call drains a warning and forwards it to
agent._emit_status - which the test stub lacked.
Harden both sides:
- tests/agent/test_system_prompt.py: _make_agent() stub gains a no-op
_emit_status, so draining a stray warning is harmless.
- tests/agent/test_prompt_builder.py: autouse fixture drains pending
truncation warnings after every test, leaving the ContextVar clean.
The order-dependent failure no longer reproduces in either ordering.
Two suites still encoded the pre-#93093 contract that the three
structural no-op branches (insufficient_messages, no_compressible_window,
empty_post_handoff_window) increment _ineffective_compression_count:
- tests/agent/test_compaction_anti_thrash.py::
TestMinimumMessagesBranch::test_too_few_messages_records_an_ineffective_pass
- tests/run_agent/test_infinite_compaction_loop.py::
TestCompressNoOpRegistersIneffective::{test_no_op_increments_counter,
test_two_no_ops_block_should_compress}
Structural no-ops are transcript-shape facts, not evidence of an
incompressible floor, so they now arm _structural_no_op_backoff_until
and leave the strike counter untouched. Update the tests to pin the new
contract (count unchanged, backoff armed via time.monotonic(),
should_compress blocked while it holds) and rename accordingly. The
outcome contract of test_two_no_ops_block_should_compress is preserved:
repeated no-ops still block further automatic compression.
Fixes#93022. A short session (protection window >= transcript) hits the
"insufficient messages" / "no compressible window" branches twice and
permanently trips the anti-thrash breaker, even though nothing was
eligible to compress - compression was never attempted, so there is
nothing "ineffective" to score. The session then rides past the
threshold with no compaction possible (recovery probes only soften,
not fix, the misclassification).
Distinguish "nothing eligible right now" from "attempted and
underperformed":
- New transient _structural_no_op_backoff_until (in-memory, 300s)
armed by _record_structural_no_op() at the three structural no-op
sites: insufficient_messages, no_compressible_window,
empty_post_handoff_window. No strikes accumulate; auto-compaction
resumes on its own once the backoff lapses or the transcript outgrows
the protection window.
- The backoff gates should_compress via
_automatic_compression_blocked_locally and surfaces in
_compression_block_reason as "structural_backoff:<seconds>".
- #40803's frozen-CLI guarantee is preserved: a transcript that can
never shrink retries at most once per backoff window instead of
every turn.
- force=True (/compress) clears an active backoff before attempting;
record_completed_compaction() lifts it - both prove the transcript
is compressible/being worked.
- Genuine attempted-but-underperformed verdicts still strike the
durable ineffective counter unchanged.
Tests: new tests/agent/test_context_compressor_structural_backoff.py;
updated the two tests that asserted the old strike-on-noop behavior.
Address review feedback:
- Add tests for deeply-nested $ref (recursion), top-level JSON array
(already wrapped, no 400 path), and $ref without '#/' prefix (stays
structured).
- Document the deliberate structural (false-positive-tolerant) detection and
its O(n) cost in the helper docstring.
Gemini 3 resolves JSON-Schema $ref/$defs pointers inside a
functionResponse.response payload and rejects unknown references with
HTTP 400 INVALID_ARGUMENT ('referenced name #/$defs/...' does not match
a display_name; see vercel/ai#14369).
tool_describe (and any tool whose result is itself a JSON Schema) returns
schema text that previously went back as a structured response, tripping
Gemini's pointer resolution. Detect such results with a $ref-pointer scan
and wrap them as opaque text instead.
Adds regression tests for the wrap path and the unchanged structured path.
The toggle was gated on having 2+ registered sources, which hid it in exactly
the local-only state the drift produces — the state where a user most needs to
change what launch restores.
Reconciliation repairs the drift at its source, but it can still fail to
persist (read-only or full userData), which leaves a window live on a source
the registry cannot name. $activeConnectionId is null there, the preferred-id
guard misses, and the restore re-homes a working connection.
Return early when a connection is live but unnameable. The registry has no
claim on a source it does not know about.
migrateV1ToRegistry runs exactly once, only when connections.json is absent.
A user who was local at that moment and pointed Settings -> Gateway at a
remote afterwards gets a live remote the registry cannot name: the descriptor
resolves to no connectionId, primary still says 'local', and the boot-time
launch pick force-switches the window onto a fresh local backend seconds after
the sessions list paints. That backend has no provider, so onboarding pops.
Reconcile on read: when the v1 global route names a remote with no matching
registry entry, register it and adopt it as primary/last-used, then persist so
the repair happens once. Narrow on purpose — an already-registered route is
left alone even when primary names something else, because that is the user's
pick in the Connections panel, not drift.
Replaces the hand-edit-connections.json workaround users have been trading.
The showAllProfiles browse-mode flag is persisted to localStorage, but
every restart it was force-collapsed anyway: initializeConnectionsRegistry
restores the last-used source via selectConnection, and selectConnection's
post-activation path unconditionally ran $showAllProfiles.set(false).
That collapse is correct for a user click on the connection picker (a
concrete-source action), but the silent boot restore is not a user action.
Gate both reset sites on pendingTarget === null && activeConnectionId ===
null (the fresh-boot state) so the persisted preference survives restart,
while any user-initiated switch still collapses browse mode.
Regression tests cover both directions: boot restore preserves true, a
user switch collapses it.
Fixes#93197
The salvaged hardening tests matched main.ts source text to assert that
fetchConnectionStatus reaches for a bearer and that Apply preflights before
persisting. A rename breaks them while a real auth regression that keeps the
substrings passes.
Make the preflight a first-class option on applyConnectionConfigAtomically so
its ordering is observable, and assert it through the seam: preflight runs
before either write, and a rejected preflight leaves both stores and the
activation untouched.
tests/stress/ was dead weight: its own conftest set
collect_ignore_glob = ["*.py"], so pytest has never collected a single
file from it, the advertised --run-stress flag was a permanent no-op,
and no CI workflow ever invoked the scripts (#93135). Rather than wire
a nightly lane for scripts that were never verified end-to-end, remove
the suite. Kanban concurrency behavior remains covered by the regular
tests under tests/hermes_cli/ and tests/gateway/.
Closes#93135.
Salvage follow-up for #88993: the preflight was defined but never
called from the prerequisites stage or the full-install path, so the
Fedora node-gyp failure (#93063) would still occur. Wire it in after
check_node in both sequences.
npm install inside install_node_deps() builds native addons (e.g. node-pty) via node-gyp, which needs a C/C++ compiler. That was never checked, so a missing g++ only surfaced as a generic "npm install failed or timed out" deep inside npm's own output — and because install_node_deps failing short-circuits the rest of main() via `|| return`, users end up with no `hermes` command and no clue why.
Profile-spawned workers (kanban bots, cron jobs) can inherit a PATH of
only version-manager dirs — observed in the wild as one nvm node dir
repeated 7x. The uv-installed browser-use binary is a POSIX sh
trampoline that resolves dirname/realpath through PATH, so it died
with 'realpath: not found … exec: /python: not found' (exit 127)
before its own Python ever started.
_base_subprocess_env now floors the child PATH via browser_tool's
_merge_browser_path (the agent-browser backend already guards the same
hazard), degrading to appending FHS bin dirs if that import is ever
unavailable. Windows is a no-op (.cmd shims don't trampoline).
Verified: unit tests + real uvx browser-use --version under a
nvm-only-PATH worker env, rc 127 -> rc 0.
get_due_jobs() fires purely off the stored next_run_at <= now, with no
check that the stored instant is still an occurrence of the schedule's
current expression. A direct jobs.json edit that narrows schedule.expr
(e.g. daily "0 7 * * *" -> weekdays "0 7 * * 1-5") keeps the stored
next_run_at computed under the old expression, so the job fires on days
the new expression excludes. The within-grace fire and the catch-up
"run once now" path both inherit the wrong instant.
Add a best-effort stale-schedule guard on the fire path: when the stored
next_run_at is not an occurrence of the current cron expression,
re-anchor it via compute_next_run() from the current expression and skip
the fire. Non-cron kinds, missing expr, croniter unavailability, and
malformed input all report a match so the fire path keeps its existing
semantics. Recomputation uses the current expression, so the re-anchor
converges and cannot defer a valid job forever.
Fixes#93049
When the managed openai-audio gateway is unavailable,
_resolve_openai_audio_client_config() raises a ValueError that names the
blocker (and, for managed-Nous users, the `hermes tools` remediation).
The boolean probe in _get_provider's explicit-openai branch flattened
that into False, so the log claimed "no API key available" and the
transcription result returned the all-provider install hint -- pointing
operators at unrelated setup instead of their managed route (#93045).
Resolve the config directly in the branch so the warning names the real
blocker, and let the dispatch's "none" fallback surface the
selection-specific error for an explicit openai choice. No fallback is
added: an unavailable selection still resolves to "none", it just
reports why.
resolve_runtime_provider() documents target_model as the explicit model
override for mid-session switches and auxiliary slots, but the custom
provider path (_resolve_named_custom_runtime) never received it and
silently substituted the provider's configured default_model instead.
This made auxiliary slots such as auxiliary.background_review silently
run the provider's default model rather than the configured one — e.g. an
ocx-proxy slot configured for gemini-flash actually executed
cursor/claude-sonnet-5, hitting upstream rate limits.
Pass target_model through to the custom runtime resolver and prefer it
over the provider's default model in both the pooled and non-pooled
credential paths.
Mark the write-only tmux load-buffer call as resolve-on-exit so a daemonized tmux server cannot retain inherited stdio and force a false timeout after the direct child has succeeded.
Completes the call-site acceptance item from #93134 as a companion to #93148.
Co-authored-by: JoaoMarcos44 <87440198+JoaoMarcos44@users.noreply.github.com>
The timeout handler only called settle(124) when resolveOnExit was true.
In the default path the promise waits for 'close', which requires every
inherited stdio handle to close — a daemonized grandchild that kept the
pipes open meant 'close' never fired, and after the timeout SIGTERM
(which only reaches the direct child) nothing settled the promise. The
await hung forever: the clipboard path (setClipboard -> tmuxLoadBuffer
-> osc.ts spawn without resolveOnExit) leaked a pending promise whenever
a spawned tool forked a stdio-inheriting daemon (#93134).
Settle(124) unconditionally in the timeout handler. The settled-guard
makes it a no-op when the child's own 'exit'/'close' won the race, so
normal timeout behavior is unchanged; in the daemon case it becomes the
only exit and returns the same 124 the close path would have.
Also un-skips the documented-hang regression test, with a 30s daemon
sleeper so it genuinely outlives the timeout (and vitest's own 5s test
timeout — before the fix the test fails by timing out, not asserting),
plus an elapsed bound.
Adversarial-review fixes for the #93057 snapshot-compaction PR:
- Fail-closed detachment: only re-enable compression after
bind_session_state successfully severs the engine's parent binding.
A failed rebind keeps the historical compression_enabled=False
behavior and warns, instead of running compaction against a
compressor still bound to the parent's SessionDB (#38727 re-open).
- Warm-cache parity: defer both compression gates (turn-prologue
preflight + pre-API pressure check) until the fork's first provider
response, so the first request replays the full snapshot as the
intended cached read and compaction applies from the second request
on — matching the documented budget mental model.
- Tests: regression for the rebind-failure fail-closed path (red on
pre-fix code) and the existing threshold-crossing test reworked to a
two-request review asserting the warm first request + compacted
second request. 116 tests green across all touched suites; ruff
clean.
Detach the review fork's compressor from the parent SessionDB/session_id
and re-enable in-memory-only compaction for oversized snapshots, instead
of the historical compression_enabled=False guard that left the fork's
replayed transcript unbounded (350k-384k input tokens per request, 1.49M
total across one 8-request review). Add an aggregate input-token budget
(auxiliary.background_review.max_input_tokens, default 600k) so repeated
tool calls cannot recreate an unbounded transcript; the tool loop stops
before the provider call that would cross it.
Closes#93057