On a cloud VM the instance-metadata service hands live IAM/service-account
credentials to any local process with no auth, so a fetch against it is
credential exfiltration unless the operator expects it — yet
detect_dangerous_command() auto-approved `curl` against the link-local
metadata IP, metadata.google.internal, and the Alibaba endpoint. Add one
DANGEROUS_PATTERNS entry covering 169.254.169.254 (AWS/Azure/GCP/OpenStack),
its AWS IPv6 form fd00:ec2::254, metadata.google.internal, and Alibaba's
100.100.100.200. The host literals have no other use, so their appearance in
a command is the signal regardless of HTTP client; lookarounds keep other
169.254.x.x link-local addresses and longer host/dotted strings out.
This prompts for approval (legit uses exist on real cloud VMs); it is NOT a
hardline block. Deterministic containment-escape detection at the approval
layer, same class as the existing credential-path detectors.
The package-manager uninstall patterns were bare \b-anchored, so quoted
prose (`git commit -m "document npm uninstall usage"`) prompted while the
real `npm --prefix DIR uninstall x` slipped past because the option group
did not allow an operand. Use the file's _CMDPOS anchor like every other
command-name rule and let each global option take one operand.
Found by independent review before merge.
`npm uninstall -g`, pnpm/yarn remove, `pip uninstall` and `brew uninstall`
remove software outside the project yet matched no dangerous-command
pattern, so the agent ran them without asking (#10199). Add one
"package manager uninstall" rule per manager; installs and updates stay
unprompted.
Hand-ported from PR #64175 (the patterns moved from tools/approval.py to
tools/approval_detection.py after it was opened).
Adapt the command-position, bounded-candidate and launcher-option work from
embwl0x's #76063 to the current detection owner, then add executable basename
projection from Rohith Pariki's #104338. Parse raw quote state before applying
existing text normalization so quoted arguments do not become commands.
Cover shell payloads and literal env split-string carriers, retain path-specific
rules and whole-command globs, and document the supported normalization rather
than claiming an OS capability sandbox. Related: #104308, #76037, #76063,
#104338, #78521, #86711. No automatic closing directives: the older carriers
also contain broader case syntax and git-option work not included here.
Co-authored-by: embwl0x <embwl0x@users.noreply.github.com>
Co-authored-by: Rohith Pariki <rohithpariki@gmail.com>
Independent review of the first fix found two defects. Bounding the grep lexer
removed the old "malformed payload" floor, and that floor had been the ONLY thing
blocking a hardline command hidden after a newline inside a double-quoted
$(grep ...) substitution: _mask_quoted_newlines treated the newline as quoted
DATA and replaced it with a space, so the later command-start pass saw
`/dev/null reboot` as an operand and the public guards approved a hardline
command with no callback. A substitution body inside double quotes is
executable; the masker now recurses into $(...) and backtick bodies with a fresh
quote state so a newline there stays a command boundary. The witness now blocks
as "system shutdown/reboot" (the correct verdict), as do the ; / && / | / nested
/ backtick variants.
Second: the lexer stopped at ANY backtick as if it closed an enclosing
substitution, so a grep with a backtick OPERAND lost its operands and was
reported malformed (a new false positive). A backtick opened after the grep is
an operand substitution; only an unmatched one closes the enclosing command.
Tests: the six hardline-inside-substitution shapes block as themselves; the
public guard blocks with zero approval callbacks; the backtick operand, a data
pattern and a quoted-newline commit message stay allowed. Approval/hardline
suites: 637 passed; the one failure (real-binary sort payload probe) fails
identically on main.
The quoted-grep scanner tokenizes from the grep's position to decide which
quoted operand is inert data. _shell_tokens_with_spans lexed from there to
the END of the top-level segment, so for a grep nested in a substitution
such as
sed -n "$(grep -n X f | cut -d: -f1),+3p" f
it read past the closing ")" into the enclosing command, met the outer
closing double quote as an opener, ended in quote state, returned None
("malformed"), and detect_hardline_command reported the whole command as
being on the unconditional blocklist. Nothing about the command was
dangerous. In the 1,393-agent refactor run this fired 546 times across the
workers (every one a false positive; 318 of 323 retries succeeded by
rewording), and three times on the maintainer's own session in one day.
The lexer now stops at the end of the simple command it was asked to lex:
an unquoted list separator or pipe or newline, or the ")" / backtick that
closes the substitution it sits inside (tracking $(...) depth opened after
start so a nested substitution's closer is not mistaken for the outer one).
Genuinely unbalanced quoting still returns None and still fails closed;
every hardline pattern is unchanged.
Tests: six substitution shapes lex clean and are not hardline; the lexer
yields exactly the grep's own words; unterminated quoting still blocks;
rm -rf / , $(rm -rf /) and shutdown still block.
tools/approval.py no longer re-exports sibling names (approval_context/prompt/floors/detection/
human_wait/smart/gateway_wait); it imports only what it uses. Siblings reference sibling-defined
names directly (module-attribute reads on tools.approval_context so patching the defining module
still works); only facade-owned state (_lock, _gateway_queues, _permanent_approved, _denied,
_denial_breaker_addendum, _gateway_notify_cb) is still read back through tools.approval.
approval_detection calls its own _command_detection_variants instead of late-binding through the facade.
Follow-up to b8f99bfc43: the seam edits for hermes_state_repair.py and tools/approval_detection.py
were overwritten by a concurrent squad's write before that commit landed (only their docstring
restores got in). Re-apply: _repair_conn/_open_exclusive/_db_opens_cleanly look up
_connect_repair_durable via hermes_state at call time; detect_dangerous_command/
detect_hardline_command look up _command_detection_variants via tools.approval, so patching the
facade (as tests/state/test_state_db_wal_unlink_race.py and tests/hermes_cli/test_approvals_test.py
do) reaches the call again, as on BASE 63279301bc.
Reviewer P2 (kshitijk4poor): tests patch hermes_cli.models.get_cached_nous_inference_base_url
but models_pricing.pricing_cache_scope read its own module global, so the patch never reached
the call and the test passed on the default-endpoint fallback. Same seam-erosion class audited
across /tmp/rf/patch_traps.json (777 candidates) with an AST reachability check + a per-test
call-count probe (facade vs defining module) against BASE 63279301bcb; three seams actually
bypassed their patch on HEAD but not on BASE:
- hermes_cli.models.get_cached_nous_inference_base_url <- models_pricing.pricing_cache_scope
- hermes_state._connect_repair_durable <- hermes_state_repair._open_exclusive/_repair_conn/_db_opens_cleanly
- tools.approval._command_detection_variants <- approval_detection.detect_{dangerous,hardline}_command
Each now looks the name up through its facade at call time (the pattern hermes_state_repair
already used in live_writer_holds_db), restoring BASE's patchability.