Commit Graph

2418 Commits

Author SHA1 Message Date
alt-glitch 62b2d78025 fix(tool-search): bridge batch barrier, listing truncation, source indexing (salvage #92693, part 1)
Four fixes to the tool-search deferral layer, split from PR #92693 (the
availability-cache staleness fix ships separately):

1. The parallel batch planner now peels the tool_call bridge wrapper and
   decides admission on the underlying tool — supports_parallel_tool_calls
   works again when deferral is active. Unparseable wrappers stay
   sequential barriers; bridged calls get exactly the admission the same
   call gets direct. tool_search/tool_describe lookups batch concurrently.
2. _short_desc no longer truncates listing lines at 'e.g.', hostnames, or
   version strings — a sentence terminator must be followed by whitespace.
3. BM25 indexes the source label (e.g. 'linear' for mcp-linear), so
   service-name queries reach tools whose own name omits the service; the
   dead 'mcp' prefix token is stripped.
4. Substring-fallback docstring corrected (token misses, not zero-IDF).

Salvaged from #92693 by @alt-glitch with authorship preserved.
2026-08-25 15:24:15 -07:00
Gille 7c5c994397 fix(teams): request supported transcript content format 2026-08-26 02:54:35 +05:30
Teknium ba9fc55e16 feat(web): cache_exempt_hosts — always-live fetches for staging/tunnel sites
Sites under active development but tested over the public internet
(Vercel previews, ngrok tunnels, staging domains) are public DNS, so
the local-dev never-cache rule can't catch them. web.cache_exempt_hosts
lists hosts whose pages are always fetched live: exact, "*.wildcard",
or domain-suffix matching (label-boundary aware — mysite.dev covers
preview.mysite.dev but never evilmysite.dev). Checked on both store
and lookup, so adding an exemption takes effect immediately even for
entries cached before the config change.
2026-08-25 04:21:45 -07:00
Teknium f0381ee4ae fix(web): never cache local development URLs in the extract cache
Dev servers, hot-reload builds, and chat-GUI artifact previews live on
localhost/private addresses and change on every save — a 20-minute
cached copy would show a stale build exactly when freshness is the
point of fetching. The extract cache now declines loopback, private,
link-local, *.local, *.localhost, and single-label LAN hostnames on
both put and get. Hostname heuristics only (no DNS) — this is a
freshness carveout, not a security boundary; SSRF enforcement is
unchanged in tools/url_safety.py.

Public URLs keep the full TTL.
2026-08-25 04:21:45 -07:00
Teknium 8adef09be8 fix(web): extract cache serves only after policy + provider gates; rescue and format/provider isolation
Review fixes for #94618 (all three blockers reproduced by the reviewer
through the real web_extract_tool):

1. Cache lookup moved AFTER provider resolution and strict-selection
   validation, and gated per-URL on the website blocklist policy — a
   blocklist-blocked or misconfigured-backend call now behaves exactly
   as it would without a cache instead of serving cached content.
2. Rescue-served extract batches are never cached (mirrors the search
   memo's exclusion), keeping one-shot rescue one-shot.
3. Cache entries now get dedicated per-(url, format, provider) files
   instead of sharing the URL-keyed truncate-store file — html and
   markdown (or two backends') copies of one URL no longer overwrite
   each other, and switching extract backends within the TTL never
   serves the old backend's rendering.

Also from review: per-process index tmp filename (cross-process writers
can no longer truncate each other mid-write) and held flight locks are
never evicted from the bounded lock table (eviction could have allowed
a duplicate paid request).

New regression tests for formats/provider keying; E2E harness extended
with policy-block, strict-selection, rescue-two-call, and dual-format
scenarios — 6/6 pass; original 13/13 still pass.
2026-08-25 04:21:45 -07:00
Teknium 04603fc040 feat(web): TTL result caching for web_search + web_extract
Repeat searches (same normalized query + provider) within a 20-minute
TTL are served from an in-process memo, and concurrent identical
queries are single-flighted so a parallel subagent fan-out pays for
one vendor request instead of N. Requested limits bucket up to
10/20/50/100 so near-identical requests share an entry; callers get
their requested count sliced from the bucket.

Repeat extracts of the same URL are served from the existing
cache/web full-text store (previously written for read_file paging
but never read back), via a small JSON sidecar index. Disk-backed, so
CLI, gateway, cron, and subagents share it. Cached extracts re-run
the normal truncate pipeline, so per-call char_limit still works.

Both caches sit after every safety gate (secret-URL, SSRF, policy,
provider resolution) and directly around the paid vendor call — hits
skip only the network request. Only successful responses cache;
rescue-served responses are never cached (one-shot rescue must stay
one-shot). Config: web.cache_enabled (default on),
web.cache_ttl_minutes (default 20, clamped 1-1440).

Idea credit: query coalescing + num-bucketing pattern observed in
Apodex FrontierAgent (Apache-2.0).
2026-08-25 04:21:45 -07:00
Jan-Stefan Janetzky fb1ec36a4b fix(mcp): treat tools.include: [] as an explicit empty whitelist
_normalize_name_filter([]) returns an empty set, which is falsy, so
_should_register fell through to "no filter" and registered every tool
— the exact opposite of what _apply_tool_selection wrote when the user
unchecked everything in the install checklist ("contributes nothing
until reconfigured"). Whitelist mode is now keyed on the include key
holding a valid filter shape (str/list/tuple/set) rather than on set
truthiness, at both the live-discovery and cached-manifest sites.
Invalid include values keep the old warn-and-ignore behaviour.
2026-08-25 04:21:37 -07:00
Teknium d736f5d53f fix(docker): digest-suffix shared-container identity labels so distinct keys never collide
Review finding on #94633: _sanitize_label_value is lossy ('team/workspace'
and 'team_workspace' both sanitize to 'team_workspace'; >63-char keys
truncate identically), and container reuse is label-keyed — so two teams
with DIFFERENT shared keys could silently attach to one running container
(filesystem, processes, env) while their host sandboxes stayed separate.
Shared-key labels now carry a sha256 digest suffix of the raw key
(deterministic across processes; plain profile labels unchanged for
backward compat). Docs also state the first-creator-wins rule for image/
mounts on a shared container. Adds adversarial collision tests.
2026-08-25 04:00:27 -07:00
Teknium 82b32f32ef feat(terminal): wire shared-container key into profile-scoped resolver and MEDIA delivery
Follow-up on @fangliquanflq's opt-in (#84775): after the profile-scoping fix
(#94560) the container cache key is resolved in _resolve_container_task_id,
so the shared key must unify profiles there too — 'shared:<key>' for every
session of every opted-in profile AND for CLI/no-session runs. Delivery adds
the shared sandbox layout as the first translation candidate. Empty key
keeps strict per-profile isolation; SSH ignores the key entirely.
2026-08-25 04:00:27 -07:00
fangliquanflq 7a67bd07a7 feat(docker): support shared container identities 2026-08-25 04:00:27 -07:00
Teknium 76e306c458 refactor(tools): remove expired BFL FLUX 3 promo core tools (migration v39); FLUX 3 stays via video_gen/FAL for subscribers (#94599)
* refactor(tools): remove expired bfl_flux3_* promo tools; FLUX 3 rides the video_gen provider surface

* test: relay-cutover migration asserts >= v38, not the version literal
2026-08-25 02:45:10 -07:00
Teknium ce9b9a6351 feat(computer_use): guide models from full-screen grabs to interactive lanes
Full-screen captures carry no element tree, so the CaptureResult now has a
'note' field surfaced in the tool summary telling the model to call
capture(app='<AppName>') or capture(app='desktop') when it needs to act on
what it sees. Schema description updated to distinguish app='screen'
(composited full-screen image) from app='desktop' (shell surface with
clickable elements); docs + regression tests (14, sabotage-verified) added.
2026-08-25 02:31:34 -07:00
Teknium 15f7b7293c fix(terminal): persistent Docker containers are profile-scoped, not per-session
Commit a270c4ade's session-key fallback in _resolve_container_task_id was
added to stop cross-profile SSH environment reuse, but it wasn't backend-
gated: persistent Docker silently fragmented into one container per gateway
session, breaking the product contract (one long-lived container per profile,
shared by CLI and every session of that profile). #93950's vanishing MEDIA
attachments were downstream damage.

- persistent Docker (container_persistent: true) now keys to the profile:
  literal 'default' for the default profile (same container as CLI),
  'profile:<name>' for named profiles
- SSH and non-persistent Docker keep session scoping (the original leak fix
  and the #82731 isolation contract are untouched)
- gateway MEDIA translation follows the profile layout and keeps the legacy
  bug-window per-session sandboxes as fallback candidates, trying each until
  the file resolves — old sessions self-heal, no migration
- /root/.hermes credential-surface refusal preserved across all layouts
2026-08-25 02:30:38 -07:00
kshitijk4poor c8c3f4c448 fix(approval): machine-readable outcome parity on the gateway tails + sudo human-wait exclusion (#85125 2e) 2026-08-25 13:27:20 +05:30
Jony 335c60ecdd fix(skills): preserve review marks across contexts 2026-08-24 23:55:39 -07:00
Leegenux 6ce7ab8bfb feat(browser): make snapshot threshold configurable 2026-08-24 21:51:44 -07:00
pierrenode bf8b28f27a fix(tools): route browser snapshot storage through the symlink-safe writer
Today's spill/cache-writer hardening (tools/spill_safety.py,
write_text_exclusive/ensure_spill_dir with O_CREAT|O_EXCL|O_NOFOLLOW)
migrated tools/web_tools.py::_store_full_text() — which writes to the same
cache/web directory with the same content-hash filename scheme — but left
its near-identical sibling, tools/browser_tool.py::_store_full_snapshot(),
on the pre-fix plain open()/write_text() pattern. A pre-planted symlink at
the content-hash path redirected the write onto an arbitrary user-owned
file, same as the sites that commit fixed.

Reproduced live: with a symlink planted at the exact
browser-snapshot-<digest>.txt path (predictable from the snapshot content
hash), the pre-fix write followed the link and overwrote the link's
target with the (secret-redacted but otherwise user/page-controlled)
snapshot content.

Fix mirrors _store_full_text's exact usage: ensure_spill_dir(private=False)
+ write_text_exclusive(private=False, overwrite=True) — not private since
cache/web is bind-mounted into remote backends whose container UID must
read it; overwrite=True because re-snapshotting the same page state
legitimately reuses the same content-hash name (the overwrite path
lstat-unlinks the link itself, never following it to write through).

Added a regression test planting a symlink at the exact digest path and
asserting the link's target is untouched (only the link itself gets
safely replaced by a real file). Mutation-verified: with the fix stashed,
the pre-fix code wrote the snapshot content into the symlink's target
file, reproducing the vulnerability exactly.
2026-08-24 21:45:56 -07:00
Teknium a75ea37dc5 feat: browser snapshots drop LLM summarization — truncate-and-store like web_extract; auxiliary.web_extract slot removed
web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.

- tools/browser_tool.py: remove _extract_relevant_content and
  _get_extraction_model; oversized snapshots always truncate at line
  boundaries, store the full tree to cache/web, and append a read_file
  pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
  pattern as session_search/PR #27590), cli.py defaults + env bridge,
  gateway/run.py bridged keys, hermes config display, hermes model picker,
  dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
  zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
  mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
  to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
  LLM path is gone and stored files are secret-redacted
2026-08-24 20:11:18 -07:00
Teknium 48f69e51d3 fix(signal): chunk long standalone sends and cover both delivery paths (salvage #57929 + #67279)
Follow-up to lkz-de's adapter chunking commit: long Signal messages no
longer truncate on ANY delivery path.

- tools/send_message_tool.py: register Signal's 8000-char limit in
  _MAX_LENGTHS (imported from the adapter module so the two paths can't
  drift) so hermes send / cron standalone / MCP sends split via the
  shared truncate_message() pass instead of signal-cli rejecting them.
  Standalone-path idea credited to @5L-hermes01 (#67279).
- tests: regression test proving standalone Signal sends chunk at the
  adapter limit with no truncation footer (fails on pre-fix main).
- docs: Long Messages section on the Signal page (en + zh-Hans).

Both fixes verified by sabotage A/B (tests fail with the respective
half reverted to origin/main) and a real-import E2E: 27k-char message
with emoji + cross-boundary bold + code blocks -> 4 chunks, all styles
in-range UTF-16, lossless reassembly.
2026-08-24 20:03:24 -07:00
kshitij 41447a6d70 Merge pull request #94187 from kshitijk4poor/fix/85125-4b-terminal-treekill
fix(terminal): sweep setsid descendants after local timeout group-kill (#85125 4b)
2026-08-25 03:37:05 +05:30
kshitij 8d29a55bed Merge pull request #94188 from kshitijk4poor/fix/85125-4d-treekill-consolidation
refactor(deadline): consolidate site-local tree-kills onto agent.deadline.kill_process_tree (#85125 4d)
2026-08-25 01:48:09 +05:30
BlackishGreen33 c73d721b1d fix(computer-use): recreate CUA session suspect after MCP timeout (#74799)
An MCP call_tool deadline hit left the cua-driver session wedged for
all later computer-use calls. Mark the session suspect on a
concurrent.futures.TimeoutError and tear down + recreate it before the
next non-lifecycle call; healthy sessions are never restarted.

Fail-closed: the timed-out action may still have taken effect on the
remote screen, so it is never silently replayed — the error result
carries structuredContent.code=timeout_outcome_unknown with
next_step=fresh_state.

Informed by #74877 by BlackishGreen33.

Co-authored-by: BlackishGreen33 <s5460703@gmail.com>
2026-08-25 01:47:21 +05:30
kshitijk4poor 9990bcb8ce fix(terminal): sweep setsid descendants after local timeout group-kill (#85125 4b)
LocalEnvironment._kill_process kills the process GROUP (SIGTERM ->
1s wait -> SIGKILL -> 2s wait), but a descendant that called setsid
escapes the group and survives — the #71148 orphan class, terminal
flavor (issue #84967's local sibling).

Fix: snapshot the descendant set via psutil BEFORE the first SIGTERM
(children reparent to init once the wrapper dies, so a later parent
walk finds nothing — same snapshot-before-signal design as
agent/deadline.py kill_process_tree), then after the existing group
escalation completes, SIGKILL any snapshotted survivor whose pgid is
no longer the (now-dead) group. The TERM->KILL grace window for
in-group members is preserved (interrupts use this path too), the
Windows branch is untouched, and the snapshot is fully guarded — a
broken psutil never breaks the kill path (unit-tested).

Tests: live_system_guard_bypass acceptance test spawning a setsid
grandchild and forcing the timeout path (RED on unmodified file,
GREEN after), plus a psutil-failure unit test.

Docker design note (#84967 open question 1, condensed; full note at
/tmp/4b-docker-design-note.md): the docker backend inherits
base.py:1378 _kill_process, which only proc.kill()s the HOST-side
`docker exec` client — the in-container tree (child of containerd-
shim, not the client) survives every timeout entirely. Option A,
`docker exec <cid> kill -- -<pgid>` with TERM->KILL escalation using
a PGID captured at command start, is surgical and preserves container
state but needs a live container + shell and still misses in-container
setsid escapees. Option B, container restart, is absolute (PID-
namespace teardown kills everything) but destroys all in-container
state mid-session and punishes every other consumer of the shared
persistent container. Recommendation: Option A as a best-effort
_kill_process override (degrade to today's behavior on failure);
reserve restart for the existing container-gone recovery path.
2026-08-25 01:34:56 +05:30
kshitijk4poor 547f985286 refactor(deadline): consolidate site-local tree-kills onto agent.deadline.kill_process_tree (#85125 4d)
Per-site decisions:

1. hermes_cli/_subprocess_compat.py kill_process_tree(proc) -> None:
   MIGRATED. Body now delegates to agent.deadline.kill_process_tree(proc.pid)
   via a function-local import; keeps the swallow-everything fail-open
   contract and the (proc) -> None signature (agent/shell_hooks.py imports
   it by name; _kill_git_process_tree alias preserved). The old body is kept
   verbatim as _legacy_kill_process_tree and used as fallback when the
   delegation import/call fails. A final proc.kill() is retained on the
   happy path so Popen bookkeeping sees the exit (matches old behavior).

2. tools/browser_tool.py _kill_process_tree(proc): MIGRATED, same pattern
   (delegate + _legacy_kill_process_tree fallback). Behavior delta: the old
   body sent SIGTERM then SIGKILL with zero grace between them; the shared
   primitive sends SIGKILL only. With no grace period the observable effect
   is identical, and the psutil descendant sweep now also reaches
   agent-browser's setsid'd daemon grandchild, which killpg alone missed.
   tests/tools/test_browser_npx_warmup.py's TestKillProcessTree repointed at
   the legacy fallback (its assertions describe the fallback's internals).

3. tools/code_execution_tool.py _kill_process_group(proc, escalate):
   MIGRATED. It was a plain parent+descendants terminate (then wait 5s +
   kill when escalate=True) — expressed as two delegated calls:
   kill_process_tree(pid, sig=SIGTERM), then on escalate-timeout
   kill_process_tree(pid, sig=SIGKILL). Delegation failure degrades to
   proc.kill(), mirroring the old psutil-failure fallback. Delta: the old
   body terminated children before the parent; the shared primitive
   signals the group atomically (child is a session leader via
   start_new_session=True) plus an identity-aware descendant sweep —
   strictly wider coverage, same signals.

4. gateway/status.py: KEPT BOTH SITES.
   - terminate_pid (~l305) taskkill wrapper: NOT migrated. Its contract is
     incompatible with the shared primitive — it must RAISE OSError with
     taskkill's stderr on non-zero exit (callers branch on that), falls back
     to os.kill on FileNotFoundError, and its POSIX branch is deliberately a
     single-PID SIGTERM/SIGKILL, not a tree kill. Wrapping the bool-returning
     fail-soft primitive would invert the error contract.
   - reap_gateway_children (~l2029): NOT migrated. It operates on a
     pre-snapshotted child list from a parent that is already dead
     (psutil.Process(pid) on the parent would fail), and every signal is
     wrapped in identity/ownership checks the primitive lacks: is_running()
     identity, zombie skip, and the skip-if-ppid-still-equals-parent guard,
     plus SIGTERM -> wait_procs -> SIGKILL staging and a reaped-count return.
     The coupling is the feature; migrating would delete the safety logic.

5. scripts/run_tests_parallel.py _kill_process_tree (~l253): NOT migrated.
   Dev tooling that intentionally kills by CAPTURED pgid because the direct
   child is usually already reaped (psutil/pid-based primitive cannot find
   it), and it avoids the psutil import on the test-runner hot path. Its
   docstring already documents why psutil is the wrong tool there.

New tests: tests/agent/test_treekill_consolidation.py — delegation +
raise-swallowing tests per migrated wrapper, consumer-identity checks, and
a live end-to-end probe (setsid grandchild dies through the compat wrapper,
zero survivors).
2026-08-25 01:34:56 +05:30
kshitijk4poor 7dde1b8b0b fix(mcp): resolve tool-call timeouts via the unified deadline layer (#85125 2g)
Both readers of the per-server MCP tool timeout (the connection's run()
and the cache-path registration) read config.get("timeout", 300) as
their own private resolution. Route them through _resolve_tool_timeout:
per-server mcp_servers.<name>.timeout still ALWAYS wins (most specific),
then timeouts.mcp.tool_call from the unified timeouts: section, then
the unchanged 300s default. Values pass through resolve_timeout's
platform clamp; resolution failure falls back to the historical default.

Default-behavior invariance pinned by contract tests (nothing
configured -> exactly 300, per-server beats section, section beats
default, invalid/failed resolution falls back).
2026-08-24 17:12:15 +05:30
Teknium d8d1e18ab9 test(terminal): harden watch_patterns lifetime cap — delivered-only counting, Nth-delivery promotion, docstring
Follow-ups on top of the cherry-picked #93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
  lifetime budget; the cap trips exactly at the Nth DELIVERED match and
  promotes to notify_on_complete with the watch_disabled summary queued
  right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
  the global breaker drops the final match, so the user always learns why
  watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
  was already updated by #93532).

Refs #93513
2026-08-24 03:22:48 -07:00
chelsealong b3730153c3 fix(terminal): cap watch_patterns notifications over a process's lifetime
Per-session rate limiting only counts consecutive strike windows, so a
pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS
(e.g. a service restarted repeatedly over a day) never trips the
existing strike-limit disable — each match lands in its own clean
cooldown window. Every one of those matches still forces a full-context
agent turn, which stalls the event loop on large sessions (#93513).

Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many
watch_match notifications over its whole life, disable watch_patterns
and fall back to notify_on_complete, reusing the existing disable path.
2026-08-24 03:22:48 -07:00
Teknium 42a6d761d2 fix(bot-relay): add shutil.which step to CLI resolution and pin utf-8 decoding on delivery subprocess
Salvage hardening on top of #93601 (with #93597 covering the same core
mechanisms) for #93590:

- _hermes_cli(): after the venv-sibling check (hermes.exe on win32),
  try shutil.which('hermes') before the bare-name fallback, so
  environments with a PATH but no venv sibling resolve exactly what an
  interactive shell would. Platform test switched os.name -> sys.platform
  ('win32') per repo convention.
- tui_gateway/methods_bot_relay.py deliver: pin encoding='utf-8',
  errors='replace' on both subprocess.run sites — without them the
  child's UTF-8 output is decoded with the locale codec (cp1252/GBK on
  Windows), mangling non-ASCII replies or raising on undecodable bytes.
- Regression tests: shutil.which resolution step, bare-name fallback
  with which=None, and encoding-pin assertions in the deliver transport
  test.

Refs #93590, #93597, #93601
2026-08-24 03:21:37 -07:00
liuhao1024 85cd576b06 test(bot-relay): match delivery CLI by basename in argv filters
CI runners have a real hermes sibling next to the venv python, so
local_delivery_command now resolves an absolute path there — the exact
argv filters in the retry-policy fakes and the relay-methods pins must
match by basename instead of the literal "hermes", mirroring the
_delivery_lock matcher.
2026-08-24 03:21:37 -07:00
liuhao1024 c099ef05de fix(bot-relay): Windows path SyntaxError in waiter + PATH-less delivery ENOENT
Two failures on a Windows desktop install relaying to a remote gateway
(#93590):

1. waiter_command embeds the reply path in generated python -c source
   with !r. repr escapes each backslash, but the Windows execution layer
   folds \\ back to \, so \U in C:\Users\... parses as a unicode escape
   and SyntaxErrors the whole waiter script. Raw-string literals keep
   the folded single backslash a literal; POSIX paths have no
   backslashes so the prefix is a no-op there, and \' inside a raw
   literal still cannot terminate the string, keeping the #93091
   injection defense intact.

2. local_delivery_command hardcoded "hermes", relying on PATH — absent
   in service contexts (systemd units, desktop launchers, non-login SSH
   shells), so delivery died with ENOENT. It now resolves the CLI next
   to this gateway's own interpreter (venv bin/Scripts sibling,
   hermes.exe on Windows) with a bare-name fallback. The #93091
   per-profile turn-lock recognition in bot_mode_dm now matches the CLI
   element by basename (split on both separators) so resolved absolute
   paths still take the lock instead of silently bypassing it.

Fixes #93590
2026-08-24 03:21:37 -07:00
Teknium a7aa814c42 fix(tools): widen the command-position anchor to the whole hardline class
#93392 was not just one pattern: every hardline rule with a bare \b anchor
fired on its token anywhere in the command line, including inside quoted
prose handed to echo, git commit -m, or gh --body. Anchor the
command-name-token rules and quote-mask the positionless ones:

- dd-to-block-device and kill -1 get the same _CMDPOS anchor as the
  format/rm/shutdown families, keeping their argument tails.
- redirect-to-block-device and the fork bomb have no command-name token to
  anchor (`>` appears mid-command; the bomb is a function definition), so
  they now match a quote-masked variant (_mask_quoted_prose) where quoted
  string content is blanked. $() and backtick spans inside double quotes
  stay raw (the shell executes them), and any command whose command-position
  words include a shell carrier (sh/bash/zsh/ksh/dash -c, eval, source, .)
  is scanned unmasked -- quoting is not a bypass. bash/sh -c payloads also
  still surface as raw detection variants via _execution_flag_findings.

Regression tests cover both directions for every touched pattern: quoted
prose passes, and every true-positive shape (bare, ; && | separators,
sudo/env prefix, $(), backticks, sh -c/bash -c/eval payloads) stays on the
unconditional floor.
2026-08-24 03:20:14 -07:00
liuhao1024 8163c8731b fix(tools): anchor the mkfs hardline pattern to command position
mkfs was the only HARDLINE_PATTERNS entry without a _CMDPOS anchor, so
the unconditional floor blocked any command that merely mentioned the
token inside quoted prose — `echo "does this workflow use mkfs
anywhere?"` was refused outright (#93392) instead of running the echo.

Anchor mkfs to command position like every sibling entry (rm root-
delete, shutdown family, dd): it matches at the start of a command,
after separators, or behind sudo/env/exec/nohup/setsid wrappers, and
no longer fires on argument-position mentions. The quote-aware
_mark_command_starts pass already keeps separators inside quoted
strings from looking like command starts, and \b still protects
mkfs_helper-style names.
2026-08-24 03:20:14 -07:00
Teknium ec44116d59 test(tools): shared sanitizer contract + singularity overlay coverage
Behavior-contract tests for sanitize_task_id_for_path (colon/separator
removal, verbatim pass-through for existing safe ids, determinism,
collision-freedom incl. the a:b vs a_b digest case, traversal and
oversized-id bounds) and for the singularity persistent overlay path
(sanitized, verbatim for safe ids, distinct dirs for colon-vs-underscore
ids).

Co-authored-by: chelsealong <chelsealong@126.com>
Co-authored-by: Parker Fawcett <259203091+Parker-Fawcett@users.noreply.github.com>
2026-08-23 21:12:32 -07:00
HexLab98 eef7a10756 test(docker): cover session-key sandbox paths and their collision boundary
Drives the real DockerEnvironment constructor with a Telegram DM session key
and asserts every persistent -v spec is a two-field bind whose source holds no
colon — the assertion that reproduces exit 125 on the unfixed path.

The derivation's own contract is covered separately: ids that already work stay
verbatim (no sandbox migration), docker's separator and the path separators
never survive, ids differing only in rewritten characters keep distinct
directories, the mapping is stable across calls so cross-process container
reuse still resolves, pathological keys stay inside the per-component length
limit, and "."/".."/empty cannot resolve to the docker sandbox root.
2026-08-23 21:12:32 -07:00
Teknium c584d15cdc feat(bots): typed failure reasons reach the sending agent on A2A calls (#93091)
message_agent callers previously got provider prose (a raw 401
paragraph, a missing-provider essay) and could not branch on the
failure class. Now the #93091 item-1 reason enum rides the whole relay
roundtrip:

- Desktop relay drain forwards bot_relay.deliver's error.data.reason
  into bot_relay.reply (and prefers it for the attention badge over
  free-text re-parsing);
- write_reply already persisted reason / classified fallbacks;
- the sender-side waiter prints "[reason: <code>]" ahead of the free
  text, so the completion notification the sending agent receives is
  machine-branchable.

Additive everywhere: healthy replies unchanged, reasonless errors
classify to a code, old consumers keep working.
2026-08-23 20:07:21 -07:00
Adolanium 2912c36aa4 fix(gateway): stop multiplex allowlist leak and bot-relay python -c injection
_auth_env fell through to os.environ on a scoped miss, so one profile
could inherit another profile's allowlists and allow-all flags.

bot_relay.waiter_command put connection_id into python -c source. A
quote in the id broke the waiter. A crafted id could run extra Python
in the sender gateway.
2026-08-23 20:00:30 -07:00
Teknium 57649294be test(bots): turn-lock fake Proc gains stdout/stderr attrs
_run_delivery now captures output to drive the retry policy; the
turn-lock test's minimal _P fake predates that contract. Sibling-test
blast radius fix, no behavior change.
2026-08-23 20:00:18 -07:00
Teknium b274b346d8 feat(bots): retry session policy — resume transient turns, compress-and-resume on context overflow (#93091 item 5)
Maintainer ruling (2026-08-23): a retried bot turn never mints a fresh
session. retry_action() maps the #93091 item-1 reason enum to one of
resume / compress_then_resume / none:

- transient classes (runtime_offline, delivery_timeout, rate limit,
  server error) re-run the same Bot Chat session once;
- context_overflow also re-runs the same session — the retried turn
goes through the pre-API compaction pass in conversation_loop.py,
  which compacts the over-threshold transcript first (the one
  sanctioned context mutation); no fresh-session escape hatch exists;
- auth/quota/config/model classes never auto-retry.

Wired at both delivery surfaces (fix the class, not one site):
bot_relay.deliver (relay handler) and _run_delivery (local
message_agent runner). Failed deliveries now carry the classified
reason in the structured error payload (error.data.reason).

Sabotage-verified: with the retry blocks removed, 3 consumer tests
fail; with them present, 22/22 pass.
2026-08-23 20:00:18 -07:00
Teknium 081cdd9911 fix(terminal): subagents no longer hijack the tty with an interactive sudo prompt
delegate_task children run on worker threads of the parent process and
inherit the process-wide HERMES_INTERACTIVE=1 the CLI sets at startup.
_transform_sudo_command's interactive gate therefore fired inside
children with no sudo callback registered, falling through to the raw
/dev/tty password prompt: a password box printed mid-TUI from a
background thread, parallel children racing for the tty, and each child
blocked for the full 45s timeout.

Gate the prompt (and the sibling 'you will be prompted again' message
after an auth failure) on agent.delegation_context.is_delegated_child_context(),
the ContextVar set around every child run and propagated through
contextvars.copy_context onto the executor thread. Children now behave
as headless for sudo: configured SUDO_PASSWORD, the session cache, and
the NOPASSWD probe still work; otherwise the command fails gracefully
with a subagent-specific tip.

A/B verified: 3 regression tests fail on merge-base, 7/7 pass at head.
2026-08-23 19:00:47 -07:00
Teknium ec06e706f1 test(cron): e2e regression — scrubbed child env resolves bare hermes under minimal parent PATH
Exercises the real build_subprocess_env()/_resolve_hermes_bin_dir chain (no
helper mocks) under a simulated systemd/cron minimal PATH, the exact call
path cron/scheduler._run_job_script uses. Companion to #93082.
2026-08-23 18:27:15 -07:00
UniversePeak b0001f45a2 fix(cron): keep hermes console script on child PATH 2026-08-23 18:27:15 -07:00
justcarlosm 478a09c06b fix(browser): floor browser-use CLI subprocess PATH with sane system dirs
Profile-spawned workers (kanban bots, cron jobs) can inherit a PATH of
only version-manager dirs — observed in the wild as one nvm node dir
repeated 7x. The uv-installed browser-use binary is a POSIX sh
trampoline that resolves dirname/realpath through PATH, so it died
with 'realpath: not found … exec: /python: not found' (exit 127)
before its own Python ever started.

_base_subprocess_env now floors the child PATH via browser_tool's
_merge_browser_path (the agent-browser backend already guards the same
hazard), degrading to appending FHS bin dirs if that import is ever
unavailable. Windows is a no-op (.cmd shims don't trampoline).

Verified: unit tests + real uvx browser-use --version under a
nvm-only-PATH worker env, rc 127 -> rc 0.
2026-08-23 18:25:35 -07:00
liuhao1024 5d8b031514 fix(stt): surface the selection-specific error for explicit openai STT
When the managed openai-audio gateway is unavailable,
_resolve_openai_audio_client_config() raises a ValueError that names the
blocker (and, for managed-Nous users, the `hermes tools` remediation).
The boolean probe in _get_provider's explicit-openai branch flattened
that into False, so the log claimed "no API key available" and the
transcription result returned the all-provider install hint -- pointing
operators at unrelated setup instead of their managed route (#93045).

Resolve the config directly in the branch so the warning names the real
blocker, and let the dispatch's "none" fallback surface the
selection-specific error for an explicit openai choice. No fallback is
added: an unavailable selection still resolves to "none", it just
reports why.
2026-08-23 18:25:35 -07:00
fangliquanflq c1c0efa375 fix(code-exec): preserve interrupt cancellation source 2026-08-23 18:25:19 -07:00
BotUser 679e07a074 fix(gateway): close order-dependency + missing-verb gap in launchctl lifecycle guards
The gateway-lifecycle guards in cron/lifecycle_guard.py (Branch B, the
unconditional hard-block used by cron creation and the terminal tool when
_HERMES_GATEWAY=1) and tools/approval.py's launchctl rule both matched
`launchctl <verb> ... hermes[.-]?gateway` as a single sequential regex,
requiring the hermes-gateway label to appear literally AFTER the verb.

A shell command that builds the label earlier in the string — e.g. a
for-loop reading labels from a list defined before the actual launchctl
call — defeats that ordering entirely:

    for item in 'ai.hermes.gateway-apollo:...' 'ai.hermes.gateway:...'; do
      label=${item%%:*}; plist=${item#*:}
      launchctl bootout "gui/$uid/$label"
      launchctl bootstrap "gui/$uid" "$plist"
    done

The literal text "hermes.gateway" only ever appears in the for-list,
never after "bootout" — so `[^\n]*\bhermes[.\-]?gateway` never matches at
the verb's position, even though the command unambiguously targets the
gateway's own launchd label.

cron/lifecycle_guard.py's verb list also didn't include `bootout` at all
(present in tools/approval.py's list and covered by its own test suite —
`launchctl bootout ai.hermes.gateway` is explicitly asserted as dangerous
there — so the omission in the sibling file looks like list drift between
the two guards rather than an intentional exclusion).

`bootout` is the verb that actually deregisters a launchd job (unlike
kickstart/stop, which just bounce a still-registered one), so a command
using it evades both guards, then removes the service from launchd with
no supervisor left to bring it back — worse than a simple restart-loop.

We hit this for real: a gateway self-restart (triggered from a chat
request to change the default model) used a raw terminal `launchctl
bootout`/`bootstrap` loop across 4 launchd labels instead of the normal
`hermes gateway restart` path. It slipped past both guards, self-bootout
killed the process mid-drain before its own follow-up bootstrap could
run, and all 4 gateway profiles ended up fully deregistered from launchd
with zero user approval (approvals.mode: manual was configured) until
someone manually re-bootstrapped them.

Fix: both guards now check "a launchctl lifecycle verb appears somewhere
AND a hermes-gateway label appears somewhere", independent of order, and
cron/lifecycle_guard.py's verb list gains bootout/kill/disable/remove to
match tools/approval.py's existing set. Internal recovery code
(hermes_cli/gateway.py's own `subprocess.run(["launchctl", "bootout",
...])` calls) is unaffected — these guards only scan shell-command
strings composed by the agent's terminal/cron tools, not the CLI's
trusted internal subprocess argument lists.

Adds regression tests in both test files reproducing the exact incident
command (label built in an earlier for-loop segment, referenced only via
`$label` at the point of the verb).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 17:56:36 -07:00
thenabbu acf8245607 fix(tools): pass single_query_deny_message to the ssh-config write approval gate
Commit 1596148ff made single_query_deny_message a required keyword-only
parameter of _run_approval_gate() and updated its two callers inside
tools/approval.py, but missed the third caller: the SSH-config write
guard in tools/file_tools.py (_check_approval_required_write,
pattern_key="ssh_config_write").

Any gated write to an SSH client config therefore raised
  TypeError: _run_approval_gate() missing 1 required keyword-only
  argument: single_query_deny_message
instead of routing through the human-approval flow.

- Pass the kwarg with a single-query-specific deny message that points
  operators at approvals.single_query_mode: approve.
- Add a regression test asserting the gate call passes every required
  kwarg (fails on unpatched main).

Fixes #93201
2026-08-23 17:54:20 -07:00
Axel Vanni 6d501c2958 fix(cron): make gateway lifecycle matching shell-token aware (#80269)
The hard block matched raw command text, but a shell resolves quote
splicing (`kick"start"`) and backslash escaping (`kick\start`) into the
literal verb before execution. So `launchctl kick"start" -k
gui/501/ai.hermes.gateway` ran exactly as the blocked `kickstart` form
while both the non-bypassable block and the approval detector missed it —
leaving an approval-bypassing gateway self-lifecycle operation reachable.

contains_gateway_lifecycle_command now runs a second pass over
shlex-tokenized command segments, where quotes and escapes are already
resolved. It stays anchored on a hermes-gateway identifier, so prose and
non-gateway hermes services are unaffected. Because this function is the
single choke point _contains_unsafe_gateway_action calls at every
recursion level, referenced-script and `sh -c` payload scanning inherit
the fix.

tools/approval.py had the same gap for quote splices: backslash escapes
are stripped by _normalize_command_for_detection, but quote splicing in an
ARGUMENT position is not touched by _deobfuscate_shell_word_for_detection
(scoped to command-position words, deliberately — widening it would let
quoted prose match the destructive patterns). It now delegates to the
fixed guard as a last check, so an ordinary pattern match still wins and
keeps its more specific reason string.

Tests: quoted, single-quoted and backslash-spliced verbs across the
launchctl/systemctl/hermes branches, the spliced gateway identifier
itself, a splice nested in an `sh -c` payload (resolves one level deeper,
asserted at the recursive entry point terminal_tool actually calls), plus
negative cases proving prose and non-gateway labels stay unblocked.

Verified on Windows: no regressions — the 10 remaining failures across
tests/tools/test_approval.py, tests/hermes_cli/test_gateway_restart_loop.py
and tests/cron are identical on the unmodified baseline (POSIX file modes,
symlink privileges, and /bin/bash script paths).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 16:28:56 -07:00
Axel Vanni 320d884d88 fix(cron): cover bootout/remove/disable in the gateway lifecycle guard
Branch B of _GATEWAY_LIFECYCLE_PATTERN enumerated launchd verbs but omitted
`bootout` - the modern replacement for the `unload` it already listed, and
the paired inverse of the `bootstrap` it already listed. `remove` (legacy
sibling of bootout) and `disable` (what makes an unload durable) were
missing for the same reason.

This matters because the two enforcement layers are not interchangeable. In
tools/terminal_tool.py under _HERMES_GATEWAY == "1":

  - the cron.lifecycle_guard hard block is documented as applying
    unconditionally ("force=True cannot help here")
  - detect_dangerous_command below it is explicitly skipped when force=True

detect_dangerous_command already flags all three verbs, so the default path
was covered - but with force=True inside the gateway they reached execution
while stop/unload/kickstart did not. SIGTERM then propagates to the child
before the command completes and the service may never come back, which is
the state described in #74973.

The label anchor (\bhermes[.\-]?gateway) is unchanged, so unrelated services
such as `launchctl bootout gui/501/ai.hermes.update-checker` stay runnable.

Adds TestLifecycleGuardLaunchctlParity, which pins the one-directional
invariant: anything the bypassable approval layer flags, the unbypassable
hard block must also catch. Deliberately not equality - the hard block is
legitimately stricter (it also covers load/restart, which the approval layer
leaves alone). Verified failing on the parent commit for exactly bootout,
remove and disable.

Closes #80260

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 16:28:56 -07:00
Brooklyn Nicholson 165d1849e2 fix(approval): stop the CLI and ACP offering a scope the protected gate discards
The protected agent-instruction gate grants one operation and persists
nothing, but only the TUI/desktop and Runs transports were taught that.
The prompt_toolkit panel, the input() fallback, and the ACP editor menu
still rendered "Allow for session", so a user editing SOUL.md tapped it,
got re-prompted on the next write, and read the gate as broken.

Thread allow_session through prompt_dangerous_approval so a caller that
re-asks every time collapses every surface to once/deny, and cover the
producer-to-transport contract end to end.
2026-08-23 17:45:47 -05:00
kshitij c460e87d10 fix(bot-mode): review follow-ups for the turn lock
- Drop the false fairness claim from acquire_turn_lock's docstring (LOCK_NB
  probe + sleep retry gives no arrival-order guarantee; only the budget is).
- logger.debug once when the lock degrades to a no-op on fcntl-less
  platforms so silent serialization loss stays diagnosable.
- Document the real worst-case deliver handler hold (120s lock wait + 600s
  turn = ~720s) where clients tune their timeouts against it.
- Pin non-reentry: local_delivery_command must stay a raw 'hermes -p' argv —
  wrapping it in --run-delivery would make the child contend with its
  parent's own flock and fail every relay delivery with target_busy.
- De-flake: the cross-profile test's upper-bound wall-time assert tolerates
  loaded CI runners; the wait-duration message assert matches ~Ns generally.
2026-08-24 02:39:06 +05:30