* refactor(video_generate): capability-gated dynamic schema — 6 optional args render only when the active provider/model honors them; fleet capability declarations + declaration<->implementation contract tests
* fix(video_gen): H3/Grok/Happy-Horse/Gemini audio is ALWAYS-ON native, not absent — new audio_native family key + audio_always_on capability surfaces as description line (maintainer catch)
* test(video_gen): duration-span test pins the active-model contract — resolved family's real window, short families not inflated, union fallback still spans 30s
Pillow is an optional dependency in this codebase (every other PIL use in
tools/vision_tools.py imports lazily and falls back). Both salvaged
validators now distinguish 'PIL missing' (pass through, header-only
sniff) from 'decode failed' (reject), so a Pillow-less install keeps
working instead of rejecting every PNG.
Follow-up to salvaged #53307 (@CannibalKush) and #76896 (@HaiyiMei).
Second review round: honoring base64Encoded recursively let any nested
dict spoof {"base64Encoded": true, "data": "<secret>"} past the redactor
(Runtime.evaluate returns arbitrary by-value JSON), and two carriers were
missed entirely — Network.streamResourceContent returns unflagged binary
bufferedData, and Network.getRequestPostData's postData was not covered.
Replace the ambient field sets with per-method exact result-path specs:
_CDP_ALWAYS_BINARY_PATHS for declared-binary paths (screenshots, PDFs,
streamResourceContent, beginFrame screenshotData, the nested
CacheStorage.requestCachedResponse.response.body) and
_CDP_FLAGGED_BINARY_PATHS for paths whose carrier object's base64Encoded
sibling gates the exemption (Network/Fetch.getResponseBody body, IO.read
data, getRequestPostData postData). Path suffixes propagate only into the
matching subtree, so base64Encoded is type information solely on trusted
carrier objects — never ambient trust in nested JSON.
Architecture-review follow-up: the method-scoped binary_payload flag skipped
redaction for every string anywhere in the result of the two listed methods,
and the same corruption stayed reachable through Network.getResponseBody /
Fetch.getResponseBody / IO.read / Network.streamResourceContent. Make the
exemption field-scoped instead: an explicit schema exempts exactly
Page.captureScreenshot.result.data and Page.printToPDF.result.data (carriers
with no flag of their own), and any dict whose base64Encoded sibling is
exactly True exempts its body/data/bufferedData string (the protocol's own
discriminator — text bodies with base64Encoded: false stay redacted). Every
other string in every result keeps full secret redaction.
_redact_cdp_output applied redact_sensitive_text(force=True) to every string
in CDP results, including the base64 screenshot/PDF payload of
Page.captureScreenshot and Page.printToPDF. The Fernet pattern (gAAAA + base64
alphabet) matches arbitrary spans inside such payloads wherever gAAAA follows a +
or /, collapsing them to first6...last4: decoded PNGs came out corrupt (valid header,
CRC failures mid-IDAT, no IEND), the persisted full copies in tool_result_storage
were redacted too, and vision_analyze then embedded corrupt images that the provider
rejected with 400 invalid_image - killing resume sessions with a misleading provider
error. Skip redaction for the two binary-payload methods: the payload is binary,
not free text, so there is no secret to protect there. Every other method keeps
full redaction.
Follow-up to @AIalliAI's #60750 commits: with python_environ_get_secret
downgraded to medium, the high-severity python_os_environ pattern still
fired on the same os.environ.get("...KEY") line, re-escalating the verdict
the downgrade intended to avoid. Exempt every .get() form (non-secret =
config read; secret-shaped = scored medium by the dedicated pattern), and
apply the same medium grade to the sibling os.getenv() secret pattern —
same shape, same rationale (#60709 point 2).
The original ^(?!\s*#) prefix only skipped full-line comments starting
with '#'. An inline comment like:
cfg = environ.get('HOME') # os.environ available
still triggered python_os_environ because the regex matched the code part
before the '#'.
Two complementary fixes:
1. Replace ^(?!\s*#) with ^[^#\n]* in the regex — this rejects any line
where a '#' comment marker appears anywhere before os.environ.
2. Add _compute_docstring_lines() — a state machine that pre-computes
lines inside triple-quoted strings (docstrings) and skips them during
pattern matching. Also handles single-line self-contained docstrings.
6 new regression tests covering: inline comments, multi-line docstrings,
single-line docstrings, full-line comments, and a verification that real
bare dict(os.environ) code still triggers. All 85 tests pass.
Five targeted fixes for #60709 (reported by @mvanhorn):
1. ruby_env_secret: scope ENV[] to case-sensitive Ruby constant
((?-i:ENV)) — no longer matches Python env[key] dict access.
2. python_environ_get_secret: downgrade critical→medium — reading
an API key via os.environ.get() is normal auth, not exfiltration.
3. python_os_environ: skip comment lines with ^(?!\s*#) — no longer
flags os.environ references in docstrings or code comments.
4. deception_hide: downgrade critical→high + negative lookahead for
UX guidance context (unless/except/until/confirm/diagnose/verify).
5. oversized_skill: downgrade high→low + raise cap 1MB→5MB — large
skills are legitimate; structural size is informational only.
All 80 existing tests pass. 6 new verification tests added for each fix.
Review-fold from the 3-angle simplify pass:
- sed -Ei / -iE / --in-place now match the shell-critical tier (the
bare '\s-i\b' token missed combined short flags and the GNU long
form); read-only sed stays unflagged. Regression tests added.
- agent_config_contract joins plugin_guard's CODE_EXEMPT_PATTERN_IDS:
content-contract prose in plugin code files (docstrings/comments)
is the same false-positive class the existing agent_config_mod
exemption suppresses. Doc/config files keep the full pattern set.
Efficiency reviewer: 1.24x full-scan cost (+3.4ms/file, install-time
only), worst-case adversarial line 55us — no ReDoS exposure.
Follow-up hardening on top of #92249's tiered scoring:
- Shell-critical tier now also catches tee, and cp/mv with the config
file in destination position (cp/mv reads and .bak backups excluded).
A single '>' redirect must be preceded by a word/quote character so
markdown blockquotes and '->' arrows no longer match.
- Prose tier catches mid-line imperatives behind directive markers
('you must modify...', 'please update...', 'make sure to append...'),
which previously bypassed the line-start anchor.
- Prose instructions aimed at AGENT config files score critical again:
project-skill quarantine acts only on 'dangerous', so high/caution
silently converted 'quarantined' into 'allowed' for exactly the
sentence shape persistence attacks use (concern raised in #88952).
Hermes/other-agent config prose stays high/caution (setup docs
legitimately instruct config.yaml edits).
- New content-contract tier ('AGENTS.md should contain ...') at
high/caution — the shape is shared by authoring guides and attacks.
- .claude/settings and .codex/config gain the same shell-critical tier.
Verified against a 595-skill corpus: 0 skills blocked by these tiers
(main blocked 44 legitimate ones), all mattpocock repro skills from
#92021 install, and 20/20 attack corpus lines keep their verdicts.
Enough1122's review on #92249 asked whether language write APIs
(Python open('w')/write_text/os.replace/shutil, Node fs.writeFileSync/
appendFile) are covered by the agent-config persistence tiers. They are
not: those tiers score shell redirection, sed -i, and imperative prose
only; language-API calls surface just the informational *_ref finding.
Static regexes cannot tie a dynamically-built path to the config-file
destination without executing the skill, so this is a documented scope
cut rather than missing coverage — runtime install gates remain the
backstop. Scoring behavior is unchanged, so SCANNER_VERSION stays at
skills-guard-v2 and cached verdicts remain valid.
Refs #92249
The skills-guard-v1 scanner flagged ANY mention of AGENTS.md / CLAUDE.md /
.cursorrules / .clinerules as critical/persistence. Any critical finding
forces a dangerous verdict, and community installs cannot be overridden
with --force — so legitimate meta-skills that merely DISCUSS agent config
files (authoring guides, setup docs, cross-references) were permanently
blocked. Three popular community skills were hit in the wild.
skills-guard-v2 scores the persistence category in three tiers:
- Mechanical persistence (shell redirection or sed -i targeting an agent
config file) stays critical -> dangerous. An unambiguous write path.
- Modification language in imperative position (verb at line/bullet start
within 80 chars of the filename) is high -> caution. Regexes cannot
separate "Edit AGENTS.md to inject instructions" from descriptive prose,
but imperative verbs are the shape real instructions take. Caution keeps
the install confirmable instead of irreversibly blocked.
- Bare references drop to low/informational for auditability without
driving the verdict.
The verb-proximity shape matches the existing convention in
tools/threat_patterns.py, and the tiering mirrors how allowed_tools_field
was already handled. The pattern id agent_config_mod is preserved so
plugin_guard.CODE_EXEMPT_PATTERN_IDS stays valid; hermes_config_mod /
other_agent_config get parallel _shell / _ref splits fixing the whole bug
class. SCANNER_VERSION bumps to v2 so cached v1 dangerous verdicts are
invalidated and re-scanned on next install attempt.
* fix(execute_code): limits line teaches spillover — big stdout is saved, not lost (follow-up to #97043)
* fix(execute_code): drop the editorializing tail from the limits line (maintainer review)
* refactor(execute_code): integrate kernel persistence into the core description (712 -> 654 tok/call, -8%)
* fix(execute_code): honest interpreter note — Hermes's own python is the common case; project venv only when VIRTUAL_ENV/CONDA_PREFIX is active
* refactor(code-execution): retire kernel_mode — session kernels always on for local runs (remote per-call is a tracked gap, not a mode)
* test(code-execution): env-filtering probes use reset=true — kernel env is frozen at spawn, so env rules are only observable on a fresh kernel
* test(code-execution): kernel-aware fixes for mode/pythonpath suites — reset=true on frozen-at-spawn probes, per-test kernel disposal, abort-after-capture fake Popen
* test(code-execution): strict-mode cwd is a behavior contract (staging tmpdir, not session cwd) — kernel stages in hermes_kernel_*, per-call in hermes_sandbox_*
cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
The forward dialed a hardcoded 127.0.0.1. The api_server adapter binds
extra.host -> API_SERVER_HOST -> 127.0.0.1, so mirror that chain when
dialing. Wildcard binds (0.0.0.0/::) listen on loopback, so keep dialing
loopback for those; bracket bare IPv6 literals.
The CLI has no 'trigger' subcommand ('trigger' is only an alias of the
cronjob TOOL's run action). Point operators at the real remediation:
start the gateway; its ticker owns relay-fronted delivery and fires the
job on schedule.
A manual 'hermes cron run' on a relay-fronted target has no live relay adapter
and no standalone sender, so it now forwards to the running gateway's
POST /api/jobs/{id}/run (marks due for the gateway ticker, which delivers via
the live relay adapter). Gateway unreachable -> the accurate 'start the gateway
or use cron trigger' error. Native topologies are untouched.
The two halves of tips were behind one switch, which meant the app
volunteering commentary at idle shipped on by default. Split them along
the line that matters: the rotation talks unprompted, so it now waits to
be asked for, while an agent tip stays ungated like the tour it mirrors
— Hermes raises one mid-conversation, in answer to something the user
said.
Drops the tool's config gate along with the config key it read. The
renderer mirrored that key with config.set, which has no branch for it
and answered "unknown config key" into a swallowed catch, so the opt-out
never reached the backend in the first place.
The quiet sibling of `tour`, in the same `desktop_ui` toolset and reading the
same `tour(action='targets')` discovery call: one bubble with an arrow, for a
sentence that would be clearer with a finger on the thing it's about. Dimming
the whole app to say "the model name is a button" is the wrong weight.
Fire-and-forget rather than a round-trip, because a tip is not a question and
blocking the turn on one would stall the reply it belongs to. The renderer
enforces the user's opt-out itself, so a stale config read can never put a
bubble on a screen that asked for none.
* feat(tools): session-persistent kernels for execute_code (kernel_mode: session)
execute_code spawns a fresh Python process per call, so every multi-step
data task re-loads its inputs: a CSV parsed in call one is gone by call
two, and scripts route state through temp files to survive. Hermes
already rewards programmatic tool calling (execute_code-only turns
refund the iteration budget), which makes the missing half — state that
survives between calls — the bottleneck.
Add opt-in `code_execution.kernel_mode: session`: one persistent kernel
per (task, mode, interpreter, cwd, tool-set). Variables, imports, and
loaded data persist across calls; `reset=true` discards state on demand.
The default `per-call` keeps today's behavior byte-for-byte.
Safety posture is unchanged by design: the child env comes from the same
builder as the per-call path (extracted, not duplicated, so the secret
scrubbing / PYTHONPATH hygiene cannot drift), the RPC server is the same
`_rpc_server_loop` with the same token and a per-cell tool budget, and
output passes the same ANSI strip + secret redaction. A timed-out or
interrupted cell kills the whole kernel tree and the next call respawns
— a wedged kernel can never hang the agent. The kernel env is frozen at
spawn; the schema and config comment say so.
Wire protocol: NDJSON requests on the kernel's stdin; responses framed
on stdout behind a per-kernel random sentinel, with unframed bytes
(fd-level output from user-spawned subprocesses) attributed to the
serialized current cell. The generated RPC client reconnects once when
HERMES_RPC_PERSISTENT=1, because a kernel legitimately outlives the RPC
server's 300s idle window between cells.
Tested on macOS 15 (Apple Silicon), Python 3.11: 13 new tests in
tests/tools/test_code_kernel.py (persistence, reset, error-keeps-kernel,
timeout-kills-kernel, sys.exit ends kernel, subprocess fd passthrough,
schema surface, mode fallback) plus the existing
test_code_execution.py / test_code_execution_modes.py suites (81 passed).
* fix(tools): session kernels get a stable owner, bounded lifetime, and per-cell RPC authority
Addresses the blocking review on the session-kernel design: two
authority/lifecycle boundaries were wrong.
1. Ownership and bounded lifetime. The kernel key's first component is
now the conversation's approval session key (_resolve_owner), not the
per-turn task id run_agent mints per top-level invocation — so state
genuinely survives across user turns of one conversation, and delegated
subagent sessions isolate naturally under their own keys (the task id
remains only the last-resort owner for embeds/tests with no session
context). Lifetime is bounded on four edges: kernels are disposed at the
same session boundary that clears the owner's approval/yolo state
(tools.approval.clear_session -> shutdown_kernels_for_owner), reaped
after code_execution.kernel_idle_timeout seconds idle (default 1800,
swept on every entry), capped process-wide at
code_execution.max_session_kernels live children (default 4, LRU
evicted), and still torn down by reset/death/atexit as before. The
ownership + disposal + idle-reap + cap shape deliberately carries
forward the lifecycle invariants of the earlier session-persistent
implementation in #88637 by @z80dev.
2. Per-cell RPC authority. The serving thread no longer freezes the
spawning cell's context/callbacks for the kernel's life. Each cell
installs a CellAuthority — captured on the calling thread exactly as
propagate_context_to_thread would for a per-call RPC thread — before its
request is written, and retires it on every settle path; _rpc_server_loop
gains a dispatch hook the kernel uses to route each tool call through
the CURRENT cell's context, callbacks, and task id. A call arriving with
no active cell is refused. Interpreter state persists; RPC authority
does not.
Composition with the per-script static guard (see the config note): a
persistent namespace lets cell N+1 invoke objects cell N created, which
a single-cell static scan cannot see — the runtime RPC boundary
(allow-list by name, per-cell budget, per-cell authority) is the
operative cross-cell enforcement in this mode, and the adversarial
alias test pins exactly that.
Tests (9 new): state survives across turns of one conversation;
sessions isolate; clear_session disposes the owner's kernels (and the
next turn starts fresh); the live-kernel cap LRU-evicts with evicted
children proven dead; idle kernels are reaped; a later cell's RPC runs
under that cell's approval callback; a cross-cell alias dispatches under
the CURRENT cell's authority; a settled cell's authority refuses
dispatch; each cell installs a fresh authority. 22/22 kernel tests, 81
code-execution tests, ruff clean. The 7 test-order failures in the
tools/-k-approval selection reproduce identically on the clean branch
base (pre-existing pollution, not this change).
* fix(code-kernel): delegated children get their own kernels — child contexts inherit the parent approval key, so qualify the owner with the delegation session id (live-verified leak, both directions)
---------
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
Both from the simplify pass: the comment kept only the ownership-relevant
rationale (incl. the no-double-bump note); the test's _ReadyAdapter was a
verbatim delegate around threading.Event — the exercised paths only call
is_set/clear/set, so the bare Event is behaviorally identical.
Widen the contributor's pre-call reconnect signal to the sibling site:
when the stdio subprocess dies mid-RPC the watcher race fast-fails, but
nothing cleared server.session, so the server stayed dead until the idle
keepalive probe noticed. Signal the reconnect there as well.
Also drop the explicit _bump_server_error at the pre-call gate: the
returned error payload already flows through the handler's JSON parse,
which bumps the breaker once — the explicit bump would double-count.
Two regression tests pin both sites (reconnect signaled exactly once,
no RPC attempted on a dead transport, single breaker bump).
_stdio_children_dead() returned True ('all children dead') when a child
process was ALIVE — the liveness predicate was inverted:
for pid in pids:
if not psutil.pid_exists(pid):
continue # dead — skip
return True # BUG: an ALIVE child reported as 'all dead'
Consequence: every tools/call on a stdio MCP server with a healthy
subprocess failed instantly (~0.01-0.3s) with
'MCP stdio subprocess has exited; failing the call fast' (#81995
fast-fail path), while hermes mcp test kept passing (it never reaches
tools/call). Servers appeared dead regardless of restarts.
Fix: return False as soon as one tracked child is alive; True only when
every child has exited:
for pid in pids:
if psutil.pid_exists(pid):
return False # at least one child alive
return True
Also: on a genuinely-dead stdio (session object still present), signal
reconnect instead of a bare fast-fail TimeoutError, so the manager
respawns the subprocess instead of stranding the call slot.
Verified: alive child -> False, all exited -> True, no tracked pids ->
False (unknown, don't fail fast).
When send_message is invoked from the agent's worker thread (a different
event loop than the gateway's), awaiting the WeCom adapter directly can hang
because the adapter enqueues onto the gateway loop. Dispatch via
run_coroutine_threadsafe onto the gateway loop when the caller loop differs,
with caller-cancellation shielded so an already-enqueued send is not cancelled
mid-flight (which would otherwise cause a false-failure retry -> duplicate).
Recognizes WeCom native chat IDs as explicit send targets and whitelists WeCom
for media delivery. Part of the async queue design this branch introduces.
Builds on e11187208f (salvage of #72206 by @luyifan, authorship
preserved) which added the post-timeout session reset. This commit
completes the Phase 3b contract:
- suspect flag: a command timeout marks the session key suspect;
the NEXT _get_session_info for that key health-checks and recycles
the session instead of handing back the poisoned handle (#72205)
- flag cleared on fresh-session store so a healthy new session is
never spuriously recycled; cross-key leakage fixed
- wedged-vs-alive rule (#68139): after a timeout, if the daemon's
socket still answers, recycle the session only; if it is
unresponsive (or no socket exists to probe — conservative), tree-
kill via agent.deadline.kill_process_tree and evict
- negative probe: successful commands never mark or recycle
Tests: tests/tools/test_browser_suspect_recycle.py (20 tests:
mark-once, recycle-then-succeed, success-never-recycles, tree-kill
invoked on the wedged path with pid assertion, flag lifecycle).
Co-authored-by: luyifan <al3060388206@gmail.com>
The smart-approval guardian (`_smart_approve`) gates every flagged
terminal command with a synchronous auxiliary LLM call, but it never
passes `timeout=` and logs nothing on the normal path. In production a
stalled provider response silently froze the agent turn for 62 minutes
with zero log output; the gateway kill-switch eventually fired, and only
an unrelated error surfaced afterwards (#82846; watchdog-style fix in
#72500). The call was invisible by design — nothing logs at the hang
point.
Changes in tools/approval.py:
- Resolve the same configured timeout the client would use internally
(`auxiliary.approval.timeout` via `_get_task_timeout("approval")`) and
pass it explicitly to `call_llm`, so the deadline cannot be lost if the
internal default resolution changes or is misconfigured.
- Log the assessment call and its duration (DEBUG), and promote the
failure branch from DEBUG to WARNING with elapsed time + exception
class, so a wedged guardian call is visible in the logs instead of
silent.
- Failure still returns "escalate" (fail open to the human/pattern
gate) — behavior unchanged, observability only.
Complements #72500 (watchdog hard ceiling) rather than duplicating it:
explicit timeout is the root-cause hardening, logging closes the
silence gap; the watchdog remains the safety net if the SDK-level
timeout itself is defeated.
Tests: explicit timeout forwarded to call_llm (revert-fails), failure
logs WARNING + escalates. 49 approval-adjacent tests pass; one unrelated
test_approval.py failure is pre-existing (fails on clean main too).
Apply-ready delta distilled by @andrexibiza: deterministic watcher-consumer
tests (watcher times out while a child is alive, resolves when all are dead),
psutil-unavailable fail-open pin, and probe-failure fail-open handling in
_stdio_children_dead (unknown is never proof that every child exited).
Local: 8 passed on tests/tools/test_mcp_stdio_children_dead.py
_stdio_children_dead returned True ('all children dead') on the first LIVE
pid — the intended False was dead code right below it. Every spawn path
that captures child PIDs (observed in hermes -z oneshots) then failed the
#81995 pre-call fast-fail with 'TimeoutError: MCP stdio subprocess ... has
exited' on every tools/call while the subprocess was demonstrably alive.
Long-lived gateway/dashboard sessions were unaffected only when
_stdio_child_pids was empty (the not-pids short-circuit).
Return False on the first live pid and drop the unreachable line.
Refines the Windows path per three requirements:
1. Only when the toggle is set — closing is offered only if
browser.real_profile_autoclose is on.
2. Blocked when locked — snapshot_real_profile NEVER kills; a locked profile
always returns the [profile-locked] signal and the copy is refused. A later
attempt that is still locked blocks again (no loop, no auto-kill).
3. Ask approval to close — closing is an explicit, user-approved step:
(new CLI subcommand) runs
close_browser_holding_profile only when the agent has the user's OK. The
locked error tells the agent to ask first, then run it, then retry.
- browser_connect: snapshot blocks with _PROFILE_LOCKED_PREFIX (autoclose-armed
message offers the close; off message says fully-quit); no in-snapshot kill.
- main.py: subcommand (identity+binding-verified
tree kill via close_browser_holding_profile); added to _BUILTIN_SUBCOMMANDS.
- browser_tool: surfaces the locked signal + the exact approved-close command.
- Docs/config: toggle arms + agent asks + blocked-if-still-locked.
Tests: snapshot blocks-not-kills with autoclose on AND off; process matcher
identity/binding. 73 real-profile tests pass. Windows live E2E (proof): locked
blocks fast without killing → approved close terminates Chrome → snapshot then
copies a valid DB; autoclose-off blocks with quit guidance.