hermes chat -q sets HERMES_INTERACTIVE=1 (for interactive sudo prompts) but
runs one turn with no user waiting to answer approval prompts. Previously a
dangerous command triggered the interactive gate, waited the full 300s
timeout, then failed closed — and the agent was effectively forced to work
around the block, often silently auto-approving via execute_code (which
auto-approves in non-gateway mode).
Add approvals.single_query_mode (default deny, mirror of cron_mode):
deny — block dangerous commands and execute_code deterministically with
a clear 'no user present' message (no 300s wait)
approve — auto-approve dangerous commands/execute_code in -q mode
cli.py marks the session with HERMES_SINGLE_QUERY_SESSION; the shared gate
(_run_approval_gate, check_all_command_guards, check_execute_code_guard)
treats -q as a deterministic non-interactive context when that marker is set.
execute_code, the -q escape hatch, now honors single_query_mode instead of
auto-approving headlessly. Includes tirith parity in the combined guard and
docs. Fixes#86878.
An MCP server exposing a native tool named read_resource (or
list_resources/list_prompts/get_prompt) collided with the auto-generated
resource/prompt utility of the same name. The registration collision
handler flagged the pair as ambiguous and skipped BOTH entries, so the
server's own tool became silently unavailable on every gateway boot.
Resolve this specific native-vs-utility collision in favour of the native
tool: keep it and drop the shadowed utility, which is only convenience
sugar for servers that expose no such tool of their own. The conservative
skip-everything path still applies to genuinely ambiguous collisions (two
or more native tools normalizing to one name), which we cannot
disambiguate. Add a regression test covering the native-tool-wins path.
Fixes#87112
`cargo build 2>&1 | tail -20` exits with tail's 0 even when the build
failed — bash without pipefail reports the last pipeline command's
status, and `cmd || echo failed` swallows the status the same way. The
model reads exit_code: 0 as a strong success signal and can conclude a
build passed while the visible output says it failed (community report,
Windows Rust builds; not platform-specific).
Two-part fix, mirroring OpenCode's prompt-side approach plus a
result-side backstop they don't have:
- Tool description now forbids piping builds/tests through
tail/head/cat (output is already auto-truncated + spilled to a file)
and warns that pipes/|| fallbacks mask exit codes.
- New annotate_masked_success() in tools/terminal_hints.py: when
exit_code == 0, the command shape can mask an upstream status
(top-level pipe into a passthrough consumer, or || echo/true), AND
the output carries strong tool-specific failure shapes (rustc,
cargo, pytest, gcc, npm, make, ninja), attach an advisory 'hint'
telling the model to treat the run as failed and re-run bare.
exit_code itself is never modified. Search/content heads
(grep/rg/echo/printf/...) are excluded to avoid false positives on
pipelines whose output legitimately contains error text.
E2E-verified through the real terminal tool path: hint fires on masked
cargo-style failures, silent on bare commands, clean pipes, and
grep/printf pipelines. 42 targeted tests pass.
Some MCP OAuth providers (notably Supabase) return a client_secret from
dynamic client registration but omit token_endpoint_auth_method. The MCP
SDK defaults the missing method to "none", so the token exchange omits
client_secret and the server rejects it (HTTP 422 "Required parameter:
client_secret"), looping the browser consent page.
This resolves the whole class, not just one provider:
- Storage layer (HermesTokenStorage): coerce secret-bearing client info
with missing/none auth method to client_secret_post on both read and
write, persisting the corrected shape.
- Both live provider paths (tools/mcp_oauth.py HermesOAuthClientProvider
and tools/mcp_oauth_manager.py HermesMCPOAuthProvider): coerce
in-memory client info immediately before token exchange and refresh.
- Accept the full 2xx range on token and refresh responses (Supabase
returns 201 Created), instead of the SDK's exact-200 check.
- Redact token response bodies from error messages and logs on
malformed responses.
The Figma-specific request-time default (apply_oauth_provider_defaults)
remains; this generalizes the same bug class for every DCR provider.
Fixes#29680. Supersedes #34274 and #35700 (201-only variants).
KDE/Qt apps report [0,0,0,0] bounds for elements that are perfectly
clickable by index (live QA: all 50 of kcalc's zero-rect elements,
including every radio button). Serializing that as a plausible rect
invites a model to derive coordinate=[0,0] and click the screen corner.
- _element_to_dict: zero rect -> bounds: null
- _format_elements: '@ bounds-unknown (click by element index)' instead
of the fake rect in the summary line
- malformed bounds fail open (unchanged serialization)
Live-proven on real kcalc (cua-driver 0.20.0): 50 elements now null, 0
zero-rect leftovers, summary annotated, real rects preserved, and a
null-bounds radio button still clicks fine by index.
Live complex-action QA on a real KDE desktop (kcalc + kate multi-app
flows) found two dispatch gaps:
1. Wrong-window input reported as success. Input actions deliver to the
backend's sticky target (last capture/focus_app); the app= argument
models routinely pass on the input call itself was silently dropped.
Proven live: with kcalc sticky, type(text='777', app='kate') returned
ok:true and typed 777 INTO KCALC. New guard: provable mismatch
(both names known, neither substring of the other — list_windows
names are localized/variant) refuses with input_target_mismatch and
a one-call fix instruction. Unknown current target fails open so
legacy no-app flows are untouched.
2. Near-miss unknown actions were dead ends. A model emitting 'hotkey'
got a bare unknown-action error. Suggestion map now names the real
action ('did you mean key?') without aliasing — we never repair bad
model output, we just point at the schema.
Also documents the verified-lost-keystroke rung in the computer-use
skill: KTextEditor (Kate/KWrite) discards synthetic X keystrokes at the
toolkit level — foreground type reports ok but AX shows nothing arrived,
and a raw XTest control fails identically outside our stack. Guidance:
after one verified-lost round trip, switch to file/DBus I/O instead of
looping the ladder.
Live proof on the fixed build: mismatch refused, kcalc display clean,
same call after capture(app=kate) succeeds, 'hotkey' suggests 'key'.
11 new tests; 158 sibling tests green.
A headless Mac or asleep built-in panel leaves ScreenCaptureKit with 0
shareable displays while TCC grants pass — health_report stays ok and
every capture silently returns 0x0 (#67165). Guard at the report seam
(_apply_display_count_guard, both real and fallback paths): flips the
screen_capture_capability check to fail with recovery actions (wake
display / HDMI dummy / virtual display) and downgrades ok -> degraded.
The empty-discovery reason ladder gains the matching darwin rung.
Composed from #52949 (sujeet111) and #67259 (webtecnica); both PRs
predate the doctor rewrite and the envelope normalization on main, so
this reimplements their shared intent at the current seams.
Co-authored-by: Sujeet <64351924+sujeet111@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
capture() read any non-None pid/window_id as a request for exact-window
targeting. Several providers emit every declared schema property on every
tool call, zero-filling unused optional integers, so those calls arrive as
pid=0, window_id=0. The exact-target branch was then entered, the caller's
app= was discarded, _positive_int(0) returned None for both ids, and the
capture failed with a message pointing at pid/window_id. For that class of
model capture(app=...) and frontmost capture never worked at all.
Normalize non-positive ids to None before the branch decision so dispatch
falls through to app/frontmost discovery. Malformed non-numeric ids are
deliberately not treated as placeholders: they still reach the existing
validation error instead of being silently ignored.
Fixes#81333
Two live-QA findings from a locked KDE desktop (real cua-driver 0.20.0):
1. capture() with zero discovered windows returned a bare
'capture mode=ax 0x0' — no hint that the desktop session was LOCKED,
which freezes renderers and hides windows. New
_empty_discovery_reason() names the dominant causes in order: locked
session (loginctl LockedHint probe, fail-safe), missing DISPLAY,
else a pointer at hermes computer-use doctor. Surfaced through the
existing window_title -> summary path, so the model and the user see
it inline.
2. _call_tool_via_cli retried 'daemon is not running' 4x with ~3.5s of
backoff sleeps — a permanent condition for that invocation (the CLI
transport needs the machine-wide daemon; Hermes' MCP runtime does
not). Now fails fast on the first attempt with a message naming the
split. Transient empty output (EAGAIN congestion) keeps the retry
loop — pinned by test.
Live-verified on the locked desktop: capture now reports the lock and
the unlock action; CLI fallback errors immediately with the transport
explanation. 7 new tests; 191 sibling tests green.
cua-driver >= 0.17 splits the semantic_v2 snapshot payload: action-bearing
refs live in the `refs` array while `content_refs` carries every node
with EMPTY action lists. _ref_map only absorbed content_refs, so every
click/pointer/type ref was registered with no declared actions and all
typed-browser mutations failed with browser_ref_stale.
Merge refs + content_refs + snapshot.refs with set union so action info
is never dropped by an empty content entry. Verified against the 0.17
split format, the legacy refs-only format, the transitional dict format,
and the snapshot.refs fallback.
Live-tested against the real cua-driver 0.19.3 binary (Linux x86_64):
- bounded serve flags corrected: the daemon accepts
--session-policy/--approve-session-policy, not the docs'
--capability-manifest names (which it rejects). Verified end-to-end:
a bounded daemon with a real policy file starts and reports running.
- browser-approve verified real but interactive-only (refuses without a
TTY) and its token is a legacy compatibility path disabled by default
on current drivers (per the live browser_prepare schema). Kept as a
passthrough; no longer presented as the primary route.
- NEW primary standard-mode route, verified live: launch the runtime
with cua-driver's trusted-launcher grant. config opt-in
computer_use.grant_existing_profile: true appends
--grant existing-profile to the standard-mode MCP spawn (MCP
initialize verified accepting the flag). Default false = attachment
keeps failing closed. Never applied to bounded/unrestricted daemons.
- Skill, system prompt, tool schema, and docs updated to the verified
ladder: config grant > bounded manifest > YOLO; token = legacy.
Completes the typed cua_browser_* route (PR #74166 lineage) with the
authorization surface that makes existing-profile attachment and
repeatable bounded automation reachable by real users:
- hermes computer-use browser-approve: CLI passthrough that mints
cua-driver's five-minute single-use attachment token for one exact
(pid, window_id). The user, never the model, is the token source.
- approval_token passthrough on cua_browser_prepare (schema + dispatch +
browser_route), forwarded only for existing_profile and only as a
non-empty string.
- computer_use.permission_mode: bounded + capability_manifest config:
private per-session embedded daemon launched with
--capability-manifest/--approve-capability-manifest; missing manifest
fails loudly. 'unrestricted' is deliberately NOT a config value —
it stays bound to the explicit per-session YOLO toggle.
- Skill + system-prompt + docs guidance for the three authorization
rungs and the isolated-profile-first default.
E2E-verified against a temp HERMES_HOME: real config resolution to
bounded, loud failure without a manifest, real argparse path driving a
fake cua-driver binary, standard default preserved.
Follow-up to #86916. That fix gave named sessions their own daemon
(socket/log/pid) and their own provider browser — but on a SHARED local
Chrome / CDP browser, a fresh named daemon still attaches to the first
existing page, the same page a sibling daemon may hold. A named session
that never calls new_tab() could still stomp another's tab.
browser_exec now prepends a small preamble to the model's code for named
sessions on shared browsers: once per daemon process (marker keyed by
uid + BU_NAME + daemon pid), it creates a fresh tab via
Target.createTarget and switch_tab()s onto it before any model code
runs. Private per-name browsers (provider-keyed bu-named-<name>, or
direct-API Browser Use cloud) skip the preamble via an internal env
sentinel popped before launch — there's nobody to collide with, and the
extra tab would leak.
Best-effort by design: if the preamble's CDP calls fail, behavior
degrades to pre-fix, never blocks the exec.
E2E against a shared headless Chrome with the STOCK harness: two named
sessions issuing bare js() writes (no new_tab) kept distinct state
(EDGE-A/EDGE-B read back intact); the sabotage run without the preamble
reproduced the clobber (both read EDGE-B). Removes the dependency on the
upstream browser-harness tab-isolation PR for correctness.
session=<name> previously set BU_NAME and then skipped backend resolution
entirely — the parameter was documented as cloud-only, so all local/CDP
work funneled through the single default daemon and one IPC socket, and
concurrent sessions (parallel subagents, simultaneous chats) clobbered
each other's browser connection. Reported by @shantanugoel on X.
Now a named session composes with whatever browser source is configured:
- BU_NAME still namespaces the harness daemon (per-name IPC socket, log,
pid — upstream already isolates these), for local Chrome and CDP.
- The /browser connect CDP override is now exported for named sessions
too; previously a named daemon ignored it and fell back to scanning
local Chrome profiles.
- On provider backends (Browserbase, Firecrawl, Nous gateway), the name
keys its own provider browser via the shared _get_session_info cache
(bu-named-<name>), so each name gets its own cloud browser, the same
name reuses one across calls and tasks, and unnamed calls keep the
per-task key.
- Direct-API Browser Use cloud configs keep the native named-daemon path
(provider resolution would double-session and double-bill).
Tool schema/description updated so models reach for session=<name> for
parallel work on any backend, not just cloud.
E2E: two named sessions against a real headless Chrome (real browser-use
CLI, BU_CDP_URL) ran concurrently, set distinct page state, and read it
back intact; sabotage run confirms the new tests fail without the fix.
Follow-up to the salvaged #86862: surface reclaim counts at warning level
(mirrors the scheduler tick's reap handling from #86853) and keep a debug
trace when the best-effort recovery itself fails, instead of a bare pass.
Fixes#86721.
`hermes cron run <job_id>` (a one-shot CLI invocation) dispatches
manual runs via the same background-delegation path as an agent's
`cronjob(action='run')` tool call (tools/cronjob_tools.py's
_try_dispatch_background_run -> dispatch_async_delegation(role=
"cron_run", runner=_runner, ...)). The runner thread lives in the
calling process's shared daemon executor. When the one-shot process
exits right after printing "Triggered job: ...", the in-flight runner
dies mid-execution, leaving its cron/executions.db row permanently
stuck at status='claimed' -- every subsequent `hermes cron run` on the
same job then reports "Ran now: failed" because of the still-claimed
row.
cron/executions.py already has the exact self-heal this needs:
recover_interrupted_executions() correctly identifies and reclassifies
'claimed'/'running' rows whose owner process has provably exited
(_owner_is_live checks PID existence AND matches process start-time,
so a reused PID isn't mistaken for the original live owner) to
'unknown', unblocking the job for a fresh claim. But it was only ever
called once, at the long-lived scheduler ticker's own startup
(cron/scheduler.py:379's self.recover_interrupted()) -- a one-shot CLI
invocation has no equivalent "startup" moment of its own, so this
self-heal never ran for it.
Added a call to recover_interrupted_executions() at the top of
_try_dispatch_background_run, right after the async-delivery-supported
gate and before any claim attempt for the current job -- mirroring
exactly what the long-lived scheduler already does at its own
startup, just triggered per one-shot invocation instead of once at
daemon startup. Wrapped in try/except: pass (best-effort; a failure
here must not block the actual dispatch this function exists for).
Traced (but did not attempt to fix) the deeper "why does the runner
die with the process at all" question -- that's the harder problem
options 1/2 in the issue describe (route to the persistent scheduler,
or block the one-shot process until completion). This fix addresses
the more urgent, more clearly-scoped symptom: a stranded stale claim
permanently blocking ALL future manual runs of the affected job, which
is option 3 from the issue and the one with an existing, already-
correct implementation just needing to be wired into this call site.
Added 3 regression tests to a new file, following the established
real-subprocess dead-owner pattern already used in
tests/cron/test_execution_ledger.py (a genuinely-dead PID, not a
mock, matching the real-world failure mode exactly): a sanity test
confirming the stale claim sits unrecovered without the fix; a direct
test of recover_interrupted_executions() reaping such a claim; and a
unit test on _try_dispatch_background_run itself confirming recovery
is called before any claim attempt. Verified as a genuine regression
by reverting the fix and confirming the unit test fails with recovery
never having been called.
35/35 pass across the new test file plus tests/cron/test_execution_ledger.py
and tests/tools/test_cronjob_run_background.py (no regression).
Main extracted a session_search() wrapper (owned-DB lifecycle) around
_session_search_impl after #82595 was opened; the cherry-picked detail
parameter landed on the impl only. Append it to the wrapper with the
same positional-compatibility contract and pass it through.
Creating a project whose resolved primary path already belongs to a
non-archived project now raises a clear ValueError naming the existing
project (create_project) — duplicated projects each seeded an identical
copy of the repo subtree, multiplying the duplicate-lane bug per copy.
The agent-facing project_create tool is idempotent instead: it re-activates
the existing project rather than erroring. allow_duplicate_path=True keeps
deliberate duplicates possible. Also updates the legacy non-git lane-id
expectation to the branch-style id introduced for #53329.
The orphan reaper had two gaps that let agent-browser daemons accumulate
indefinitely inside a single long-lived hermes process:
1. `_reap_orphaned_browser_sessions()` ran exactly once, before the cleanup
loop started, so a leak appearing after boot could never be recovered.
2. `owner_alive is True` skipped unconditionally. In-memory session tracking
is lost on any exception path between spawn and registration, but the
owner PID stays up — so such a daemon was skipped forever.
The daemon-side `AGENT_BROWSER_IDLE_TIMEOUT_MS` is not a backstop for (2):
it does not fire when the daemon itself is wedged, e.g. after Chrome's
framework was replaced underneath it by an auto-update.
Observed on macOS: five agent-browser daemons (96 Chrome processes) built up
over 10 days inside an 18-day-uptime hermes process, holding roughly 5 CPU
cores busy and driving the load average past 100. Four of those processes
were still running a Chrome framework version that had since been replaced
on disk, spinning at ~85% CPU each.
Changes:
- Re-run the reaper every `BROWSER_ORPHAN_REAP_INTERVAL` (300s) from inside
the cleanup loop. Cycle 0 preserves the existing startup reap.
- When the owner is alive but the session is untracked, fall back to idle
age: reap past `BROWSER_ORPHAN_GRACE_SECONDS`, defined as
`max(1h, 20 x inactivity_timeout)`. Unknown age fails safe.
- Add `_socket_dir_idle_seconds()` — the newest mtime under a session's
socket dir. Every browser command writes `_stdout_<cmd>` / `_stderr_<cmd>`
there, making it a last-activity marker that survives hermes restarts and
does not depend on in-memory bookkeeping surviving an exception path. It
scans directory entries rather than reading the directory mtime alone:
command names repeat, and rewriting an existing `_stdout_click` updates
that file's mtime but not the directory's, so a dir-mtime-only check would
report a busy session as idle and reap it.
Sessions still present in `_active_sessions` are never touched at any age,
and the new path still goes through `_verify_reapable_browser_daemon`, so
the anti-spoof / anti-PID-recycle guarantees from #14073 are unchanged.
Adds 9 tests: idle-age unit tests (including the dir-mtime regression),
spared/reaped/fail-safe cases for a live owner, the identity-guard gate on
the new path, and a periodic-reap test asserting more than one reap per
cleanup-thread lifetime.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The vision tools' call_kwargs hardcode max_tokens caps (2000 for
vision_analyze/browser_vision, 4000 for video analysis), truncating
descriptions of complex images at the cap. The centralized aux client
already omits max_tokens by default (#34845) so providers use their
model max output; these three call sites were the leftovers that
bypassed that policy.
Remove the hardcoded caps entirely — the aux client handles the
mandatory-max_tokens Anthropic wire via _resolve_anthropic_messages_max_tokens
(model output ceiling) and Gemini native omits maxOutputTokens (65K ceiling),
so no wire needs an explicit cap.
delegation.max_concurrent_children caps how many delegated children run in
parallel per batch (and concurrent background delegation units). The old default
of 3 needlessly serialized independent fan-outs (e.g. reviewing/investigating N
PRs or issues at once), so large batches ran in slow chunks of 3.
Raise the shipped default to 10, which sits at/below the existing high-cost
advisory threshold (>10), so the default never trips the warning. Each child
still consumes API tokens independently, so this is a throughput/latency win the
user pays for in parallel token spend — the floor stays 1 and there is no
ceiling, so anyone can tune it down or up.
- config_defaults.py: default 3 -> 10; _config_version 36 -> 37.
- delegate_tool.py: _DEFAULT_MAX_CONCURRENT_CHILDREN 3 -> 10 (+ docstring).
- config_migrations.py: _migrate_to_37 lifts configs pinned at exactly the old
default 3 to 10 (deliberate non-3 overrides preserved; unset inherits 10).
- cli-config.yaml.example: documented default updated.
Verified: default/fallback read 10, version 37, and the migration lifts 3->10,
preserves an explicit 5, and leaves unset untouched.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
The guard's "use a separate worktree or temporary clone" advice sent
agents to /tmp by default. /tmp is RAM-backed tmpfs on most distros, and
parallel salvage clones each running npm ci (~1.6GB per clone) filled a
32GB tmpfs to 97% during a 15-subagent campaign, ENOSPC-ing sibling test
runs. The message now recommends `git clone --shared <root> ~/.hermes/scratch/<task>`
(honoring HERMES_HOME), warns that dependency installs belong on real
disk, and tells the agent to delete the clone once the branch is pushed.
Follow-up on the salvaged #85764 commits, addressing review findings:
- _session_left_live_context now allowlists end_reason == 'compression'
or a fresh reset (_FRESH_RESET_END_REASONS) instead of accepting any
non-None end_reason. The wide predicate let 'branched' parents — whose
transcript /branch verbatim-copies into the child — surface as
same-lineage recall hits, returning content already in the caller's
live context (verified empirically vs main).
- _FRESH_RESET_END_REASONS is now derived from the canonical
hermes_state_common._RESET_END_REASONS (plus CLI 'new_session') instead
of a third hand-maintained copy, per that tuple's anti-drift comment.
Import verified cycle-free.
- Browse drops the Python re-check of parent_session_id rows:
list_sessions_rich (include_children=False) already applies the
canonical _LISTABLE_CHILD_SQL classifier, and the Python re-check
re-hid legacy pre-marker reset children the SQL same-key heuristic
deliberately admits. _has_reset_from_marker (now orphaned) removed.
- Tests: branched-parent exclusion regression guard (mutation-checked:
fails on the overbroad predicate) + legacy pre-marker reset child
browse guard. 48/48 pass.
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).
- Close partially initialized SessionDB connections on every constructor
exception path via a finally block guarded by an initialization-complete
flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
thread, close on worker exit, reject new enqueues after shutdown starts,
and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
API disconnect failures, shutdown recovery, RetainDB late enqueue, and
foreign-loop async clients.
Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A bare SessionDB() resolves the launch profile's default state.db, but
parents can hold non-default per-profile handles (tui_gateway opens
SessionDB(db_path=<profile_home>/state.db) for non-launch profiles and
hands them to agents via _transfer_db_to_agent). A child of such a
parent would write its transcript into the WRONG database — cross-
profile leakage that breaks parent_session_id lineage and
session_search. Open the dedicated handle at the parent handle's
db_path instead (AsyncSessionDB forwards .db_path via __getattr__, so
the gateway wrapper path works too). Regression test verified RED on
the pre-fix code.
Review follow-up (cc3f18197): if AIAgent() raises inside _build_child_agent
the freshly-opened dedicated handle has no owner and no child close() will
ever run — release it on the exception path so the sqlite fds don't
outlive the failed spawn. Also pin the degradation contract with a test:
a parent without a SessionDB still yields session_db=None children.
Cron run_job closes its per-job SessionDB in its finally block while a
fire-and-forget background delegation subagent is still flushing on a
daemon thread. The child shared the parent's SessionDB object, so every
subsequent flush hit the closed handle ('NoneType' object has no
attribute 'execute') and the child's whole transcript was silently
dropped. The same teardown-while-child-alive shape exists on gateway
session end and /new mid-delegation.
Each child now opens its own SessionDB connection (owned flag set at
construction so child.close() releases it), so no parent teardown can
close the child's handle out from under it.
Regression test proves the child gets a distinct live handle that
survives the parent's close().
Teknium review on #75732: releasing the pending clarify whenever
resolve_text_response_for_session returned False also cancelled
retryable multi-select invalid selections (out-of-range numbers,
unrecognised comma-lists).
Classify rejected typed replies in clarify_gateway:
- rejected_prose → cancel clarify, fall through busy routing (deadlock break)
- rejected_selection → keep clarify armed so the user can retry
Add native multi-select gateway regressions for both paths.
Gateway /new ends the predecessor with session_reset and leaves the child
empty, but discovery/browse/scroll still treated the whole lineage as
already in context. Allow ended reset parents through and keep live
delegation children excluded.
A delegated subagent that exhausts its per-child iteration budget
(delegation.max_iterations) still returns a summary, so the result carries
status='completed' even though the child's exit_reason is 'max_iterations' and
its work was cut off mid-task. The parent then reads 'completed', trusts the
partial summary, and only discovers the truncation by parsing the prose (where
the child happens to mention 'hit the iteration limit'). That wastes parent
turns and risks acting on incomplete work.
exit_reason is already computed authoritatively and threaded to every
parent-visible surface; it just wasn't reflected anywhere the parent reads at a
glance. This surfaces it:
- delegate_tool.py: add a parent-visible boolean 'truncated' (= exit_reason ==
'max_iterations') to each task entry, alongside the existing exit_reason.
- process_registry._format_async_delegation: for both the batch and single-task
paths, when truncated -> use a warning icon, append
'TRUNCATED: hit max_iterations — work may be incomplete' to the header/Status
line, and prefix the summary with an unmissable truncation notice. status
semantics are left unchanged (stays 'completed') so existing icon/summary
branch logic and ~10 tests asserting status=='completed' stay valid.
Tests: single-task truncated -> banner; single-task clean -> no banner; batch
marks only the truncated task, not its clean sibling. 23/23 in the async-
delegation suite.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
The post-write lint runs `npx tsc --noEmit <file>` on a single .ts file with
no `-p tsconfig`. tsc ignores tsconfig.json for explicit file args, so for any
file in a real TS project it floods phantom diagnostics — unresolved path
aliases (@/... -> TS2307) and ambient globals (Window.hermesDesktop -> TS2339)
that the project config defines. The delta filter sees the same phantom errors
pre- and post-edit, finds no NEW ones, and returns the misleading
'pre-existing lint errors ... the file is still broken' on a perfectly correct
one-line edit — wasting the caller's turns chasing nonexistent breakage.
Existing code already skips shell tsc when an LSP server claims the file, but
that only fires with LSP configured+enabled (not the default), leaving the
common LSP-disabled case fully exposed.
Fix: when an ancestor tsconfig.json exists (local host only), skip the per-file
shell tsc for .ts — its verdict carries zero signal for project files. Real
diagnostics still come from the LSP tier or an explicit `tsc -p tsconfig.json`.
Best-effort ancestor walk; remote/sandbox backends fall back to running the
linter as before. (.tsx already returns skipped via the ext-not-in-LINTERS
branch.)
Tests: ancestor-tsconfig .ts -> skipped even with LSP off; standalone .ts with
no ancestor tsconfig -> shell tsc still runs. 7/7 in the LSP-skip suite.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
memory replace/remove target an entry by old_text but supply the replacement
via content — an asymmetric pairing. Callers naturally reach for new_text to
mirror old_text (it's exactly the patch tool's old_string/new_string shape),
which left content empty and errored 'content is required'. The failure also
rendered tersely, making the cause easy to miss and costing a retry.
Accept new_text as an alias for content on both shapes:
- single-op: coalesce content = content or new_text in memory_tool() + a
new_text param wired through the registry handler.
- batch ops: content = op.content or op.new_text (and in the approval-gate
preview builder).
- schema: document the alias on content and add a new_text property to the
single-op params and batch item props so strict validators accept it.
content wins if both are set.
Tests: new_text alias on single add/replace, batch add/replace, and
content-wins-over-new_text. 43/43 across memory tool + schema suites.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
Session-boundary cleanup (gateway/run.py: run-finalization and prompt
delivery-failure paths) calls clear_session to cancel pending clarifies.
It unconditionally overwrote every entry's response with the empty
cancellation sentinel, even entries already resolved by a button callback
or text intercept. A waiter that had already observed the resolved event
would then return the empty sentinel instead of the real answer, silently
discarding the user's response on /new, gateway shutdown, or cached-agent
eviction.
First-writer-wins: clear_session now cancels only entries whose event is
not yet set; already-resolved entries keep their response. Mirrors the
guard resolve_gateway_clarify gained for the same contract. Reported by
doryani-ai on PR #75732; regression test covers the button-then-cleanup
interleaving.
- The gateway api_server fire webhook acknowledges 202 only after a
durable claim + execution row exist (admission failure stays retryable
as 503; a live claim answers 200 duplicate), then dispatches the
claimed snapshot with the live runner adapters (delivery parity with
the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
split hooks) keep being driven through their own hook. Capability
detection now credits claim_fire AND fire_claimed overrides, so
Chronos is correctly classified split-aware (its re-arm lives in
fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
unscoped reconcile would disarm other profiles' armed one-shots in the
shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
through every entry point, composing with upstream's manual-run
heartbeat (#76502) and background dispatch.
Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
Correlate approval requests, reject stale responses, replay pending approvals after reconnect or session resume, and preserve fail-closed timeout behavior.
Restores pass-through behavior for cron delivery and react/unreact that
was lost when d409f6748 routed them through resolve_send_target. Stored
cron job targets the channel directory doesn't recognize (e.g.
telegram:ops-room on a fresh install, photon group GUIDs) used to go to
the adapter verbatim; after d409f6748 they were silently dropped. Same
for react on platform-native ids.
Adds an opt-in pass_unresolved_references flag to resolve_send_target,
passed only by cron and react. Model-facing send tool stays strict.
Plugin platforms with a parser stay strict for all callers. The optional
validator still has the final say over passed-through ids.
Follow-up fixes on salvage:
- Update test_cron_relay_delivery_guards.py mock lambdas to accept **kw
(file added to main after PR branch point; lambdas didn't accept the
new keyword argument)
- Consolidate duplicated pass-through blocks into _pass_through_unresolved
local helper
Fixes#85128
Co-authored-by: Adolanium <Adolanium@users.noreply.github.com>
delegation.max_iterations is the per-subagent tool-call budget. The old
default of 50 truncated substantial delegated work: leaf agents spend
~15-20 turns on reconnaissance before producing output, then ran out of
budget mid-task and returned 'completed but unfinished' summaries. 250
gives real delegated work room to finish.
Changes:
- config_defaults.py: delegation.max_iterations 50 -> 250; _config_version 35 -> 36
- tools/delegate_tool.py: DEFAULT_MAX_ITERATIONS fallback 50 -> 250 (kept in
sync with the shipped default to prevent drift)
- config_migrations.py: _migrate_to_36 lifts configs still pinned at exactly
the OLD default 50 -> 250 on update, so existing installs inherit the new
headroom. Any other explicit value (deliberate override) is preserved;
unset inherits 250 at read time.
- cli-config.yaml.example: doc the new default
The cap is per-child and children run concurrently (max_concurrent_children
default 3), so this raises worst-case fan-out cost; delegation.child_timeout_seconds
(default 0 = off) remains available as a wall-clock guardrail, and users can
still pin a lower max_iterations explicitly.
Verified: migration lifts 50->250, preserves a deliberate 120, leaves unset
untouched (3/3); DEFAULT_CONFIG reads version=36, max_iterations=250, fallback=250.
The browser-use CLI runs under its own Python (uv tool / uvx), which
can differ from Hermes's venv interpreter. PYTHONPATH/PYTHONHOME
inherited from the agent process point at Hermes's venv
site-packages, and a child interpreter honors them ahead of its own —
so the CLI imported compiled C-extensions (pydantic_core) built for
the wrong interpreter and crashed with ABI mismatch /
ModuleNotFoundError (issues 83427, 84841, 86006, 86104; hits the
desktop backend on py3.14 and any shell exporting PYTHONPATH).
Strip both vars in _base_subprocess_env() — the CLI manages its own
environment and never needs Hermes's import path.
Salvaged from PR 83471 by Benjamin (@n1majne3), the earliest of two
independent fixes (also PR 84022 by @jklance16, PYTHONPATH-only);
regression test covers both vars and preserves unrelated env.
Lazy import inside _check_binary_document_write was unnecessary —
binary_extensions is a leaf module already imported at line 15.
Hoisted has_opaque_document_extension and is_pdf_path to the existing
module-level import. /simplify-code finding.
OPAQUE_DOCUMENT_EXTENSIONS was missing 10 extensions that read_file
auto-extracts via anydoc: .docm, .xlsm, .xlsb, .pptm, .ppsx, .ppsm,
.pps, .pot, .rtf, .epub. Each has the same corruption path: read_file
shows extracted text, model writes it back, container is destroyed.
Flagged by @egilewski on PR #82818 — proven live for .docm (text write
left a non-zip corpse). Added bytes-untouched regression test for .docm.
Port from nearai/ironclaw#7109: read_file auto-extracts .docx/.xlsx/.pptx
(and PDF via anydoc) to readable text, so a model plausibly believes it
holds the file's contents and writes the edited text back with
write_file/patch — silently destroying the document container. Proven
live on main: write_file over a valid .docx left a non-zip corpse, and a
text write over an existing .pdf clobbered the %PDF header.
- tools/binary_extensions.py: OPAQUE_DOCUMENT_EXTENSIONS +
has_opaque_document_extension() + is_pdf_path() (pure string checks)
- tools/file_tools.py: _check_binary_document_write() — opaque container
formats (doc/docx/xls/xlsx/ppt/pptx/odt/ods/odp) always rejected; .pdf
rejected only when overwriting an existing regular file (new-PDF
creation stays allowed, matching the upstream split guard). Wired into
write_file_tool and patch_tool (replace + V4A Update/Add headers;
Delete/Move skip the guard since they write no text).
- tests/tools/test_binary_document_write_guard.py: guard unit tests +
end-to-end write_file/patch coverage incl. bytes-untouched assertions.
Everything Browser Use is now managed by Hermes: the canonical binary
is the one install_cli() provisions into HERMES_HOME/bin, and every
resolution and provisioning site prefers it.
- _find_cli(): probe order flipped to managed bin -> PATH ->
user-level tool dir (then uvx across the same order). A user's own
uv tool install can no longer shadow the Hermes-managed copy with a
drifted version; side installs only matter when we have nothing.
- install_cli(): a browser-use on PATH no longer short-circuits the
install — only the managed copy does, so selecting any backend
provisions the copy Hermes controls and updates.
- _ensure_browser_use_cli() (hermes tools): drops its own PATH check
and always delegates to install_cli(), the single owner of the
managed-copy policy.
- install.sh / install.ps1: same short-circuit fix — only
HERMES_HOME/bin/browser-use counts as installed.
Follow-up to #86240 and #86320: with every non-Camofox backend
selection installing the CLI, managed-first closes the remaining
version-drift/shadowing class instead of guarding single sites.
Tests updated to pin the new contract: managed beats PATH and
user-local; PATH install does not satisfy install_cli; helper always
delegates. E2E-verified precedence chain with real files and a real
degraded-PATH install attempt.