Commit Graph

2233 Commits

Author SHA1 Message Date
Chen Jin ed0a8a480e fix(skills): resolve skills dir before relative_to so junction installs work (#86971) 2026-08-16 02:02:34 -07:00
Vivaan Dhawan 1596148ff2 fix(approval): deterministic approvals.single_query_mode for -q sessions
hermes chat -q sets HERMES_INTERACTIVE=1 (for interactive sudo prompts) but
runs one turn with no user waiting to answer approval prompts. Previously a
dangerous command triggered the interactive gate, waited the full 300s
timeout, then failed closed — and the agent was effectively forced to work
around the block, often silently auto-approving via execute_code (which
auto-approves in non-gateway mode).

Add approvals.single_query_mode (default deny, mirror of cron_mode):
  deny    — block dangerous commands and execute_code deterministically with
            a clear 'no user present' message (no 300s wait)
  approve — auto-approve dangerous commands/execute_code in -q mode

cli.py marks the session with HERMES_SINGLE_QUERY_SESSION; the shared gate
(_run_approval_gate, check_all_command_guards, check_execute_code_guard)
treats -q as a deterministic non-interactive context when that marker is set.
execute_code, the -q escape hatch, now honors single_query_mode instead of
auto-approving headlessly. Includes tirith parity in the combined guard and
docs. Fixes #86878.
2026-08-16 01:53:44 -07:00
PRATHAMESH75 d6f18cd7db fix(mcp): prefer server-native tool over generated utility on name collision (#87112)
An MCP server exposing a native tool named read_resource (or
list_resources/list_prompts/get_prompt) collided with the auto-generated
resource/prompt utility of the same name. The registration collision
handler flagged the pair as ambiguous and skipped BOTH entries, so the
server's own tool became silently unavailable on every gateway boot.

Resolve this specific native-vs-utility collision in favour of the native
tool: keep it and drop the shadowed utility, which is only convenience
sugar for servers that expose no such tool of their own. The conservative
skip-everything path still applies to genuinely ambiguous collisions (two
or more native tools normalizing to one name), which we cannot
disambiguate. Add a regression test covering the native-tool-wins path.

Fixes #87112
2026-08-16 01:52:26 -07:00
Teknium 8ad055414b fix(terminal): warn when exit_code 0 masks a piped build/test failure
`cargo build 2>&1 | tail -20` exits with tail's 0 even when the build
failed — bash without pipefail reports the last pipeline command's
status, and `cmd || echo failed` swallows the status the same way. The
model reads exit_code: 0 as a strong success signal and can conclude a
build passed while the visible output says it failed (community report,
Windows Rust builds; not platform-specific).

Two-part fix, mirroring OpenCode's prompt-side approach plus a
result-side backstop they don't have:

- Tool description now forbids piping builds/tests through
  tail/head/cat (output is already auto-truncated + spilled to a file)
  and warns that pipes/|| fallbacks mask exit codes.
- New annotate_masked_success() in tools/terminal_hints.py: when
  exit_code == 0, the command shape can mask an upstream status
  (top-level pipe into a passthrough consumer, or || echo/true), AND
  the output carries strong tool-specific failure shapes (rustc,
  cargo, pytest, gcc, npm, make, ninja), attach an advisory 'hint'
  telling the model to treat the run as failed and re-run bare.
  exit_code itself is never modified. Search/content heads
  (grep/rg/echo/printf/...) are excluded to avoid false positives on
  pipelines whose output legitimately contains error text.

E2E-verified through the real terminal tool path: hint fires on masked
cargo-style failures, silent on bare commands, clean pipes, and
grep/printf pipelines. 42 targeted tests pass.
2026-08-15 23:48:15 -07:00
5Hyeons 9975180a59 fix(mcp): handle DCR clients with secrets across all OAuth paths
Some MCP OAuth providers (notably Supabase) return a client_secret from
dynamic client registration but omit token_endpoint_auth_method. The MCP
SDK defaults the missing method to "none", so the token exchange omits
client_secret and the server rejects it (HTTP 422 "Required parameter:
client_secret"), looping the browser consent page.

This resolves the whole class, not just one provider:

- Storage layer (HermesTokenStorage): coerce secret-bearing client info
  with missing/none auth method to client_secret_post on both read and
  write, persisting the corrected shape.
- Both live provider paths (tools/mcp_oauth.py HermesOAuthClientProvider
  and tools/mcp_oauth_manager.py HermesMCPOAuthProvider): coerce
  in-memory client info immediately before token exchange and refresh.
- Accept the full 2xx range on token and refresh responses (Supabase
  returns 201 Created), instead of the SDK's exact-200 check.
- Redact token response bodies from error messages and logs on
  malformed responses.

The Figma-specific request-time default (apply_oauth_provider_defaults)
remains; this generalizes the same bug class for every DCR provider.

Fixes #29680. Supersedes #34274 and #35700 (201-only variants).
2026-08-15 23:37:15 -07:00
Teknium 9b9dd6eeab fix(computer-use): zero-rect AX bounds serialize as unknown, not a position
KDE/Qt apps report [0,0,0,0] bounds for elements that are perfectly
clickable by index (live QA: all 50 of kcalc's zero-rect elements,
including every radio button). Serializing that as a plausible rect
invites a model to derive coordinate=[0,0] and click the screen corner.

- _element_to_dict: zero rect -> bounds: null
- _format_elements: '@ bounds-unknown (click by element index)' instead
  of the fake rect in the summary line
- malformed bounds fail open (unchanged serialization)

Live-proven on real kcalc (cua-driver 0.20.0): 50 elements now null, 0
zero-rect leftovers, summary annotated, real rects preserved, and a
null-bounds radio button still clicks fine by index.
2026-08-15 22:36:52 -07:00
Teknium b7f6280259 fix(computer-use): refuse wrong-window input on app= mismatch; hint near-miss actions
Live complex-action QA on a real KDE desktop (kcalc + kate multi-app
flows) found two dispatch gaps:

1. Wrong-window input reported as success. Input actions deliver to the
   backend's sticky target (last capture/focus_app); the app= argument
   models routinely pass on the input call itself was silently dropped.
   Proven live: with kcalc sticky, type(text='777', app='kate') returned
   ok:true and typed 777 INTO KCALC. New guard: provable mismatch
   (both names known, neither substring of the other — list_windows
   names are localized/variant) refuses with input_target_mismatch and
   a one-call fix instruction. Unknown current target fails open so
   legacy no-app flows are untouched.

2. Near-miss unknown actions were dead ends. A model emitting 'hotkey'
   got a bare unknown-action error. Suggestion map now names the real
   action ('did you mean key?') without aliasing — we never repair bad
   model output, we just point at the schema.

Also documents the verified-lost-keystroke rung in the computer-use
skill: KTextEditor (Kate/KWrite) discards synthetic X keystrokes at the
toolkit level — foreground type reports ok but AX shows nothing arrived,
and a raw XTest control fails identically outside our stack. Guidance:
after one verified-lost round trip, switch to file/DBus I/O instead of
looping the ladder.

Live proof on the fixed build: mismatch refused, kcalc display clean,
same call after capture(app=kate) succeeds, 'hotkey' suggests 'key'.
11 new tests; 158 sibling tests green.
2026-08-15 22:18:09 -07:00
Teknium 460d345642 fix(computer-use): flag macOS zero-display capture in doctor + discovery reason
A headless Mac or asleep built-in panel leaves ScreenCaptureKit with 0
shareable displays while TCC grants pass — health_report stays ok and
every capture silently returns 0x0 (#67165). Guard at the report seam
(_apply_display_count_guard, both real and fallback paths): flips the
screen_capture_capability check to fail with recovery actions (wake
display / HDMI dummy / virtual display) and downgrades ok -> degraded.
The empty-discovery reason ladder gains the matching darwin rung.

Composed from #52949 (sujeet111) and #67259 (webtecnica); both PRs
predate the doctor rewrite and the envelope normalization on main, so
this reimplements their shared intent at the current seams.

Co-authored-by: Sujeet <64351924+sujeet111@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
2026-08-15 18:20:55 -07:00
0xGr1mm d303e18d09 fix(computer_use): ignore placeholder pid/window_id ids
capture() read any non-None pid/window_id as a request for exact-window
targeting. Several providers emit every declared schema property on every
tool call, zero-filling unused optional integers, so those calls arrive as
pid=0, window_id=0. The exact-target branch was then entered, the caller's
app= was discarded, _positive_int(0) returned None for both ids, and the
capture failed with a message pointing at pid/window_id. For that class of
model capture(app=...) and frontmost capture never worked at all.

Normalize non-positive ids to None before the branch decision so dispatch
falls through to app/frontmost discovery. Malformed non-numeric ids are
deliberately not treated as placeholders: they still reach the existing
validation error instead of being silently ignored.

Fixes #81333
2026-08-15 18:20:55 -07:00
Teknium 95aa709606 fix(computer-use): diagnose empty window discovery; fail fast on dead-daemon CLI fallback
Two live-QA findings from a locked KDE desktop (real cua-driver 0.20.0):

1. capture() with zero discovered windows returned a bare
   'capture mode=ax 0x0' — no hint that the desktop session was LOCKED,
   which freezes renderers and hides windows. New
   _empty_discovery_reason() names the dominant causes in order: locked
   session (loginctl LockedHint probe, fail-safe), missing DISPLAY,
   else a pointer at hermes computer-use doctor. Surfaced through the
   existing window_title -> summary path, so the model and the user see
   it inline.

2. _call_tool_via_cli retried 'daemon is not running' 4x with ~3.5s of
   backoff sleeps — a permanent condition for that invocation (the CLI
   transport needs the machine-wide daemon; Hermes' MCP runtime does
   not). Now fails fast on the first attempt with a message naming the
   split. Transient empty output (EAGAIN congestion) keeps the retry
   loop — pinned by test.

Live-verified on the locked desktop: capture now reports the lock and
the unlock action; CLI fallback errors immediately with the transport
explanation. 7 new tests; 191 sibling tests green.
2026-08-15 17:11:53 -07:00
Teknium 951ae62ffc test(computer-use): pin 0.17+ split refs/content_refs merge behavior
Regression test for the _ref_map merge (salvaged from #79515): the live
0.19.3 driver splits action refs into refs[] while content_refs re-lists
every node with empty actions; the empty entries must not clobber the
action-bearing ones. Caught live: every typed click refused with
browser_ref_stale until the merge fix.
2026-08-15 15:36:19 -07:00
Teknium 20cf326bd1 fix(computer-use): align browser authorization with live-verified cua-driver 0.19.3 contract
Live-tested against the real cua-driver 0.19.3 binary (Linux x86_64):

- bounded serve flags corrected: the daemon accepts
  --session-policy/--approve-session-policy, not the docs'
  --capability-manifest names (which it rejects). Verified end-to-end:
  a bounded daemon with a real policy file starts and reports running.
- browser-approve verified real but interactive-only (refuses without a
  TTY) and its token is a legacy compatibility path disabled by default
  on current drivers (per the live browser_prepare schema). Kept as a
  passthrough; no longer presented as the primary route.
- NEW primary standard-mode route, verified live: launch the runtime
  with cua-driver's trusted-launcher grant. config opt-in
  computer_use.grant_existing_profile: true appends
  --grant existing-profile to the standard-mode MCP spawn (MCP
  initialize verified accepting the flag). Default false = attachment
  keeps failing closed. Never applied to bounded/unrestricted daemons.
- Skill, system prompt, tool schema, and docs updated to the verified
  ladder: config grant > bounded manifest > YOLO; token = legacy.
2026-08-15 15:04:32 -07:00
Teknium 48dd9c87cf feat(computer-use): user-facing authorization for cua-driver browser attachment
Completes the typed cua_browser_* route (PR #74166 lineage) with the
authorization surface that makes existing-profile attachment and
repeatable bounded automation reachable by real users:

- hermes computer-use browser-approve: CLI passthrough that mints
  cua-driver's five-minute single-use attachment token for one exact
  (pid, window_id). The user, never the model, is the token source.
- approval_token passthrough on cua_browser_prepare (schema + dispatch +
  browser_route), forwarded only for existing_profile and only as a
  non-empty string.
- computer_use.permission_mode: bounded + capability_manifest config:
  private per-session embedded daemon launched with
  --capability-manifest/--approve-capability-manifest; missing manifest
  fails loudly. 'unrestricted' is deliberately NOT a config value —
  it stays bound to the explicit per-session YOLO toggle.
- Skill + system-prompt + docs guidance for the three authorization
  rungs and the isolated-profile-first default.

E2E-verified against a temp HERMES_HOME: real config resolution to
bounded, loud failure without a manifest, real argparse path driving a
fake cua-driver binary, standard default preserved.
2026-08-15 15:04:32 -07:00
Teknium c9a806e9d2 fix(browser): pin named sessions to their own tab on shared browsers
Follow-up to #86916. That fix gave named sessions their own daemon
(socket/log/pid) and their own provider browser — but on a SHARED local
Chrome / CDP browser, a fresh named daemon still attaches to the first
existing page, the same page a sibling daemon may hold. A named session
that never calls new_tab() could still stomp another's tab.

browser_exec now prepends a small preamble to the model's code for named
sessions on shared browsers: once per daemon process (marker keyed by
uid + BU_NAME + daemon pid), it creates a fresh tab via
Target.createTarget and switch_tab()s onto it before any model code
runs. Private per-name browsers (provider-keyed bu-named-<name>, or
direct-API Browser Use cloud) skip the preamble via an internal env
sentinel popped before launch — there's nobody to collide with, and the
extra tab would leak.

Best-effort by design: if the preamble's CDP calls fail, behavior
degrades to pre-fix, never blocks the exec.

E2E against a shared headless Chrome with the STOCK harness: two named
sessions issuing bare js() writes (no new_tab) kept distinct state
(EDGE-A/EDGE-B read back intact); the sabotage run without the preamble
reproduced the clobber (both read EDGE-B). Removes the dependency on the
upstream browser-harness tab-isolation PR for correctness.
2026-08-15 04:23:26 -07:00
Teknium bb4f680f22 fix(browser): named browser_exec sessions compose with every backend
session=<name> previously set BU_NAME and then skipped backend resolution
entirely — the parameter was documented as cloud-only, so all local/CDP
work funneled through the single default daemon and one IPC socket, and
concurrent sessions (parallel subagents, simultaneous chats) clobbered
each other's browser connection. Reported by @shantanugoel on X.

Now a named session composes with whatever browser source is configured:

- BU_NAME still namespaces the harness daemon (per-name IPC socket, log,
  pid — upstream already isolates these), for local Chrome and CDP.
- The /browser connect CDP override is now exported for named sessions
  too; previously a named daemon ignored it and fell back to scanning
  local Chrome profiles.
- On provider backends (Browserbase, Firecrawl, Nous gateway), the name
  keys its own provider browser via the shared _get_session_info cache
  (bu-named-<name>), so each name gets its own cloud browser, the same
  name reuses one across calls and tasks, and unnamed calls keep the
  per-task key.
- Direct-API Browser Use cloud configs keep the native named-daemon path
  (provider resolution would double-session and double-bill).

Tool schema/description updated so models reach for session=<name> for
parallel work on any backend, not just cloud.

E2E: two named sessions against a real headless Chrome (real browser-use
CLI, BU_CDP_URL) ran concurrently, set distinct page state, and read it
back intact; sabotage run confirms the new tests fail without the fix.
2026-08-15 04:05:57 -07:00
Teknium e729055a56 test(delegate): regression for #86632 — cron sync fallback must return after child completes
End-to-end #86632 reproduction: a real AIAgent child (mocked LLM) with the
post-turn skill-review trigger armed, dispatched through
delegate_task(background=True) on a session runtime where async delivery is
unsupported and no origin session id is bound (cron, post-#66617) — forcing
the synchronous fallback. Asserts (1) delegate_task returns the child's
result, and (2) the automatic background-review fork never spawns inside the
delegated child (the wedge site: the fork replayed the conversation on the
child's finalize path, and _child_future.result(timeout=None) never returned;
heartbeat went stale after 15 idle cycles and the cron watchdog killed the
job).

Verified RED on pre-fix main (fork spawns and wedges), GREEN with the
_delegate_depth guard in AIAgent._spawn_background_review.

Fixes #86632
2026-08-15 02:44:38 -07:00
blunkjamie-dev 4415f917b4 test(session-search): lock positional parameter prefix
(cherry picked from commit 73592200c69a4f0b6d7c290ce45832847df608e2)
2026-08-15 00:34:29 -07:00
blunkjamie-dev 6e1bdc0a18 perf(session-search): adapt discovery result hydration
(cherry picked from commit 60a3530444f65c2cdcd4e5b983e4aa380ed651c7)
2026-08-15 00:34:29 -07:00
andy c99a45b28e fix(browser): reap leaked agent-browser daemons whose owner is still alive
The orphan reaper had two gaps that let agent-browser daemons accumulate
indefinitely inside a single long-lived hermes process:

1. `_reap_orphaned_browser_sessions()` ran exactly once, before the cleanup
   loop started, so a leak appearing after boot could never be recovered.

2. `owner_alive is True` skipped unconditionally. In-memory session tracking
   is lost on any exception path between spawn and registration, but the
   owner PID stays up — so such a daemon was skipped forever.

The daemon-side `AGENT_BROWSER_IDLE_TIMEOUT_MS` is not a backstop for (2):
it does not fire when the daemon itself is wedged, e.g. after Chrome's
framework was replaced underneath it by an auto-update.

Observed on macOS: five agent-browser daemons (96 Chrome processes) built up
over 10 days inside an 18-day-uptime hermes process, holding roughly 5 CPU
cores busy and driving the load average past 100. Four of those processes
were still running a Chrome framework version that had since been replaced
on disk, spinning at ~85% CPU each.

Changes:

- Re-run the reaper every `BROWSER_ORPHAN_REAP_INTERVAL` (300s) from inside
  the cleanup loop. Cycle 0 preserves the existing startup reap.

- When the owner is alive but the session is untracked, fall back to idle
  age: reap past `BROWSER_ORPHAN_GRACE_SECONDS`, defined as
  `max(1h, 20 x inactivity_timeout)`. Unknown age fails safe.

- Add `_socket_dir_idle_seconds()` — the newest mtime under a session's
  socket dir. Every browser command writes `_stdout_<cmd>` / `_stderr_<cmd>`
  there, making it a last-activity marker that survives hermes restarts and
  does not depend on in-memory bookkeeping surviving an exception path. It
  scans directory entries rather than reading the directory mtime alone:
  command names repeat, and rewriting an existing `_stdout_click` updates
  that file's mtime but not the directory's, so a dir-mtime-only check would
  report a busy session as idle and reap it.

Sessions still present in `_active_sessions` are never touched at any age,
and the new path still goes through `_verify_reapable_browser_daemon`, so
the anti-spoof / anti-PID-recycle guarantees from #14073 are unchanged.

Adds 9 tests: idle-age unit tests (including the dir-mtime regression),
spared/reaped/fail-safe cases for a live owner, the identity-guard gate on
the new path, and a periodic-reap test asserting more than one reap per
cleanup-thread lifetime.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 00:33:01 -07:00
adikpb 688abc585f test(vision): assert no max_tokens cap in browser and video aux kwargs
Sweeper follow-up: the browser-screenshot and video kwargs captures now
also assert max_tokens is absent, protecting the central auxiliary
no-cap policy against refactors that would restore the hardcoded caps.
2026-08-15 12:53:37 +05:30
adikpb ec470d9db2 test(vision): assert vision aux calls carry no max_tokens cap
Covers the max-tokens-knob contract: vision call_kwargs omit max_tokens
entirely (configured values, defaults, and even an explicit
auxiliary.vision.max_tokens config entry must never be forwarded), so
providers use their full output budget.
2026-08-15 12:53:37 +05:30
Teknium 62014d8dd3 fix(guard): steer live-checkout block message to disk-backed scratch clones
The guard's "use a separate worktree or temporary clone" advice sent
agents to /tmp by default. /tmp is RAM-backed tmpfs on most distros, and
parallel salvage clones each running npm ci (~1.6GB per clone) filled a
32GB tmpfs to 97% during a 15-subagent campaign, ENOSPC-ing sibling test
runs. The message now recommends `git clone --shared <root> ~/.hermes/scratch/<task>`
(honoring HERMES_HOME), warns that dependency installs belong on real
disk, and tells the agent to delete the clone once the branch is pushed.
2026-08-14 22:34:51 -07:00
Tachi b67f202184 fix(update): honor lazy install opt-out during restore 2026-08-14 22:33:44 -07:00
Tachi 979a20052f fix(update): preserve activated extras across runtime rebuilds 2026-08-14 22:33:44 -07:00
kshitij ce658e82ff fix(session-search): narrow lineage escape to reset/compression; trust SQL child classifier
Follow-up on the salvaged #85764 commits, addressing review findings:

- _session_left_live_context now allowlists end_reason == 'compression'
  or a fresh reset (_FRESH_RESET_END_REASONS) instead of accepting any
  non-None end_reason. The wide predicate let 'branched' parents — whose
  transcript /branch verbatim-copies into the child — surface as
  same-lineage recall hits, returning content already in the caller's
  live context (verified empirically vs main).
- _FRESH_RESET_END_REASONS is now derived from the canonical
  hermes_state_common._RESET_END_REASONS (plus CLI 'new_session') instead
  of a third hand-maintained copy, per that tuple's anti-drift comment.
  Import verified cycle-free.
- Browse drops the Python re-check of parent_session_id rows:
  list_sessions_rich (include_children=False) already applies the
  canonical _LISTABLE_CHILD_SQL classifier, and the Python re-check
  re-hid legacy pre-marker reset children the SQL same-key heuristic
  deliberately admits. _has_reset_from_marker (now orphaned) removed.
- Tests: branched-parent exclusion regression guard (mutation-checked:
  fails on the overbroad predicate) + legacy pre-marker reset child
  browse guard. 48/48 pass.
2026-08-15 10:25:19 +05:30
joaomarcos 39e480c051 fix(state): close leaked SessionDB connections on exception paths (#83226)
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).

- Close partially initialized SessionDB connections on every constructor
  exception path via a finally block guarded by an initialization-complete
  flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
  CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
  thread, close on worker exit, reject new enqueues after shutdown starts,
  and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
  API disconnect failures, shutdown recovery, RetainDB late enqueue, and
  foreign-loop async clients.

Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 21:41:26 -07:00
kshitij 70704b962f fix(delegate): child's dedicated SessionDB must follow the parent's db_path
A bare SessionDB() resolves the launch profile's default state.db, but
parents can hold non-default per-profile handles (tui_gateway opens
SessionDB(db_path=<profile_home>/state.db) for non-launch profiles and
hands them to agents via _transfer_db_to_agent). A child of such a
parent would write its transcript into the WRONG database — cross-
profile leakage that breaks parent_session_id lineage and
session_search. Open the dedicated handle at the parent handle's
db_path instead (AsyncSessionDB forwards .db_path via __getattr__, so
the gateway wrapper path works too). Regression test verified RED on
the pre-fix code.
2026-08-14 21:39:31 -07:00
thatssoheil eef43e3160 fix(delegate): close the dedicated SessionDB if child construction fails; test degradation
Review follow-up (cc3f18197): if AIAgent() raises inside _build_child_agent
the freshly-opened dedicated handle has no owner and no child close() will
ever run — release it on the exception path so the sqlite fds don't
outlive the failed spawn. Also pin the degradation contract with a test:
a parent without a SessionDB still yields session_db=None children.
2026-08-14 21:39:31 -07:00
thatssoheil 65e005d0e7 fix(delegate): subagents get a dedicated SessionDB, not the parent's (#81267)
Cron run_job closes its per-job SessionDB in its finally block while a
fire-and-forget background delegation subagent is still flushing on a
daemon thread. The child shared the parent's SessionDB object, so every
subsequent flush hit the closed handle ('NoneType' object has no
attribute 'execute') and the child's whole transcript was silently
dropped. The same teardown-while-child-alive shape exists on gateway
session end and /new mid-delegation.

Each child now opens its own SessionDB connection (owned flag set at
construction so child.close() releases it), so no parent teardown can
close the child's handle out from under it.

Regression test proves the child gets a distinct live handle that
survives the parent's close().
2026-08-14 21:39:31 -07:00
deacon-botdoctor 9bff109783 fix(gateway): cancel native clarify only on free prose
Teknium review on #75732: releasing the pending clarify whenever
resolve_text_response_for_session returned False also cancelled
retryable multi-select invalid selections (out-of-range numbers,
unrecognised comma-lists).

Classify rejected typed replies in clarify_gateway:
- rejected_prose → cancel clarify, fall through busy routing (deadlock break)
- rejected_selection → keep clarify armed so the user can retry

Add native multi-select gateway regressions for both paths.
2026-08-14 21:24:36 -07:00
deacon-botdoctor 28b62b069a fix(gateway): keep first clarify resolution 2026-08-14 21:24:36 -07:00
HexLab98 1a06e70e1e test(session-search): cover /new-reset discovery, scroll, and browse
Regress the #85756 chain: a session_reset parent must surface from the
empty child, title/scroll must follow, browse must list that parent, and
live delegation children must stay hidden.
2026-08-14 21:22:56 -07:00
Teknium dc2fe99ecf feat(delegation): mark max_iterations-truncated subagent results for the parent (#86641)
A delegated subagent that exhausts its per-child iteration budget
(delegation.max_iterations) still returns a summary, so the result carries
status='completed' even though the child's exit_reason is 'max_iterations' and
its work was cut off mid-task. The parent then reads 'completed', trusts the
partial summary, and only discovers the truncation by parsing the prose (where
the child happens to mention 'hit the iteration limit'). That wastes parent
turns and risks acting on incomplete work.

exit_reason is already computed authoritatively and threaded to every
parent-visible surface; it just wasn't reflected anywhere the parent reads at a
glance. This surfaces it:

- delegate_tool.py: add a parent-visible boolean 'truncated' (= exit_reason ==
  'max_iterations') to each task entry, alongside the existing exit_reason.
- process_registry._format_async_delegation: for both the batch and single-task
  paths, when truncated -> use a warning icon, append
  'TRUNCATED: hit max_iterations — work may be incomplete' to the header/Status
  line, and prefix the summary with an unmissable truncation notice. status
  semantics are left unchanged (stays 'completed') so existing icon/summary
  branch logic and ~10 tests asserting status=='completed' stay valid.

Tests: single-task truncated -> banner; single-task clean -> no banner; batch
marks only the truncated task, not its clean sibling. 23/23 in the async-
delegation suite.

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-14 21:19:15 -07:00
Teknium eeb2d23b86 fix(memory): accept new_text as an alias for content (#86642)
memory replace/remove target an entry by old_text but supply the replacement
via content — an asymmetric pairing. Callers naturally reach for new_text to
mirror old_text (it's exactly the patch tool's old_string/new_string shape),
which left content empty and errored 'content is required'. The failure also
rendered tersely, making the cause easy to miss and costing a retry.

Accept new_text as an alias for content on both shapes:
- single-op: coalesce content = content or new_text in memory_tool() + a
  new_text param wired through the registry handler.
- batch ops: content = op.content or op.new_text (and in the approval-gate
  preview builder).
- schema: document the alias on content and add a new_text property to the
  single-op params and batch item props so strict validators accept it.
content wins if both are set.

Tests: new_text alias on single add/replace, batch add/replace, and
content-wins-over-new_text. 43/43 across memory tool + schema suites.

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-14 21:15:08 -07:00
David Metcalfe ebb2859132 fix(clarify): preserve resolved answers when clear_session races a button response
Session-boundary cleanup (gateway/run.py: run-finalization and prompt
delivery-failure paths) calls clear_session to cancel pending clarifies.
It unconditionally overwrote every entry's response with the empty
cancellation sentinel, even entries already resolved by a button callback
or text intercept. A waiter that had already observed the resolved event
would then return the empty sentinel instead of the real answer, silently
discarding the user's response on /new, gateway shutdown, or cached-agent
eviction.

First-writer-wins: clear_session now cancels only entries whose event is
not yet set; already-resolved entries keep their response. Mirrors the
guard resolve_gateway_clarify gained for the same contract. Reported by
doryani-ai on PR #75732; regression test covers the button-then-cleanup
interleaving.
2026-08-14 20:58:42 -07:00
Evgenii 1a8625abee fix(cron): harden gateway fire admission and provider compatibility
- The gateway api_server fire webhook acknowledges 202 only after a
  durable claim + execution row exist (admission failure stays retryable
  as 503; a live claim answers 200 duplicate), then dispatches the
  claimed snapshot with the live runner adapters (delivery parity with
  the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
  split hooks) keep being driven through their own hook. Capability
  detection now credits claim_fire AND fire_claimed overrides, so
  Chronos is correctly classified split-aware (its re-arm lives in
  fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
  unscoped reconcile would disarm other profiles' armed one-shots in the
  shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
  through every entry point, composing with upstream's manual-run
  heartbeat (#76502) and background dispatch.

Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
2026-08-14 20:46:50 -07:00
Teknium f8cc6d082e test(todo): make JSON-string coercion test order-agnostic
The type-coercion test pinned index order of todos, which #42649's
_normalize_order intentionally changes (in_progress lifts ahead of
earlier pending rows). Assert coercion by id instead of position.
2026-08-14 20:24:42 -07:00
LeonSGP43 0e7151ceae fix(todo): keep active step ahead of pending rows 2026-08-14 20:24:42 -07:00
VooDoo Pixels f703e70618 fix: make desktop approval routing reliable
Correlate approval requests, reject stale responses, replay pending approvals after reconnect or session resume, and preserve fail-closed timeout behavior.
2026-08-14 20:24:32 -07:00
Adolanium ed4f50de51 fix(send_message): hand unresolved cron and react targets to the adapter again
Restores pass-through behavior for cron delivery and react/unreact that
was lost when d409f6748 routed them through resolve_send_target. Stored
cron job targets the channel directory doesn't recognize (e.g.
telegram:ops-room on a fresh install, photon group GUIDs) used to go to
the adapter verbatim; after d409f6748 they were silently dropped. Same
for react on platform-native ids.

Adds an opt-in pass_unresolved_references flag to resolve_send_target,
passed only by cron and react. Model-facing send tool stays strict.
Plugin platforms with a parser stay strict for all callers. The optional
validator still has the final say over passed-through ids.

Follow-up fixes on salvage:
- Update test_cron_relay_delivery_guards.py mock lambdas to accept **kw
  (file added to main after PR branch point; lambdas didn't accept the
  new keyword argument)
- Consolidate duplicated pass-through blocks into _pass_through_unresolved
  local helper

Fixes #85128

Co-authored-by: Adolanium <Adolanium@users.noreply.github.com>
2026-08-15 07:07:40 +05:30
Benjamin 47fd2eb7c8 fix(browser): strip PYTHONPATH/PYTHONHOME from browser-use CLI subprocess env
The browser-use CLI runs under its own Python (uv tool / uvx), which
can differ from Hermes's venv interpreter. PYTHONPATH/PYTHONHOME
inherited from the agent process point at Hermes's venv
site-packages, and a child interpreter honors them ahead of its own —
so the CLI imported compiled C-extensions (pydantic_core) built for
the wrong interpreter and crashed with ABI mismatch /
ModuleNotFoundError (issues 83427, 84841, 86006, 86104; hits the
desktop backend on py3.14 and any shell exporting PYTHONPATH).

Strip both vars in _base_subprocess_env() — the CLI manages its own
environment and never needs Hermes's import path.

Salvaged from PR 83471 by Benjamin (@n1majne3), the earliest of two
independent fixes (also PR 84022 by @jklance16, PYTHONPATH-only);
regression test covers both vars and preserves unrelated env.
2026-08-14 16:22:38 -07:00
kshitij 5c988e2461 fix: cover remaining anydoc-extracted container formats
OPAQUE_DOCUMENT_EXTENSIONS was missing 10 extensions that read_file
auto-extracts via anydoc: .docm, .xlsm, .xlsb, .pptm, .ppsx, .ppsm,
.pps, .pot, .rtf, .epub. Each has the same corruption path: read_file
shows extracted text, model writes it back, container is destroyed.

Flagged by @egilewski on PR #82818 — proven live for .docm (text write
left a non-zip corpse). Added bytes-untouched regression test for .docm.
2026-08-15 02:51:59 +05:30
Teknium 6d51c831eb fix(file_tools): refuse plain-text writes that corrupt binary documents
Port from nearai/ironclaw#7109: read_file auto-extracts .docx/.xlsx/.pptx
(and PDF via anydoc) to readable text, so a model plausibly believes it
holds the file's contents and writes the edited text back with
write_file/patch — silently destroying the document container. Proven
live on main: write_file over a valid .docx left a non-zip corpse, and a
text write over an existing .pdf clobbered the %PDF header.

- tools/binary_extensions.py: OPAQUE_DOCUMENT_EXTENSIONS +
  has_opaque_document_extension() + is_pdf_path() (pure string checks)
- tools/file_tools.py: _check_binary_document_write() — opaque container
  formats (doc/docx/xls/xlsx/ppt/pptx/odt/ods/odp) always rejected; .pdf
  rejected only when overwriting an existing regular file (new-PDF
  creation stays allowed, matching the upstream split guard). Wired into
  write_file_tool and patch_tool (replace + V4A Update/Add headers;
  Delete/Move skip the guard since they write no text).
- tests/tools/test_binary_document_write_guard.py: guard unit tests +
  end-to-end write_file/patch coverage incl. bytes-untouched assertions.
2026-08-15 02:51:59 +05:30
Teknium 20e5d51bea fix(browser): managed-first browser-use CLI resolution
Everything Browser Use is now managed by Hermes: the canonical binary
is the one install_cli() provisions into HERMES_HOME/bin, and every
resolution and provisioning site prefers it.

- _find_cli(): probe order flipped to managed bin -> PATH ->
  user-level tool dir (then uvx across the same order). A user's own
  uv tool install can no longer shadow the Hermes-managed copy with a
  drifted version; side installs only matter when we have nothing.
- install_cli(): a browser-use on PATH no longer short-circuits the
  install — only the managed copy does, so selecting any backend
  provisions the copy Hermes controls and updates.
- _ensure_browser_use_cli() (hermes tools): drops its own PATH check
  and always delegates to install_cli(), the single owner of the
  managed-copy policy.
- install.sh / install.ps1: same short-circuit fix — only
  HERMES_HOME/bin/browser-use counts as installed.

Follow-up to #86240 and #86320: with every non-Camofox backend
selection installing the CLI, managed-first closes the remaining
version-drift/shadowing class instead of guarding single sites.

Tests updated to pin the new contract: managed beats PATH and
user-local; PATH install does not satisfy install_cli; helper always
delegates. E2E-verified precedence chain with real files and a real
degraded-PATH install attempt.
2026-08-14 13:51:07 -07:00
kimyxx d29abb7e6b fix(browser): discover browser-use from user-level tool directories
Desktop/TUI workers can spawn with a minimal PATH that omits
~/.local/bin, the default location where uv tool install links the
browser-use binary. _find_cli() then failed to resolve an installed
CLI and Browser Use mode silently fell back to the built-in tools.

Probe the user-level tool dir (~/.local/bin on POSIX, APPDATA/uv/bin
on Windows) between PATH and the managed HERMES_HOME/bin, for both
the browser-use binary and the uvx fallback.

Salvaged from PR #83788 by @kimyxx onto current main; tests adapted
and extended with precedence and uvx coverage.
2026-08-14 13:32:03 -07:00
kshitij 1b1975781f test: use tmp_path fixture for import test HERMES_HOME
The subprocess import test was creating .tmp-hermes-exec-ask-import/
in the repo root without cleanup. Switch to pytest's tmp_path fixture
so the temp directory is auto-cleaned and never appears as untracked.
2026-08-14 17:48:28 +05:30
xxxigm d6b4083f41 test(approval): cover CLI EXEC_ASK leak and fix slash-worker Path mock
Regression tests for silent pending_approval when ask-mode leaks into
interactive CLI, plus a Path-typed hermes_constants mock so the
slash-worker profile_home test survives per-file isolation.
2026-08-14 17:48:28 +05:30
dsad 668396c2e1 test(security): add tests for notification-path redaction
Add TestNotificationRedaction class with two tests:

1. test_completion_notification_redacts_secret — verifies _move_to_finished
   redacts API keys in completion notifications before enqueueing
2. test_watch_match_notification_redacts_secret — verifies _check_watch_patterns
   redacts secrets in watch_match notifications before enqueueing

These tests cover the gap identified in #43025 where the explicit process
tool path (poll/log/wait) was redacted but the automatic notification
delivery path was not.
2026-08-14 01:08:13 -07:00
spfcraze d2d7766750 fix(tools,gateway): format watch_overflow events instead of dropping them
format_process_notification had no case for watch_overflow_tripped /
watch_overflow_released, so a watch-pattern notification flood surfaced
as '[IMPORTANT: Background process  exited (exit code ?)]' — a phantom
exit notification for a process that never existed — while the actual
'watch flood, N notifications suppressed' summary in the event's
message field was silently dropped. The gateway delivery path was
worse: _drain_gateway_watch_events retained only watch_match and
watch_disabled, discarding overflow events entirely before formatting.
Route both event types through the message field in the shared
formatter and the gateway formatter, and retain them in the gateway
drain.
2026-08-14 01:08:13 -07:00
Teknium 2a6eef77b7 chore: enforce LF line endings for all source files via .gitattributes (#85866)
Windows contributors' tools default to CRLF; without repo-level
normalization an edit becomes a whole-file phantom diff (583-line
'change' observed today from one 2-line edit), string-match patch
tooling breaks on invisible \r, and review is polluted. Extend the
existing LF rules (shell/Docker) to every source/text extension:
normalize at check-in AND check out as LF so working trees match the
index on all platforms. *.ps1 stays CRLF (PowerShell 5.1 tooling).

git add --renormalize: exactly one tracked file had mixed endings
(tests/tools/test_windows_agent_loop_papercuts.py) — normalized here,
so no phantom diffs land on anyone's next commit.
2026-08-13 23:36:25 -07:00