Commit Graph

4742 Commits

Author SHA1 Message Date
teknium1 30b22b54ae fix(tools): bound execute_code's lifecycle probe and keep the terminal guard answerable to /stop
execute_code ran the same unbounded _is_supervised_gateway_process() probe
ahead of every cell, so the wedge #111922 bounds in terminal_tool still hung
an execute_code call (and its cron slot) forever: share the cell's deadline
and fail closed with a retryable error when the probe renders no verdict.

Moving the terminal pre-exec guard onto a deadline worker made it blind to
/stop, which keys on the tool thread's ident: record the acting-for tid in a
contextvar (copied into the worker by run_bounded_sync) so is_interrupted()
on the worker honours the tool thread's bit too.

Floor the guard's share of the deadline at 30s so a short command timeout
does not turn the guard's own cold-start cost (imports, git probes under
load) into a refusal — tests/tools/test_terminal_error_redaction.py was red
on the branch for exactly that.
2026-09-15 19:09:29 -07:00
teknium1 c832920275 fix(tools): pre-exec guard that misses the deadline refuses the command
The salvaged commit put `_pre_exec_block` behind the command's
`run_bounded_sync` deadline but let a timed-out guard fall through into
execution. The gateway-lifecycle, dangerous-workdir and self-repo checks
apply unconditionally (`force=True` cannot bypass them), so a guard that
never rendered a verdict must not let the command run unguarded: return
the terminal error envelope (`status: error`, "did not finish ... Retry
the call") instead, mirroring how the bounded `env.execute` path reports
its own expiry as a result rather than continuing.

Tests trimmed to the two invariants: a wedged guard returns a bounded
error without executing; a completed guard keeps its verdict (pass ->
execution, rejection -> its own blocked result).
2026-09-15 19:09:29 -07:00
KoNit-K c1e749d679 fix(tools): bound terminal pre-exec guards 2026-09-15 19:09:29 -07:00
John Paul Soliva a17d0409be fix(mcp): report lazily registered servers as lazy, not configured or failed
A `lazy: true` MCP server registers its tools from the schema cache and
spawns on first use. Three consumers still equated "alive" with a live
session, so a healthy all-lazy startup was reported as a total failure:

- `get_mcp_status()` fell through to `status: configured, tools: 0` for a
  lazily registered server. It now reports `lazy` with the cached tool
  count (`connected: False`); an in-flight or failed first-use connect
  still outranks it because the error is the actionable part.
- `discover_mcp_tools()`'s summary counted every name absent from
  `_servers` as failed, logging `MCP: 0 tool(s) from 0 server(s) (2
  failed)` right after registering every cached tool, and re-announced
  the same "failure" on every repeat discovery. Lazy servers are now
  reported as `(N lazy, not spawned yet)` and an already-lazy server is
  not re-announced.
- `hermes_cli/mcp_startup.py` judged a discovery run by `connected` at
  two sites, so every startup logged `Background MCP discovery completed
  with zero connected servers` and every later call re-spawned the
  discovery thread as a retry. One predicate,
  `_discovery_registered_servers`, treats a lazy registration as a
  usable outcome at both sites.
- `hermes_cli/banner.py` rendered the unknown `lazy` status through the
  red "could not connect" line; it now shows the cached tool count with
  `(lazy, starts on first use)`.

Ported from #100648 (core hunks only; the toolsets-filter predicate
branch, the Ink TUI component extraction and 13 tests were not ported).

Fixes #111717
2026-09-15 19:06:54 -07:00
KoNit-K c001881d85 fix(mcp): route /v1/runs MCP trust-gate consent through the run's approval callback
A write-capable tool on a `trust: untrusted` MCP server was denied instantly
from POST /v1/runs: request_elicitation_consent only took the gateway path
when _is_gateway_approval_context() was true, and api_server sits in
_UNATTENDED_APPROVAL_PLATFORMS (webhook-style sessions have nobody to
answer). A live /v1/runs run is the exception: it registers a gateway notify
callback and answers via approval.request -> POST /v1/runs/{id}/approval —
the same bridge 04fcf9159 keeps alive for the dangerous-command gate. Treat
an api_server session that is neither cron nor single-query as
callback-backed; a run without a registered callback still fails closed.

Salvaged from #111529 with the redundant single-query re-gate on the
generic gateway branch dropped (no real surface binds a chat platform,
HERMES_SINGLE_QUERY_SESSION and an in-process callback together).

Part of #111526
2026-09-15 19:06:27 -07:00
teknium1 341f8b4d93 fix: cap, loopback-bypass and share the MCP proxy mounts
Review follow-up on the MCP HTTP proxy PR:
- Proxy mounts win over transport= for matching URLs, so a bare
  AsyncHTTPTransport mount bypassed the 10 MiB wire-body cap whenever a
  proxy applied. Each mount is now wrapped in _make_mcp_body_cap_transport.
- Loopback MCP servers (127.0.0.1 / ::1 / localhost) were dialed through
  HTTP_PROXY unless NO_PROXY covered them; _mcp_proxy_mounts now returns
  None for is_loopback_host (agent.proxy_bypass rule).
- Dropped the fail-open try/except around the proxy transport construction;
  a proxy httpx cannot build surfaces as the server's connect error.
- The content-type preflight client now takes an explicit transport plus
  the same proxy mounts as the SDK client, so probe and handshake take the
  same route (no httpx env auto-detection divergence).
2026-09-15 19:05:57 -07:00
teknium1 ee1bfef857 fix(mcp): NO_PROXY for MCP servers uses the repo matcher; trim tests to two invariants
Follow-up to the salvaged #111796 commit:

- NO_PROXY matching goes through `agent.proxy_bypass.should_bypass_proxy` (the one
  matcher the LLM transport and the gateway adapters already use), so CIDR ranges and
  `*.host` patterns bypass the proxy for MCP servers exactly as they do for the model
  endpoint. The stdlib `proxy_bypass` stays for the OS bypass list (Windows
  ProxyOverride / macOS exceptions). Live probe: NO_PROXY=10.255.255.0/24 still routed
  the MCP request through the proxy before this commit, direct after.
- Drop the try/except around `getproxies()` / `proxy_bypass()`: the stdlib guards its
  own registry/sysconf reads and httpx calls the same functions unguarded.
- Trim the six contributor tests to two invariants (mount + NO_PROXY incl. CIDR; both
  client builders carry mounts next to the body-cap transport). Fixture uses the stdlib
  `getproxies_environment` / `proxy_bypass_environment` instead of a hand-rolled copy and
  skips when the mcp SDK is absent.
- Docs: one sentence on the MCP page about proxy resolution for HTTP/SSE servers.
- contributors/emails mapping for the PR author.
2026-09-15 19:05:57 -07:00
VictorTran1023 ceb1aa19a0 fix(mcp): restore proxy support for HTTP/SSE MCP servers
httpx auto-detects proxies only when ``transport is None``
(``allow_env_proxies = trust_env and transport is None``). The wire-body cap hands
every MCP HTTP/SSE client a custom transport, so HTTP_PROXY / HTTPS_PROXY and the
Windows-registry / macOS system proxy were silently ignored: on a network that
reaches the MCP host only through a proxy, every connect failed with
"All connection attempts failed" and the server was parked (tools never appeared).

Rebuild httpx's own proxy resolution as explicit ``mounts`` — environment first,
then the OS proxy, NO_PROXY / platform bypass honoured, socks:// normalized, and
TLS settings identical to the transport they accompany.
2026-09-15 19:05:57 -07:00
teknium1 8e11666726 fix(mcp): resolve managed Windows Node launchers (npx.cmd/npm.cmd) for stdio MCP servers
On Windows a stdio MCP server configured with `command: npx|npm|node` failed
with WinError 2 whenever the desktop/gateway PATH lacked the managed Node dir:
`_node_fallback` probed only the POSIX shape `<HERMES_HOME>/node/bin/<cmd>`
with no extension, while `scripts/install.ps1` unpacks Node directly into
`<HERMES_HOME>\node` as `npx.cmd`/`npm.cmd`/`node.exe`. It also derived the
home from raw `os.getenv("HERMES_HOME")`, so a context-local profile home
(multiplexed gateway) was ignored.

Reuse the platform-aware helpers instead of a second hand-rolled layout:
`hermes_constants.iter_hermes_node_dirs(get_hermes_home())` supplies both
managed shapes in the right order, and the module's own `_npx_bin_candidates`
supplies the `.cmd` -> `.exe` precedence (same injectable `windows=` seam the
npx-cache shortcut already uses, so the branch is testable on Linux CI).
POSIX candidates (`node/bin`, `~/.local/bin`, `/usr/local/bin`) are unchanged.

Slimmer redo of #111941 by @KoNit-K, which re-derived the Windows shape
in-place and kept the raw env read.

Fixes #111937

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 19:04:29 -07:00
teknium1 54ed7cbb7b fix: validate the skill name before opening its lock; key lock files on a digest
Review finding on #112218 (major): `_skill_lock_path` opened `<skills>/.locks/<name>.lock`
before the name was validated, so `skill_manage(action='create', name='a'*300)` raised
OSError (File name too long) and a NUL name raised ValueError instead of the handler's
JSON error, and every rejected name ('../../etc', '') left a residue lock file.

- tools/skill_manager_tool.py: lock filename is sha256(basename).lock (fixed width, no
  filesystem limit reachable; `foo` and `category/foo` still share one lock), the redundant
  `_find_skill` rglob is gone, and `skill_manage` runs `_validate_name` on the name
  (create) / basename (other actions) before the lock is opened.
- '.locks' joins the skills-dir exclusion sets (EXCLUDED_SKILL_DIRS, ledger
  _NON_PACKAGE_TOPS, learning-graph/skill-commands skip parts, curator backup excludes).
- tests: 2 invariants in TestSkillMutationLock (rejected names -> JSON + no .locks residue;
  digest-keyed lock shared across name forms), red on the old head.
2026-09-15 19:03:33 -07:00
teknium1 273986f88f fix(skills): route the skill_manage lock through the existing skill_usage lock helper
Slim follow-up to the cherry-picked #111585 (@KoNit-K):

- tools/skill_usage.py: generalize the usage ledger's `_usage_file_lock()` into
  `skill_file_lock(lock_path)` — same fcntl/msvcrt idiom, now thread-re-entrant
  via a per-thread held set (flock is not re-entrant across separate fds; a
  ContextVar would leak "held" into copy_context() timer threads).
- tools/skill_manager_tool.py: drop the third fcntl/msvcrt copy, hashlib and the
  ContextVar; the per-skill lock is `<skills>/.locks/<skill-dir-name>.lock`
  (readable, outside the skill dir so delete/recreate cannot unlink it under a
  waiting writer). Batch locks sort by lock PATH, not name, so two batches
  naming the same skills in different forms cannot deadlock.
- tools/skill_manager_batch.py: plain `with` around snapshot -> commit/rollback
  instead of manual __enter__/__exit__ bookkeeping.
- tests: trimmed to two invariants — the two-writer lost-update test on
  SKILL.md (from #111585) and a re-entrancy/exclusivity test on the helper.
  Dropped: the edit/write_file/remove_file parametrization (same dispatcher
  path as patch) and the category-dir cleanup test (lock files never lived in
  category dirs here).
2026-09-15 19:03:33 -07:00
KoNit-K 0b8b000ce9 fix(skills): serialize skill mutations 2026-09-15 19:03:33 -07:00
teknium1 55e2986dfd fix: walk a group's __cause__/__context__ in the exception node walker
Review finding: _exc_children returned only .exceptions for a group, so
_is_session_expired_error missed a session-expiry marker (or the
InterruptedError override) hanging off a group's __cause__/__context__
that main used to inspect. Groups now yield nested + chain like every
other node; _flatten_messages' "group str() is opaque" rule is unchanged.
2026-09-15 19:02:39 -07:00
teknium1 e1114bdcf9 refactor(mcp): one cycle-safe exception walker for every connect-error scan
The salvaged fix gave `_find_missing` and `_flatten_messages` each their own
visited-set loop, next to the one `_is_session_expired_error` already had —
three copies of the same idiom in one module. Collapse them into
`_iter_exception_nodes` (pre-order, left-to-right, each node once, bounded by
`_EXC_TRAVERSAL_MAX_NODES`) and read all three scans off that list. Acyclic
output is byte-identical: the missing-executable search keeps its depth-first
order and a message-less leaf still renders as its class name.

Tests move from the issue-numbered file into `tests/tools/test_mcp_tool_errors.py`
(mirror of the source module): a two-node cycle renders the real messages, and a
missing stdio binary wrapped deeper than the recursion limit with the chain
looping back to the top is still reported as the missing executable. Both are
red on origin/main (RecursionError).

Co-authored-by: Stephan Mongstad <stephan@users.noreply.github.com>
2026-09-15 19:02:39 -07:00
KoNit-K 030d4caa0e fix(mcp): bound nested connection error traversal 2026-09-15 19:02:39 -07:00
teknium1 2588c908e7 fix: recognise fences opened inside list items and blockquotes
Review finding on #112198: _mask_prose_link_destinations matched
_FENCE_LINE against the raw line, so a fence behind a CommonMark
container prefix (`- ```sh`, `1. ```sh`, `> ```sh`, nested) was not
seen and its body was scored as prose with link destinations masked.
Strip the container prefix before fence matching (open and close).
Bundled-skill rescan vs origin/main: 208 skills, 1447 findings on
both, no new/gone findings, no verdict changes.
2026-09-15 19:01:39 -07:00
teknium1 33292affd6 fix(skills): fenced blocks close only on a matching fence; temp-root rm covers //.. and ..;
Follow-up to the two cherry-picked contributor commits.

The picked fence tracker never checked for a closing fence once a block was
open (the closer test sat inside the not-in-code branch), so every prose link
after any code block was scanned verbatim again and the #111254 documentation
link exemption was lost; a fence line carrying an info string was also accepted
as a closer, which handed the scanner back to prose mode mid-block. Rewrite the
loop around CommonMark fence semantics: a block opens on 3+ backticks/tildes
indented at most 3 spaces (backtick info strings may not contain a backtick)
and closes only on a fence with the same marker, at least as long, and nothing
after it; tab- or 4-space-indented lines are code; an unclosed fence stays
code to EOF. plugin_guard inherits the behaviour through scan_file.

The temp-root exemption in destructive_root_rm now also refuses a parent
segment reached through an empty path segment or followed by a shell
separator, which the first cut let through.

Tests trimmed to one invariant per fix: the fence test covers the six code
shapes plus the prose-link-after-fence control that the picked version broke;
the rm test gains the two residual shapes.

Part of #111334
Fixes #112129
Fixes #111335
2026-09-15 19:01:39 -07:00
JulianCruzet 725713d575 fix(skills): track CommonMark fence state in prose-link masking exemption
replace the boolean fence toggle in _mask_prose_link_destinations with
proper (marker_char, opener_length) tracking so a mismatched-markdown-fence
body or an indented code block cannot re-enable prose-masking over live
command lines. closes an exploitable bypass in the community-source
install path; plugin_guard inherits the fix through scan_file.

also tighten is_indented_code to treat any tab indent (single or double)
as code, per CommonMark §4.4.
2026-09-15 19:01:39 -07:00
KoNit-K 4537869dc8 fix(skills): detect temp-root traversal deletes 2026-09-15 19:01:39 -07:00
teknium1 e133f3f607 fix: device OAuth login scans every advertised authorization server
`hermes mcp login <server> --flow device` took `authorization_servers[0]`
from the protected-resource metadata and failed when that entry was a
browser-only or issuer-inconsistent server, even though a later entry was
the issuer-bound device_code server meant for headless clients (Higgsfield
advertises exactly this shape: a PKCE server first, the device server second).

Discovery now tries each advertised server in order and binds to the first
whose metadata issuer matches its advertised URL and that offers device
authorization. Issuer validation (RFC 8414 / SEP-2468) is unchanged per
server; a single-server resource raises exactly the error it raised before,
and a multi-server resource with no usable entry reports every attempt.

The browser path (`tools/mcp_oauth_manager.py` pre-flight) is deliberately
left on the SDK's own first-entry selection: the SDK's 401-branch discovery
re-selects `authorization_servers[0]` itself, so a divergent pre-flight pick
would only desynchronise the cached metadata from what the SDK authorizes against.
2026-09-15 19:00:42 -07:00
teknium1 704f0b1c91 fix(gateway): background completions reach the chat while the launching turn is still running
`terminal(background=true, notify_on_complete=true)` appended its watcher descriptor to
`process_registry.pending_watchers`, which only the post-turn hooks drain. A process that
finished while the turn that launched it was still running (an agent sleep-polling for
hours) had no watcher task at all: the completion_queue entry sat inert, nothing was
injected, and the chat stayed mute until that turn ended (#112033).

- `_register_completion_watcher` arms the watcher on the live gateway loop at registration
  (`GatewayRunner.arm_process_watcher`, via the existing `_gateway_runner_ref` /
  `_gateway_loop` seam that send_message and cron already use); `pending_watchers` stays
  the fallback while the gateway is not serving (checkpoint recovery at startup, shutdown).
- The agent-notify branch of `_run_process_watcher` keeps its design (the agent's next turn
  is the user-facing report) but, when the launching turn is still active at process exit,
  the injection only queues a follow-up — so the concise receipt is sent to the chat right
  away instead of never. The busy check is taken before injection because the injected turn
  itself installs the adapter's session guard.

Live probe (real process, real GatewayRunner loop, fake telegram adapter, busy session):
before — pending_watchers=1 after exit, 0 watcher tasks, 0 injections, 0 receipts;
after — pending_watchers=0, watcher task armed at launch, 1 injection, 1 concise receipt.
Control (idle session): 1 injection, 0 receipts, unchanged.

Slimmer redo of #112038 by @KoNit-K: same two gaps closed, without a second scheduler
registry / loop attribute on ProcessRegistry and GatewayRunner.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:58:30 -07:00
teknium1 be67e1d31e fix(tools): tolerate an untraversable HOME when probing ~/.local/bin
CI runs the suite as an unprivileged user with HOME=/root in one fixture;
Path.is_dir() raised PermissionError from _user_local_bin_entries and the
run-env builder crashed. An unreadable home has no usable ~/.local/bin, so
treat the OSError as absent.
2026-09-15 18:49:29 -07:00
teknium1 43e7e830fd fix(tools): fold ~/.local/bin into the POSIX PATH completion siblings, tests + docs
Slim follow-up to the salvaged #111790: the helper becomes a list-returning
sibling of _managed_runtime_path_entries (same shape, same "only when it
exists" convention) and loses the Windows check the caller already performs.

Why here and not in the Electron remote spawn: propagating the login-shell PATH
that locateHermes discovered into `exec env HERMES_DESKTOP=1 … hermes serve`
would fix only the Desktop SSH surface; the terminal environment's PATH
completion is the seam every thin-PATH launcher (SSH, systemd, launchd, cron)
already goes through, so the class closes once. Windows twin out of scope.

Tests move to the mirror dir tests/tools/environments/ with an absent-dir
control; FAQ documents the terminal PATH composition.

Fixes #111778
2026-09-15 18:49:29 -07:00
KoNit-K c68e306ea4 fix(tools): include user local bin in POSIX PATH 2026-09-15 18:49:29 -07:00
teknium1 1e2cb57973 fix(approval): session teardown and interrupted leaders withdraw the prompt instead of denying it
clear_session (/new, /reset, auto-reset boundary) stamped entry.result="deny"
before waking the wait, and an interrupted coalesced leader published the
same deny to its followers, so both still rendered outcome="denied" /
"denied by user". Carry the cause on the entry (entry.cancelled) and let
_cancel_cause map a result-less wake to a withdrawn prompt; the wait still
unwinds fail-closed and the leader's own decision is unchanged.
2026-09-15 18:44:46 -07:00
teknium1 6332216384 fix(approval): withdrawn gateway approval prompts no longer read as a user deny
When a gateway approval wait ends without anyone answering — the parent's
delegate_task finishing and tearing the child down, a /stop, or the turn's
notifier being unregistered at turn end — the tool result said
"BLOCKED: Command denied by user" (outcome="denied", user_summary "You denied
this command"). The user never saw or answered the prompt, so the parent agent
went on reasoning about a refusal that never happened (#112026, #22992).

The action stays fail-closed (the command does not run, the model still gets
the NOT-consented stop text), but the attribution is now truthful:

- tools/approval_gateway_wait.py: `_cancel_cause()` reads the existing
  per-thread interrupt-cause channel (`get_interrupt_reason()`, a trusted fixed
  category — no string matching) for the interrupted state and marks a
  notifier-unregister wake (event set, result None) as "the turn ended before
  the prompt was answered". Both the direct and the coalesced-follower wait
  return `cancelled=<cause>`; the post_approval_response hook fires
  choice="cancelled" instead of "deny"/"timeout".
- tools/approval.py: a cancelled decision renders
  "BLOCKED: Command approval was withdrawn before the user answered (<cause>)."
  with outcome="cancelled" and its own user_summary; an explicit /deny is
  untouched.
- tools/delegate_tool_child_run.py: `_signal_child_stop` publishes a fixed
  tool_reason ("parent delegation ended"; the late-child mirror forwards the
  parent's own category) so a child's pending approval can tell teardown from a
  user /stop — previously it rode the default "explicit stop requested".
- tools/file_tools_write_guards.py / tools/approval_prompt.py: the protected
  instruction-file gate and MCP elicitation consume the same key instead of
  reporting "denied by the user" / "decline".

Co-authored-by: zccyman <16263913+zccyman@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:44:46 -07:00
teknium1 54d7f75590 fix(bot-mode): a pending command approval is not a failed DM delivery
terminal_tool's approval gate answers `status: pending_approval` with an
EMPTY `error` (#28323) and no `session_id`, so _spawn_delivery's specific
branch (`if parsed.get("error")`) was skipped and every unanswered
approval fell through to "Delivery to X failed to start: no process id
returned" — blaming the spawn for an approval nobody in a non-interactive
turn (api_server, `hermes peer dm`, cron) could grant.

- _spawn_delivery: the pending shape gets its own message (the runner
  command needs terminal approval nobody in this turn can grant); a
  local/peer DM adds "nothing was sent — approve it or add it to
  command_allowlist and send again". Ownership is never transferred, so
  the existing finally still reclaims the plaintext DM file.
- _try_relay_delivery: the envelope is queued on disk BEFORE the reply
  waiter spawns and the Desktop drains it independently, so ANY waiter
  spawn failure is a lost wake-up, not a failed delivery; reporting it as
  an error made the sender resend and deliver the message twice. The
  relay path now returns the shape _start_delivery's live-owner branch
  already uses (status queued + notification_error + "Do NOT resend")
  instead of inventing a new status value nothing reads.

Slimmer redo of #92971 by @jonpol01 (same diagnosis, same relay/local
split on `dm_file is None`); the source-text contract test and the
`sent_no_reply_wake` status were dropped.

Fixes #111716

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
2026-09-15 18:44:23 -07:00
fangliquanflq 27e3fc51ff fix(tools): keep async delegation results past a failed durable write and never prune live records
Two defects in tools/async_delegation.py:

- _push_completion_event called _persist_completion unguarded before
  publishing onto completion_queue. One sqlite3 error (locked/full
  state.db) dropped the completion event, left the record parked on
  "finalizing" (a permanently leaked max_concurrent_children slot) and
  let recover_abandoned_delegations later rewrite a succeeded unit as
  "unknown". The write is now try/except: the failure is logged and the
  event is still delivered, so _finalize flips the status and frees the
  slot. A lost durable row is acceptable degradation; a lost result and
  a leaked slot are not.

- _prune_completed_locked treated anything != "running" as finished,
  while the module's own _LIVE_STATES also names stalling/finalizing.
  A stalling record has no completed_at, so it sorted oldest and was the
  first eviction candidate once the retained cap overflowed; its late
  runner return then hit the missing-record path and the real result was
  dropped. The predicate is now `status not in _LIVE_STATES`.

Slim redo of #76606 (earliest fix) and #112031: the converge/shield/
delete-row machinery both PRs built around the write is dropped as
defense-in-depth; the two core hunks are ported as-is.

Fixes #76605
Fixes #112030
Co-authored-by: luckystar2026 <1393268817@qq.com>
2026-09-15 18:43:56 -07:00
teknium1 80f76edfaf fix(web): rescue eligibility asks the provider whether the ring was walked
Follow-up to the cherry-picked gateway fix: instead of re-inferring "keyless
mode" from the key env var (wrong for Firecrawl, whose managed-gateway and
self-hosted routes bypass the ring without a key), `_rescue_eligible` asks the
ring vendor's own predicate — `_use_keyless_ring()` for Firecrawl, `use_keyless`
for the others. That covers the persisted `nous` selection the contributor fix
handled AND the legacy never-configured fallback onto a ready gateway, plus
`FIRECRAWL_API_URL`. A ring vendor that actually walked the ring stays
ineligible (its failure means the ring already failed). Docs mention the
gateway route is rescued.
2026-09-15 18:43:05 -07:00
KoNit-K 911eea567f fix(web): rescue failed Nous gateway searches 2026-09-15 18:43:05 -07:00
teknium1 45a4db2225 fix(web): key extract cache on metadata.sourceURL too, pin redirect case
Follow-up to the cherry-picked "cache extracts by returned URL": Keenable and
Firecrawl report the post-redirect address in `url` and the REQUESTED URL in
`metadata.sourceURL`, so matching on `url` alone left every redirected page
uncached. Accept either field, as long as it names a URL from this batch;
anything else is served but never cached (a miss re-fetches, a mis-key poisons
the cache for the whole TTL). Docs: say the cache key is the requested URL the
provider reports, not the batch position.

Co-authored-by: nemofq <5635994+nemofq@users.noreply.github.com>
Co-authored-by: wooyongbin3-cpu <256294002+wooyongbin3-cpu@users.noreply.github.com>
2026-09-15 18:42:38 -07:00
KoNit-K ffd02f37e8 fix(web): cache extracts by returned URL 2026-09-15 18:42:38 -07:00
KoNit-K c9fa191334 fix(tools): support daemon pool workers on Python 3.14
Cherry-picked from #111814. The same feature-detected fix was proposed earlier in
#58699, #65182 and #57459 (final form).

Co-authored-by: nankingjing <76432572+nankingjing@users.noreply.github.com>
Co-authored-by: TheNeuralVault <jdkabattles@gmail.com>
Co-authored-by: gongyi <yigongsyl@gmail.com>
2026-09-15 18:41:46 -07:00
teknium1 8e16bde491 test: fold the schema-retry context test into the output-schema module
Reuse tests/tools/test_delegate_output_schema.py's _StubChild instead of a
new one-test file with its own double; the invariant (the retry turn sees
is_delegated_child_context() True and the flag is restored afterwards) is
unchanged. Trim the source comment to the WHY.
2026-09-15 18:41:19 -07:00
moep90 e002cdb92d fix(delegation): run the schema-retry turn in the delegated-child context
_validate_child_output_schema issues a second run_conversation on the child when
the first answer fails the declared output_schema. The main child turn is wrapped
in delegated_child_context; this one was not. It runs on the parent worker's
thread, where HERMES_KANBAN_TASK is set and nothing marks the execution as a
child, so every identity gate keyed on is_delegated_child_context() fails open.

The visible effect is the kanban stop guard: it nudges the child to call
kanban_complete or kanban_block. A child owns no board task and carries no kanban
toolset, so it cannot, and the nudge text ("do not narrate intent", "finish any
remaining deliverable") displaces the structured answer the retry exists to
produce. The retry then fails the same schema and delegate_task reports an error
for a child whose work was already complete.

Observed with four children, each nudged during its retry:

  [subagent-0] Kanban worker tried to exit without kanban_complete/kanban_block
  [subagent-2] Kanban worker tried to exit without kanban_complete/kanban_block
  [subagent-3] Kanban worker tried to exit without kanban_complete/kanban_block
  [subagent-1] Kanban worker tried to exit without kanban_complete/kanban_block
  4/4 - Final answer does not satisfy the declared output_schema (after 1 retry)

Wrap the retry the same way the main turn is wrapped. The context is entered and
exited around the single call, so nothing outside the retry sees it.

Signed-off-by: moep90 <volleyballlive@googlemail.com>
2026-09-15 18:41:19 -07:00
teknium1 b121a02416 fix(delegate): composite parents can grant their included toolsets to children
_expand_parent_toolsets built the parent's tool surface from each
toolset's declared `tools` only, so a composite parent's `includes` were
invisible: a child of a `debugging` parent (terminal/process_manage +
includes web/file) asking for `file` or `web` was refused, and `safe` /
`hermes-gateway` parents could grant nothing but their own name. Same
root cause as the `_strip_blocked_tools` fix in the previous commit
(#111700, "Related" section).

Both sides of the subset check now use the resolved static surface
(`resolve_toolset(name, include_registry=False)`), so a child may request
any toolset whose real tools the parent genuinely holds, and still never
gains a tool the parent lacks. Candidates that resolve to nothing are not
expanded into (they cannot be a meaningful subset).

Co-authored-by: DresvyanskiyDenis <dresvyanskiydenis@gmail.com>
2026-09-15 18:40:52 -07:00
KoNit-K 3d4c4cc24d fix(delegate): retain composite child toolsets 2026-09-15 18:40:52 -07:00
KoNit-K ae5666f7fc fix(tools): quiet expected unavailable toolsets 2026-09-15 18:40:29 -07:00
teknium1 decf8e3f26 fix(tools): keep an empty required: [] in sanitized tool schemas
`_sanitize_node` deleted the `required` key whenever the pruned list came
out empty. Four built-in tools (skills_list, todo, delegate_task,
session_search) declare `required: []`, so they left the sanitizer with no
key at all. Strict OpenAI-compatible proxies read the missing key as
`null` and 400 the whole request ("null is not of type array"), which is
non-retryable and kills the session on its first call.

An empty array is valid for every backend; the pruning was added (34c3e67)
to drop names that are not in `properties`, not to delete the key. Keep the
key with the filtered list, even when that list is empty.

Fixes #111684
Fixes #59386
Co-authored-by: Cr4ckMe <jiqing.liu@whu.edu.cn>
2026-09-15 18:39:09 -07:00
Konstantin Khlopkov 0724a6a0fd fix(tools): persist the kill outcome when the reader thread finalises first 2026-09-15 18:38:44 -07:00
teknium1 0959224313 fix(kanban): claim-less complete no longer closes a live worker's run
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).

Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.

Fixes #111764
2026-09-15 18:34:40 -07:00
KoNit-K 0aec64c874 fix(skills): honor profile-scoped readiness secrets 2026-09-15 18:33:15 -07:00
Sahil Vishnalya afe9e25c57 fix(stt): consume lazy whisper segments inside the CUDA→CPU retry guard
faster-whisper's `model.transcribe()` returns a lazy generator; ctranslate2
dlopens the CUDA runtime on the FIRST encode, which happens while the segments
are iterated in `_join_confident_segments()` — outside the try/except that
implements the CUDA → CPU fallback in `_transcribe_local`. On a host with an
NVIDIA driver but no CUDA runtime (Windows `cublas64_12.dll`, Linux
`libcublas.so.12`) the model loads fine, the error escapes the guard, and every
voice note fails with "Local transcription failed: Library cublas64_12.dll is
not found or cannot be loaded" until the user pins `stt.local.device: cpu`.

Materialize the segments inside the guarded block (first attempt and CPU retry)
so the dlopen failure reaches the existing evict-and-retry-on-CPU path.

Salvaged from #103848 by @Sahilvishnaliya (earliest fix of this class).
Trimmed during salvage: the `_CUDA_LIB_ERROR_MARKERS` additions (`cublas64_`,
`cudnn64_`, `cudart64_`) — the Windows message already matches the existing
"cannot be loaded" marker, proven by the live probe with the reporter's exact
string; the 6-test file was reduced to 2 invariant tests in the existing suite.

Fixes #111929
Fixes #105295
Part of #103793 (the CPU fallback now fires; GPU-wheel install is separate)

Co-authored-by: atmaksri <sri.atmakur@gmail.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: isoenthusiast <287677567+isoenthusiast@users.noreply.github.com>
2026-09-15 18:25:17 -07:00
KoNit-K 065bc91846 fix(checkpoints): report legacy archive deletion failures 2026-09-15 18:24:22 -07:00
teknium1 f9c3a8a186 feat: hermes update and doctor tell you when /rollback checkpoints are on and large
Checkpoints were enabled by default from 9e845a6e (2026-03-16) until #20709
(2026-05-06) flipped the default back to off; the migration of that window
wrote `checkpoints.enabled: true` into user configs, where a later default
flip cannot reach it. Users who never type /rollback have carried a GB-scale
`~/.hermes/checkpoints/store` since (one live install: 1.2 GB across 250
projects, mostly disposable worktrees and /tmp dirs), and the cap cannot
bring it down because every project keeps at least one snapshot.

Silently flipping the key back is indistinguishable from overriding a real
opt-in, so this surfaces it instead: `checkpoint_footprint_notice()` returns
one line when checkpoints are enabled AND the store is at or above
`max_total_size_mb`, naming the opt-out (`hermes config set
checkpoints.enabled false` + `hermes checkpoints clear`) and the
retention knob. `hermes update` prints it with the post-update notices;
`hermes doctor` reports it as a warning after the state.db check.
2026-09-15 12:06:08 -07:00
teknium1 3272fb35aa docs: profile-scope invariant in AGENTS.md — one process serves many profiles; out-of-turn code binds its scope
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.

Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).

Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
2026-09-15 10:59:22 -07:00
teknium1 804707bea6 fix: checkpoint store gc never runs inside a tool call or gateway startup
Symptom: `hermes update` sat for ~40s after "Refreshing cua-driver" and ended
with "Fleet version check returned no rows" (exit 1); the restarted gateway
took 26s to reach "Starting Hermes Gateway" instead of the usual 3s. The
gateway constructor was running `maybe_auto_prune_checkpoints` synchronously,
before the control socket, adapters and the code_sha stamp, and on a 1.2 GB
store its `git gc --prune=now` (a full repack) takes 20-28s — twice, because
the size-cap shrink gc'd again even when it could drop nothing.

The same defect sat on the tool-call path: `CheckpointManager._take` ran
`_enforce_size_cap`, whose `_shrink_store_to_cap` returned True without
dropping anything and triggered a 20-28s gc on the first file-mutating tool
call of every turn once the store was over the cap. That loop also re-measured
a pack size that cannot move without a gc, so a single over-cap checkpoint
dropped 20 rounds of history and flattened every project to one snapshot.

- `_take` never gcs: `_prune` and `_enforce_size_cap` rewrite refs (cheap),
  drop at most one snapshot round, and mark the store `.gc-pending`.
- `prune_checkpoints` gcs only when a ref moved (project deleted, or the
  pending marker), and its cap loop is drop -> gc -> re-measure.
- `maybe_auto_prune_checkpoints` claims the interval marker before the run
  so a failing prune costs one day, not a gc per housekeeping tick.
- `auto_prune_from_config` is the one config-driven entry point; the gateway
  calls it from the housekeeping tick (last chore), the CLI from a daemon
  thread. Nothing on either startup path waits for git.

Live A/B on a copy of a real 1.2 GB / 224-ref store: checkpoint 20.5s ->
1.2-1.6s (0 inline gc); the single repack (19.6s) now runs in the prune.
2026-09-15 10:57:16 -07:00
teknium1 d84ece48b8 fix(mcp): Figma OAuth login completes despite the omitted iss parameter
Figma's authorization-server metadata advertises
authorization_response_iss_parameter_supported and its redirect omits iss,
so the mcp SDK's RFC 9207 check discarded every valid code and login never
finished. For that one issuer the provider fills a missing iss with the
discovered issuer and warns; a mismatching iss still fails and every other
server keeps the strict rule.

Fixes #111135
2026-09-15 09:29:00 -07:00
Robin Fernandes 59fad62a40 fix(free-tier): review follow-ups — read the classifier's context, never replace a locked identity, re-inventory on retry
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
  the long-wait rate-limit check read the turn's extract_api_error_context()
  dict, which never carries welcome_refusal / welcome_route. They now read
  classified.error_context, where _nous_welcome_tier parks them; the guard
  records the classifier's reset_at. Tests drive the real classifier and the
  real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
  account (anon_account_locked) is now retired without replacement, matching
  the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
  re-inventories, so a provider connected during the cooldown keeps
  inference.
- The desktop's setup.ready listener only refreshes an untouched picker
  (oauth mode, no local endpoint, idle flow) and re-checks after the
  readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
  lock that _send re-acquires; the log is copied out first.

Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
  behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
  scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
  extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
  loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
teknium1 14bd358f5e fix: unlink the dm plaintext file too once a live delivery settled
The live branch of _run_delivery returns from _wait_live_dm before the
try/finally that removes <dm>.txt, so a settled delivery dropped its
.live.json intent but left the sibling .txt — the same plaintext — on disk.
_wait_live_dm now takes the dm file and removes both on settled; pending
and failed outcomes still keep both for the retry.

Review finding: settled live path unlinked <dm>.live.json but the sibling <dm>.txt survived.
2026-09-15 06:34:25 -07:00