Commit Graph

3009 Commits

Author SHA1 Message Date
Teknium 382060f022 feat(mcp): speak the 2026-07-28 stateless protocol
Phase 2 of the MCP 2026-07-28 migration (#69931), on top of the SDK 2.x
migration (#88180):

- Protocol-era negotiation (_negotiate_session): per-server `protocol`
  config key — auto (default, handshake-first with server/discover
  fallback on -32022/-32601), stateless (discover-first), legacy
  (handshake only). Auto is handshake-first deliberately: zero extra
  round-trips and zero behavior change for the entire existing server
  fleet, while 2026-07-28-only servers now connect via the fallback.
  All four transport call sites (stdio, SSE, new HTTP, legacy HTTP)
  route through the one choke point, so the CLI/desktop probe path
  inherits it too.
- SEP-2549 list caching: tools/list ttlMs/cacheScope hints are captured
  during discovery and bound to the lazy-startup schema cache — TTL'd
  entries expire and force a live re-probe; hint-less (pre-2026)
  servers keep the never-expires behavior. Pagination continuation now
  speaks both SDK generations (params= vs cursor=).
- SEP-837: OAuth client metadata declares application_type=native
  (config-overridable), with a fallback for 1.x-era metadata models.
  (RFC 9207 iss validation and SEP-2352 issuer-keyed credentials are
  native to SDK 2.0's OAuthClientProvider — verified, no client-side
  gap.)
- SEP-2577 deprecation posture: SamplingHandler docstring marks the
  Sampling feature as upstream-deprecated (12-month window) — kept
  fully functional, closed to new capability.
- Docs: `protocol` key in the MCP config reference.
2026-08-17 03:03:35 -07:00
Xinyu Du bbc894b0ab docs(tools): tighten subprocess env isolation docstrings
Shorter, single-source ownership explanation for
_strip_hermes_owned_pythonpath (the code-level Check comments already
carry the per-branch detail; the docstring only needs the contract).
2026-08-17 02:10:32 -07:00
Xinyu Du 98389e7895 refactor(tools): simplify subprocess env isolation coverage
Same behavior, same coverage, less boilerplate (test file 1691 -> 1512
lines; PR diff unchanged in semantics).

Production (mechanical only):
- Extract _strip_hermes_owned_pythonpath_and_runtime_markers(): the three
  builders (_make_run_env, _sanitize_subprocess_env, hermes_subprocess_env)
  ran the identical strip-then-pop-markers sequence in the same order
  (ordering is load-bearing for VIRTUAL_ENV validation); the helper makes
  that explicit once instead of three times.

Tests:
- Non-owned preservation: 11 single-shape tests -> one parametrized matrix
  (user/Nix/other-version/python2.7/pythonX.Y-contained/raw spelling/empty
  component/empty PYTHONPATH) + one runtime-shaped matrix (other-version SP,
  venv-SP descendant, repo direct child, repo deep child).
- Owned stripping: venv SP, repo root (independent parents[2] computation),
  duplicates, all-owned key removal, mixed ordering -> one matrix.
- Builder integration: _make_run_env/_sanitize_subprocess_env/
  hermes_subprocess_env venv-SP stripping -> one parametrized test;
  same for the four PYTHONHOME builders (incl. build_subprocess_env).
- Junction: same-named non-owned negative control now covers both the
  configured-root location and an unrelated location; shared
  _physical_repo_root helper; profile resolution matrix (root->named,
  profile-shaped->named no nesting, profile-shaped->default, custom root).
- Every independent proof preserved: home-level junction, repo-level
  junction, profile interaction, negative identity control, uv-base lexical
  VIRTUAL_ENV, validated/unrelated VIRTUAL_ENV, no-scrub escape hatch,
  #84500 same-env/external-env composition, PYTHONHOME removal, real
  Windows-only semantics, POSIX fail-closed backslash paths.
2026-08-17 02:10:32 -07:00
Xinyu Du d67cd58e77 fix(tools): recover repo-level junction lexical root via exact identity
Second real-world topology reported and confirmed on native Windows 11:
the repository itself is a cross-drive junction (D:\hermes\hermes-agent ->
C:\...\hermes-agent) under a real HERMES_HOME directory.  The editable
import spelling resolves to the physical location, so _hermes_repo_root is
physical while the launcher writes the lexical spelling into PYTHONPATH.
The home-relative mapping cannot express a cross-drive link (commonpath
raises on different drives), so the lexical repo root survives stripping;
and with the repo alias missing, a lexical VIRTUAL_ENV
(D:\hermes\hermes-agent\venv) also fails _validated_runtime_venv, so the
venv site-packages survives too (uv-base gateway: both entries survive).

Fix: after the existing home/profile-root mapping, try the single
deterministic candidate <lexical root>/<repo dirname> for every trusted
home candidate (configured home, plus the profile root when the configured
home is a profile path) and accept it only when strict resolve proves it is
the exact physical repo root (fail-closed: missing paths, real directories
that are not the known repo, and unrelated spellings are never aliased).
This also re-enables the VIRTUAL_ENV validation for lexical venv spellings,
so uv-base gateway site-packages cleanup follows the repo alias.

Tests: repo-level junction positive + negative control (same-named real
directory preserved), profile-home + repo-level junction combination,
lexical VIRTUAL_ENV validation after recovery (root + site-packages
stripped, user entries kept), and a no-provenance lookalike preserved.
The execute_code composition test now compares composed paths with
os.path.normcase so a Windows case-only spelling difference (resolve() vs
abspath() casing) can never fail the composition contract.
2026-08-17 02:10:32 -07:00
Xinyu Du 57d94dd8dc fix(tools): keep junction lexical root across profile re-home
Confirmed on native Windows 11 with a real junction and the real startup
chain: when the desktop/CLI spawns the backend with HERMES_HOME in the
configured (lexical) spelling and --profile / sticky active_profile is in
play, _apply_profile_override() re-homes HERMES_HOME through
resolve_profile_env(), which resolves the junction under the platform
default and returns the PHYSICAL spelling.  tools.environments.local is
imported after that mutation, so _hermes_repo_root_aliases is built from
the physical home, the lexical repo-root spelling written into PYTHONPATH
by the launcher (D:\hermes\hermes-agent) is not derivable, and the entry
survives stripping (reproduced: cases --profile default / named / sticky
active_profile / cross-drive junction all leave it in place; no-profile
strips it).

Two narrow changes, no heuristics, no new env vars:

- hermes_cli/profiles.py::resolve_profile_env: when HERMES_HOME is set,
  the configured spelling IS the launch root (junction-transparent,
  physically identical dirs); keep it instead of re-deriving the native
  default.  This is the same producer contract _preserve_hermes_home_path
  already follows.
- tools/environments/local.py::_build_hermes_repo_root_aliases: when the
  configured home is a profile home (<root>/profiles/<name>), also derive
  the root spelling lexically (parent of the profiles component, same
  rule get_default_hermes_root uses) and run the exact-ownership mapping
  against it, so the launcher's lexical root is recovered after re-home
  without ever matching arbitrary descendants of HERMES_HOME.

Regression test test_profile_rehome_keeps_junction_lexical_alias covers
junction + profile re-home + inherited lexical PYTHONPATH end to end.
2026-08-17 02:10:32 -07:00
Xinyu Du deb4953776 test(tools): make subprocess env regressions Windows-portable
The PYTHONPATH/PATH sanitization suite was written POSIX-centric and
failed on real Windows 11 (reproduced natively: 4 failures before this
change).  Fix the tests to express the true per-platform contract:

- test_other_major_version_site_packages_preserved /
  test_make_run_env_injects_hermes_bin_dir: build inputs with
  os.pathsep instead of hardcoded ':'.
- test_make_run_env_appends_homebrew_on_minimal_path: split on
  os.pathsep, neutralise Git Bash dir prepending, and assert the
  documented Windows passthrough (_append_missing_sane_path_entries is
  a no-op off POSIX) instead of the Homebrew append.
- test_make_run_env_real_launchd_path_gains_homebrew: mark
  macos_only per repo OS-marker policy (the regression is the macOS
  launchd PATH; the merge is a passthrough on Windows).
- test_configured_home_alias_matches_launcher_output: create the
  configured-home link via a helper that falls back to an unprivileged
  directory junction (cmd /c mklink /J) when symlink creation raises
  WinError 1314, and skips with a clear reason if no mechanism exists.

Also correct a stale comment in execute_code: the child is not always
the same Python as Hermes (project mode can select an external venv),
so the strip is about compatibility, not redundancy.
2026-08-17 02:10:32 -07:00
Xinyu Du 6e9eeb5413 fix(tools): harden subprocess Python runtime ownership 2026-08-17 02:10:32 -07:00
Xinyu Du 73b49f473a fix(tools): tighten Hermes PYTHONPATH ownership semantics
Adversarial review of the previous two commits (and #78917 itself)
found three ownership-boundary issues; this commit addresses them:

1. Repo direct-child over-strip (Finding A)
   No launcher injects <repo>/tools or another direct child as an
   independent PYTHONPATH entry - audited all four producers (Electron
   electron-main.mjs, gateway/run.py::_ensure_windows_gateway_venv_imports,
   cron/scheduler.py::_windows_cron_python_invocation,
   tui_gateway/host_supervisor.py).  The depth<=1 rule deleted user paths
   that merely live under the repo directory; only the EXACT repo root is
   now stripped.

2. Windows junction/symlink alias (Finding B)
   The gateway launcher renders Hermes-owned paths under the configured
   HERMES_HOME spelling (gateway_windows.py::_preserve_hermes_home_path),
   which may be a junction to another drive, so it differs lexically from
   the resolved repo root.  _hermes_repo_root_aliases now carries both the
   resolved and unresolved spellings; both are recognized as Hermes-owned.

3. Stale abstraction rename (Phase 4)
   _strip_mismatched_site_packages -> _strip_hermes_owned_pythonpath:
   the cross-version heuristic is gone, so the old name misdescribes the
   behavior (ownership-based, not version-based).

Tests: direct-child now preserved; junction alias stripped (lexical pair
monkeypatched); Windows-only real-semantics test added (POSIX test remains
a safety test); mixed-ordering, duplicate-Hermes, and no-scrub PYTHONHOME
contract tests added.  Full file: 52 passed / 16 failed (identical failure
set to base, all isolation-venv environment issues).
2026-08-17 02:10:32 -07:00
Xinyu Du 850686a515 fix(tools): sanitize inherited PYTHONHOME (#75018)
The gateway runs inside its own venv; if its PYTHONHOME leaks into
subprocesses (terminal commands, cron no_agent scripts, TTS providers),
any child interpreter redirects its stdlib search to the Hermes venv and
crashes with version-mismatch errors before importing anything.

PYTHONHOME is now part of _ACTIVE_VENV_MARKER_VARS so all env builders
(_make_run_env, _sanitize_subprocess_env, hermes_subprocess_env, and
build_subprocess_env used by cron) drop it, consistent with Hermes'
existing PYTHONHOME handling in managed_uv.py and sqlite_runtime.py.
execute_code already scrubbed it via _SAFE_ENV_PREFIXES.

Tests cover all four builders plus the marker constant.
2026-08-17 02:10:32 -07:00
Xinyu Du ced80b2a20 fix(tools): preserve user PYTHONPATH entries (#74817 follow-up)
Remove the cross-version heuristic from _strip_mismatched_site_packages:
the subprocess env builder cannot know which Python version a child will
run, so judging user PYTHONPATH entries against the backend interpreter's
version deletes legitimate paths meant for a different child Python
(e.g. /custom/lib/python3.13/site-packages while Hermes runs 3.11).

Also fix over-strip: entries merely containing a pythonX.Y path component
(e.g. /opt/tools/python3.13/bin) were stripped even though they are not
site-packages. Hermes-owned entries (repo root, own venv site-packages)
are now identified by path ownership, not by version.

Regression tests cover both cases; user paths with any pythonX.Y
component are preserved.
2026-08-17 02:10:32 -07:00
elphamale 23a86594cc fix(mcp): read the elicitation schema under the SDK's real field name
`ElicitationHandler` read `params.requested_schema`, but on the pinned
`mcp==1.28.1` the model field is spelled `requestedSchema`. The getattr
always missed and returned its `{}` default, so
`_format_elicitation_schema_summary` took its no-properties branch and the
approval prompt collapsed to the generic

    Approval requested by MCP server '<name>'.

for every request. The field names, types, and descriptions the summary
exists to surface never reached the user, so an elicitation asking for a
card number rendered identically to one asking for a nickname — consent
without the substance of what was being consented to.

Read both spellings rather than just correcting to the 1.x name: mcp 2.0
renames this field to `requested_schema` (it renamed every model field to
snake_case and kept camelCase only as a serialization alias, which
pydantic does not expose to attribute access), so a dual read is correct
on either SDK generation and does not go wrong again on the next bump.
Verified against real 1.28.1 and 2.0.0 installs.

Every existing test in tests/tools/test_mcp_elicitation.py builds a
duck-typed `SimpleNamespace` stand-in, which carries whatever field name
the test wrote and therefore cannot detect a mismatch with the real model.
Add one test that constructs the actual `ElicitRequestFormParams` and
asserts the requested field name reaches the consent description; it fails
on the unfixed tree. The cheap stand-ins are left alone elsewhere.

Found while porting the tree to the mcp 2.x SDK in #76736, but independent
of it: this reproduces on the current pin with no other changes, #76736
does not touch this line, and the two branches merge cleanly in either
order.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:56:50 -07:00
Teknium 9adc900ab2 fix(bot-mode): route bot-to-bot sends through --create-if-missing
The teammate-messaging protocol told bots to send with
`chat -c "Bot Chat" -Q -q` and, on 'No session found', fall back to a
manual two-step (send without -c, then sessions rename) — a dance the
CLI already made unnecessary when --create-if-missing landed (#86794).
Profiles that never went through the Bots-panel birth flow (CLI-created,
pre-Bot-Mode, remote-source) hit that miss on every first contact, and
background sends swallowed the error entirely (the original silent-drop
in Hermes-Bot-Mode#48 / #88059).

- tools/bot_mode_probe.py: protocol command gains --create-if-missing
- bundled plugin: Bot Chat prompt section + @mention handoff note use
  the flag; rename-dance instructions deleted

The capability epoch hashes the protocol section, so existing eternal
Bot Chat sessions pick the new instructions up on their next message via
the established once-per-change rebuild — no per-turn cache drift.

Live-verified: fresh profile with zero sessions, protocol command
created 'Bot Chat' and delivered (PONG round-trip); second send resolved
the same session by title (no duplicate); missing-title send WITHOUT the
flag still errors loudly on stderr.
2026-08-16 23:28:13 -07:00
elphamale 77ed1bbf40 fix(mcp): seed MCP-Protocol-Version from the handshake version, not the latest
The HTTP transport seeded `MCP-Protocol-Version` from LATEST_PROTOCOL_VERSION,
which on mcp 2.x is 2026-07-28 — a revision that replaced the `initialize`
handshake with a per-request envelope. But this transport connects through
`ClientSession.initialize()`, which sends LATEST_HANDSHAKE_VERSION (2025-11-25)
in the body. Header and body therefore disagreed by construction, and a
conforming 2.x server honours the header: it routed the request onto its
per-request-envelope ladder and rejected the legacy body with

    params._meta is missing the required envelope key(s):
    io.modelcontextprotocol/protocolVersion,
    io.modelcontextprotocol/clientCapabilities

Observed against a live MCP endpoint, and confirmed by probing the same
endpoint three ways: the header at 2026-07-28 is rejected, at 2025-11-25 it
succeeds, and with no header at all it succeeds.

Third defect in this migration from one cause: the 2.x bump changed what an
existing constant *means* without revisiting its uses. The header seed was
written when LATEST_PROTOCOL_VERSION was 2025-03-26 and was correct then.

Seeded from LATEST_HANDSHAKE_VERSION, imported with a fallback to
LATEST_PROTOCOL_VERSION for SDKs predating the split, where the two are the
same thing and header and body agree either way. An explicitly configured
header still wins — that override is why servers demanding a specific revision
can have one, and a test pins it.
2026-08-16 23:26:10 -07:00
elphamale 2e1d724e3e fix(mcp): accept both SDK generations' streamable-HTTP transport arity
`streamable_http_client` yields `(read, write, get_session_id)` on mcp 1.x and
`(read, write)` on 2.x. `_run_http` unpacked a fixed 3-tuple, so on 2.x every
HTTP and SSE MCP server failed its handshake with `ValueError: not enough
values to unpack (expected 3, got 2)` and parked after exhausting its retry
ladder. Only stdio servers kept working.

This is the same defect as the import gating fixed earlier in this branch, one
layer further in. That fix's own comment claimed reaching
`streamable_http_client` was "the path that does work" — reaching it was
necessary and not sufficient, and the comment asserted the half that was never
exercised. Corrected along with the code.

Unpacked positionally rather than by arity, since this file deliberately
supports both SDK generations and `get_session_id` was never used here.

The reason this survived review is worth the test it now has: the existing
coverage in test_mcp_client_cert.py fakes the transport with a 3-tuple, so it
encoded 1.x's shape into the assertion and passed on 2.x regardless. The new
test drives `_run_http` once per arity the supported SDK range actually yields,
and asserts the streams handed to ClientSession are the first two — positional,
because 1.x's third element is not a stream. Verified it fails on the 2.x case
without this change.

Found while pointing a real HTTP MCP server at a live deployment running this
branch: the server parked at startup and no tool from it ever registered.
2026-08-16 23:26:10 -07:00
elphamale 11a9dcf567 feat(mcp): migrate to the mcp 2.x SDK
mcp 2.0.0 implements MCP revision 2026-07-28 and makes three breaking
changes Hermes sits on top of: `mcp.server.fastmcp` is gone, every model
field is renamed to snake_case (camelCase survives only as a
serialization alias, which pydantic does not expose to attribute
access), and the SDK's own HTTP stack moved from `httpx` to `httpx2`.

Bump the pin across the dev/mcp/computer-use extras and port the tree:

- `mcp_serve.py` and `agent/transports/hermes_tools_mcp_server.py` move
  from `FastMCP` to `mcp.server.MCPServer`, which has the same
  decorator/add_tool surface. The hermes-tools server already
  synthesised `__signature__` from Hermes' JSON Schema, which is exactly
  what 2.0's `add_tool` reads.
- SDK model reads go through `mcp_field(obj, snake, camel)`, which reads
  both spellings. A single-spelling read fails *silently* on the other
  generation — empty tool schemas, dropped structured content, tool
  results vanishing from sampling conversations — and `mcp` is an
  optional extra users install at their own version.
- `sdk_httpx()` resolves the httpx flavour from the SDK's own transport
  module, so objects handed to `streamable_http_client`, the `sse_client`
  factory, and the OAuth metadata helpers come from the module the
  installed SDK actually imports.
- HTTP support is gated on either streamable-HTTP entry point, not just
  the deprecated alias 2.0 removed.
- OAuth: `OAuthClientProvider` lost its `timeout` argument (the
  configured `oauth.timeout` now bounds the callback waiter's own poll
  loop, where the browser round-trip was always awaited), and
  `callback_handler` must return `AuthorizationCodeResult` rather than a
  tuple. 2.0 also validates the RFC 9207 `iss` parameter, so the
  callback handler and paste fallback capture it.

`mcp`/`mcp-types` 2.0.0 are inside the 14-day `exclude-newer` window, so
two narrow `exclude-newer-package` entries unblock `uv lock`, annotated
for removal on or after 2026-08-11. `httpx2` needs no exemption: 2.7.0 is
already outside the window and satisfies mcp's floor.

Refs #69931

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 23:26:10 -07:00
Yiipu 2824899321 fix(terminal): correct off-by-one in _hermes_repo_root path resolution
_hermes_repo_root used parents[1] which resolves to tools/ instead of
the repository root. The file lives at tools/environments/local.py so it
needs parents[2] to reach the actual repo root that Electron injects
into PYTHONPATH.

The test test_repo_root_stripped reused the module constant under test
as its input. This made it pass regardless of what the constant pointed
at. The test now computes the real repo root independently from the
source file location. It fails with the old parents[1] code and passes
with the fix.

Reported by spfcraze in PR #78917 review.
2026-08-16 23:26:04 -07:00
mcjoys 43c463fa95 fix(terminal): strip Hermes-venv site-packages from terminal subprocess PYTHONPATH to prevent cross-version ABI conflicts
The Desktop Electron process injects the Hermes venv's site-packages path
(e.g. .../python3.11/site-packages) into PYTHONPATH so the Python 3.11
backend can import its packages. When this PYTHONPATH leaks into terminal
subprocesses running a different Python version (e.g. Python 3.13), 3.11
C extension modules appear on sys.path ahead of the correct 3.13 versions
and crash with ImportError (PIL _imaging, cryptography, etc.).

Replace the existing blunt pop of PYTHONPATH from _ACTIVE_VENV_MARKER_VARS
with a surgical Hermes-venv-aware filter:

- Parse each PYTHONPATH entry by path
- Strip only paths under ~/.hermes/hermes-agent/venv/.../site-packages
- Preserve the Hermes source root (needed for import hermes_cli)
- Preserve all user-set PYTHONPATH entries

The same filter is applied in all three env builders:
- _make_run_env (foreground terminal commands)
- _sanitize_subprocess_env (background/PTY spawns)
- PTY env builder

This preserves env_passthrough semantics and never silently discards the
user's own PYTHONPATH configuration.
2026-08-16 23:26:04 -07:00
Teknium 66312aec48 Port from can1357/oh-my-pi#7553: allow quoted shell metacharacters in allowlist matching
command_allowlist glob rules (e.g. 'cargo *') rejected any command whose
quoted arguments contained shell metacharacters — a cargo benchmark
regex filter like '^layer3/write/(a|b)$' disqualified the whole command
even though those characters are literal to the shell.

_has_allowlist_shell_operator is now quote-aware:
- metacharacters inside single/double quotes or behind a backslash are
  treated as literal arguments;
- $ and backtick inside DOUBLE quotes still disqualify (expansion is
  active there);
- quoted/escaped control characters still disqualify when the command
  carries a -c/-e/--command/--eval-style option that hands the payload
  to another interpreter (sh -c '...', git -c alias.x='!...' x);
- unterminated quotes disqualify (shape can't be reasoned about).

Compound commands (unquoted ; & | < > backtick $( newline) are rejected
exactly as before. hermes_cli/approvals_suggest.derive_glob picks up the
same semantics via its existing import.
2026-08-16 22:10:13 -07:00
Teknium 2e4d771c69 Inspired by Factory Droid: accept unique ID prefixes in process tool lookups
Factory Droid v0.175.0 made TaskOutput/TaskStop accept task-ID prefixes so
background tasks can be referenced without pasting the full ID. Hermes'
process tool had the same friction: every action required the exact
proc_<12-hex> session ID.

ProcessRegistry.get() now falls back to unique-prefix resolution when the
exact lookup misses: 'proc_4dae' or bare '4dae' resolves to
proc_4dae56ca81f6 when exactly one running/finished session matches.
Ambiguous or too-short (<4 suffix chars) prefixes still return None, so
callers keep their existing 'No process with ID ...' error and nothing is
ever picked arbitrarily. Exact IDs never pay the scan, and a full ID that
happens to prefix another always wins.

All process actions (poll/log/wait/kill/write/submit/close) route through
get(), so they all gain prefix support from the single change.
2026-08-16 22:09:37 -07:00
Teknium ea29702749 feat(cron): --continuity / --no-continuity flags on hermes cron create/edit
CLI parity for the continuity toggle:

- subcommands/cron.py: --continuity on create; --continuity / --no-continuity
  tri-state pair on edit (same store_const pattern as --no-agent/--agent)
- cron.py: forwarded to the cronjob tool; created/edited job summaries print
  a "Continuity: on" line
- cronjob_tools._format_job: reports continuity as an explicit boolean and
  strips the reserved 'self' entry from the reported context_from list
- cron-job.ts: form reader accepts both shapes (raw store record with 'self'
  inside context_from, or formatted record with the explicit flag)
- docs: CLI flag examples in the continuity section

E2E (real argparse -> cron_create/cron_edit -> jobs.json in temp HERMES_HOME):
create --continuity stores ['self']; edit --no-continuity clears; edit
--continuity restores; default-off unchanged. 91 cron/tool tests + 16 CLI
cron tests + vitest 10/10 pass.
2026-08-16 22:09:28 -07:00
Teknium 2e7a46cc27 feat(cron): continuity=true/false flag as the user-facing surface for self-context
Per review: expose run-to-run continuity as a boolean `continuity` flag on
cronjob create/update instead of asking users to know the reserved
context_from='self' value. The flag translates to the 'self' entry in
context_from internally (create: appends/omits; update: adds or removes
'self' while preserving other upstream refs). Schema documents the flag and
steers context_from back to job-id chaining only. Docs updated; 7 new tests.
2026-08-16 22:09:28 -07:00
Teknium 47d7661aa8 Inspired by Amp: cron self-context — context_from='self' gives recurring jobs run-to-run continuity
Amp's 'Right on Schedule' (Jul 21 2026) lets scheduled agents wake up with
their saved context and continue where they left off. Hermes cron jobs run
in isolated sessions with per-run amnesia; the existing context_from chain
mechanism only referenced OTHER jobs. This adds the special value 'self'
(and treats a job's own literal id the same way): the job's most recent
output is injected with continuity framing so recurring scouts/monitors
dedupe against what they already reported and continue where they left off.

- cron/scheduler.py: resolve 'self'/own-id in _build_job_prompt with
  continuity framing instead of upstream-job framing
- tools/cronjob_tools.py: allow 'self' through create/update validation
  (can't be validated against the store — the job doesn't exist yet at
  create time); schema description documents the value
- tests: 6 new tests incl. sabotage-verified failures without the fix
- docs: self-context section in cron.md
2026-08-16 22:09:28 -07:00
Teknium bd4b709258 Inspired by Copilot CLI: /rollback keeps user hand-edits by default
Copilot CLI 1.0.78 reworked /rewind to restore only the files the agent
changed, 'skipping any file whose contents no longer match what Copilot
last wrote'. This ports that protection to Hermes checkpoints:

- tools/checkpoint_manager.py: per-project agent-write ledger
  (sha256 of every landed write_file/patch), safe_restore_plan()
  classifier, and restore(safe=True) that reverts only agent-authored
  changes, deletes agent-created files, and preserves user hand-edits.
  Empty ledger (pre-existing stores) falls back to the classic full
  restore.
- run_agent.py: feed the ledger from _record_file_mutation_result on
  every landed mutation (zero new hooks; rides the existing verifier).
- CLI + gateway /rollback: safe mode is the default; --all/--force
  restores everything; skipped files are reported with a hint.
- 17 locales: new gateway.rollback.kept_user_edits key.
- Docs: checkpoints-and-rollback.md updated.
- Tests: 7 new cases incl. user-edit preservation, post-agent user
  tweaks, agent-created file removal, empty-ledger fallback.
2026-08-16 22:08:47 -07:00
Teknium d44a295492 Inspired by Claude Cowork: security scanning for plugin install/update
Claude Cowork (Aug 6, 2026) added skill & plugin security scanning:
third-party skills and plugins are automatically checked for malicious
content on upload/edit, returning pass/warn/fail. Hermes already scans
hub-installed skills (tools/skills_guard.py), but `hermes plugins
install` cloned and activated arbitrary Git repos completely unscanned —
and plugins run Python in-process, making them the more dangerous
surface.

- tools/plugin_guard.py: plugin-adapted scanner reusing the skills_guard
  pattern engine. Exempts the documented provider-plugin patterns (own
  requires_env API-key reads, HTTP calls with keys) on code files while
  keeping true threat signals (foreign credential-store access, reverse
  shells, destructive/persistence/obfuscation patterns, prompt injection
  in docs). Plugin-sized structural limits; VCS/venv dirs excluded.
- hermes_cli/plugins_cmd.py: scan the temp clone before it is moved into
  ~/.hermes/plugins/. safe=install, caution=confirm (interactive prompt
  or --force), dangerous=blocked (--force does NOT override). Re-scan on
  `hermes plugins update`; a dangerous updated tree is deactivated until
  the user reviews the findings. Dashboard install path returns
  structured scan_blocked/scan_findings.
- Config gate: plugins.scan_on_install (default true) in config.yaml.
- Validated against all 60 bundled plugins: 57 safe, 3 caution (real
  sudo / curl|sh content in their docs), 0 false-positive blocks.
- 15 new tests incl. E2E through _install_plugin_core with real git
  clones.
2026-08-16 22:08:37 -07:00
Teknium 341d5aebc6 Port from MoonshotAI/kimi-code#2647: read UTF-16 text files by transcoding to UTF-8
UTF-16 text files (Windows Notepad .txt, PowerShell > redirects) were
refused as binary: the terminal env decodes stdout as UTF-8 with
errors=replace, so their content arrived mangled with U+FFFD and
tripped the binary guard.

ShellFileOperations.read_file now probes the raw bytes via the
backend's Python when the binary guard fires: a BOM or the zero-byte
parity heuristic (derived from VS Code's encoding sniffer, tolerant of
mixed Latin/CJK content) identifies UTF-16 LE/BE, and the file is
transcoded to UTF-8 with CRLF normalized and the BOM stripped. Real
binaries (zeros at both parities), binary extensions, files over
10 MiB, and legacy 8-bit encodings (GBK, Big5) still refuse — a wrong
silent guess is worse than a clear refusal. Works on every shell
backend (local/docker/ssh) since the probe runs via python3 -c.

Tests run against a real LocalEnvironment (E2E, no mocks); sabotage
run confirmed 6/9 fail without the fix.
2026-08-16 22:08:28 -07:00
Teknium c031fec365 Port from MoonshotAI/kimi-code#2596/#2600: surface MCP tool-result _meta to the model, minus protocol-reserved keys
MCP tool results carry a server _meta mapping (exposed as .meta by the
Python SDK) alongside structuredContent. Servers return namespaced
machine-readable contracts there (validated payloads, browser-handoff
URLs); Hermes previously dropped the field entirely, so that data was
invisible to the agent.

Now _meta is included in the JSON tool output, after filtering
protocol-reserved keys per the MCP spec's key-name rules: a prefix is
reserved when a modelcontextprotocol or mcp label is followed by at
least one more label (modelcontextprotocol.io/..., tools.mcp.com/...).
Vendor namespaces with a trailing reserved word (com.example.mcp/...)
and unprefixed keys pass through. Non-serializable metadata drops the
extras rather than failing the call.
2026-08-16 22:08:18 -07:00
Teknium 8bbda8ff33 Port from block/goose#10746: strip invisible Unicode TAG chars from MCP content
Unicode TAG characters (U+E0000-U+E007F) render as nothing in terminals
and chat UIs but are fully visible to LLM tokenizers, making them an
ASCII-smuggling prompt-injection channel for untrusted MCP servers.

- tools/ansi_strip.py: new strip_unicode_tags() with fast path; unlike
  goose we preserve valid emoji tag sequences (U+1F3F4 base + tag spec +
  U+E007F cancel), so regional flags survive.
- tools/mcp_tool.py: applied at every MCP text ingestion point — tool
  result text blocks, embedded resource text, read_resource contents,
  get_prompt message content, and tool descriptions entering the schema.
- tests/tools/test_unicode_tag_strip.py: smuggled-instruction vectors,
  goose's test vector, emoji-tag-sequence preservation, ZWJ untouched.
2026-08-16 22:08:05 -07:00
Teknium 08d9828503 feat(approval): coalesce identical concurrent gateway approval prompts
Port from anomalyco/opencode#40869: parallel tool calls hitting the same
dangerous-command gate each enqueued their own _ApprovalEntry and fired
their own notify_cb — the user got N identical prompts and had to
/approve N times while the agent sat wedged.

_await_gateway_decision now detects an already-pending identical
approval (same command text + pattern-key set) in the session queue and
waits on the leader's event via _await_coalesced_leader instead of
re-prompting. Followers adopt session/always (persistence would auto-pass
a re-check anyway) and deny/timeout (re-asking a just-declined command is
prompt spam); a single-use 'once' makes the follower issue a fresh
prompt. Pre/post approval hooks fire with coalesced=True for followers.
2026-08-16 22:07:33 -07:00
Teknium a35625d7c5 fix(search): zero-match probes return the file paths they found, not just counts
The casing/hidden/literal probes already ran the widened search to produce
their counts, then threw away the paths and returned a hint-only warning.
Strong models pivot in one turn; weak models spiral — the A/B eval measured
qwen3-coder-30b going 3.3 -> 9.3 turns on err_case_search, retrying casing
variants the probe had already resolved.

All three probes (case-insensitive, hidden/gitignored, literal-vs-regex) now
include up to 5 matched paths (+N more) in the warning via a shared tally
helper. Fixes the class, not the site.

Closes #80522
2026-08-16 22:06:50 -07:00
Teknium efe41abde0 feat(curator): per-mutation audit ledger + single-edit rollback
Tracker #79686 P3. Every skill mutation — curator, agent, or user — now
appends one entry to the append-only JSONL ledger at
~/.hermes/skills/.curator_ledger.jsonl, with per-file before/after
manifests whose contents are stored content-addressed (sha256-deduped)
under ~/.hermes/.curator_backups/blobs/.

- tools/skill_ledger.py: append/list/get, blob store, actor derivation
  (curator|agent|user), single-entry rollback that takes a pre-rollback
  safety entry first and FAILS CLOSED when that capture fails (consistent
  with the whole-run tarball rollback hardening from #63366). Path
  containment check so a hand-edited ledger can't write outside
  HERMES_HOME.
- Hooked all three choke points: skill_manage() dispatch (all actors,
  delete intent recorded via absorbed_into/archived evidence),
  archive_skill()/restore_skill(), and curator auto-transitions (tagged
  actor=curator via a ContextVar override).
- Ledger failures never block the mutation — telemetry, not a gate.
  Config gate skills.ledger (default true).
- hermes curator ledger [--skill NAME] [--limit N] and
  hermes curator rollback <entry-id> (whole-tree snapshot rollback
  unchanged).
- Optional TTL purge of skills/.archive/: curator.archive_ttl_days
  (default 0 = never) + explicit hermes curator purge, recorded in the
  ledger with before-blobs so purges stay recoverable.
- Docs: curator.md sections on the ledger, single-edit rollback, and
  archive TTL purge.

Curator invariants unchanged: only created_by:agent skills auto-transition,
never hard-delete autonomously, pinned exempt; foreground user deletes stay
hard-delete (and are now recoverable via the ledger).

Closes #45778, #50875. Tests adapted from #50261 by @yu-xin-c.
2026-08-16 22:06:41 -07:00
Teknium 184cddb449 fix(delegation): honor pinned delegation.provider — no silent parent-fallback substitution
When delegation.provider/model is explicitly pinned, the child no longer
inherits the parent's fallback chain: a mid-run auth/429 failure on the
pin previously rerouted the quiet-mode child onto parent fallback models
with no surfaced signal. Same treatment as the existing override_provider
OpenRouter filter-clearing — explicit pins are honored or fail loudly.

Also upgrades the pinned delegation.command-missing-from-PATH case from
warning + silent transport fallback to a loud spawn refusal, both at
credential preflight and in _build_child_agent.

Fixes #80450 (tracker #79686 audit item).
2026-08-16 22:06:32 -07:00
Teknium 204302bd64 feat(terminal): interpret signal-termination exit codes for the model
Port from Kilo-Org/kilocode#12698: report signal-terminated commands with
a human-readable note instead of a bare numeric exit code.

Kilo's fix settles a signal-killed process as the conventional 128+signum
exit code so its bash tool stops hanging. Hermes already produces numeric
codes for signal deaths (subprocess -signum, or the shell's 128+signum),
but the model saw a bare exit_code=-9 or 137 and burned turns
mis-diagnosing (137 = OOM kill being the most common). This adapts the
idea to Hermes' existing exit-code semantics tier:

- _interpret_signal_exit(): maps negative codes (definite signal death)
  and the 128+signum band (hedged with 'usually') to a note naming the
  signal and its likely cause, wired into _interpret_exit_code() ahead of
  the per-command semantics table.
- Curated signal table (SIGKILL/SIGSEGV/SIGTERM/SIGABRT/...) so ambiguous
  application exit codes are never mislabeled; uncurated 128+N codes stay
  silent, SIGINT is excluded (executor's interrupt-marker path owns
  rc=130).
- Notes surface via the existing exit_code_meaning result field.

E2E verified against real SIGSEGV/SIGKILL processes.
2026-08-16 22:05:59 -07:00
Teknium 3b9a963b8e refactor(xai): lift API-key precedence into resolve_xai_http_credentials behind prefer_api_key
Rework of the #88049 inline early-return per review:

- resolve_xai_http_credentials gains an opt-in prefer_api_key flag that
  checks the explicit XAI_API_KEY first and falls back to OAuth. The key
  is read through tools.tool_backend_helpers.resolve_provider_secret
  (config -> profile secret scope -> env/.env -> credential pool) so the
  preferred path enforces the same scope policy as the existing fallback
  branch, including failing closed under a multiplexed gateway turn.
- The preferred path's base URL honors HERMES_XAI_BASE_URL then
  XAI_BASE_URL behind hermes_cli.auth._xai_validate_inference_base_url,
  mirroring the OAuth branch (a foreign origin can't exfiltrate the key).
- x_search's _resolve_xai_bearer now calls the shared resolver with
  prefer_api_key=True instead of re-implementing precedence inline (#88040).
- tools/tts_tool.py _generate_xai_tts converted to the same flag — same
  root cause for /v1/tts 403s (#87045, supersedes the inline shape in
  #87081 by @enwaiax).
- Regression tests retargeted at the tools.xai_http.get_env_value seam and
  the shared resolver; added coverage for the flag's OAuth fallback,
  HERMES_XAI_BASE_URL + origin validation, default-order stability, and a
  profile-scope-only key on the preferred path.
- Docs: x-search authentication section now states the explicit API key
  wins (metered billing implication).
2026-08-16 21:20:29 -07:00
liuhao1024 32170dd1a2 fix(x_search): prefer an explicit API key over subscription OAuth
When a paid XAI_API_KEY is configured alongside xAI SuperGrok OAuth,
_resolve_xai_bearer() took the OAuth path unconditionally. The OAuth
credential authorizes /v1/responses but answers in a degraded Grok
explanatory mode with no citations, while the API key returns real
posts - so every x_search query silently degraded (#88040).

Prefer the explicit API key when set (same shape as the TTS fix for
#87045 in #87081), keeping OAuth as the fallback when no API key is
configured. The shared resolver and every other xAI call site are
unchanged.
2026-08-16 21:20:29 -07:00
webtecnica e3c71e052d feat(delegation): record model/provider in live-transcript manifest (#telemetry) 2026-08-16 20:14:51 -07:00
Teknium 9829064f89 fix: capability fingerprint reads config via the canonical loader
The config-read guard (test_config_read_guard) correctly flagged the
probe's raw yaml.safe_load of config.yaml — raw reads miss the managed
overlay, env expansion, and normalization. Use load_config_readonly()
under a scoped HERMES_HOME override instead. E2E v3/v3b and the guard
both green.
2026-08-16 18:30:53 -07:00
Teknium 4e22d070f4 feat(agent): one-time protocol upgrade for legacy Bot Chat sessions
Bot Chats created before the epoch mechanism persisted prompts with no
protocol section and no stamp — the staleness check only fires on
stamped prompts, so pre-existing bots would never learn to message
teammates. stored_bot_chat_prompt_needs_upgrade() migrates them: one
rebuild, title-gated to Bot Chat, only when the probe would actually
emit a section (SOUL-append legacies and unmanaged installs are left
alone — rebuilding those would loop). The rebuilt prompt carries the
stamp, so the upgrade can never re-fire.

E2E v3b through the real restore path: legacy Bot Chat upgraded once
then verbatim-reused; legacy regular sessions byte-untouched.
tests/agent/ 4648/4648.
2026-08-16 18:30:53 -07:00
Teknium ea4310e76c feat(agent): capability-refresh + timeless prompts for eternal Bot Chat sessions
Bot Chats break the "new sessions come often" assumption behind
build-once system prompts: capability edits used to sit invisible until
/new or compression, and the frozen birth date became misinformation.

- tools/bot_mode_probe.py: capability_fingerprint() hashes the profile's
  capability surface (disabled skills, toolset pins, MCP config, SOUL.md,
  installed skills, Bot-Mode roster); Bot Chat prompts embed the 12-hex
  epoch stamp
- agent/conversation_loop.py restore path: stored Bot Chat prompt whose
  epoch mismatches disk → ONE rebuild (through a cleared skills-prompt
  cache so new installs appear), persisted so the next turn reuses the
  new bytes verbatim. Prompts without a stamp — every non-Bot-Chat
  session — never take the branch; probe failure fails closed to reuse
- agent/system_prompt.py: Bot Chat prompts are timeless — the
  "Conversation started:" date is dropped (timezone kept); no ticking
  fields in an eternal session
- tui_gateway: _sync_bot_capabilities at turn start rebuilds the live
  agent (tool definitions are construction-baked) when the fingerprint
  moves, same session id/history, with a user-visible notice

Cache stance: this is the /model exception applied to capabilities — a
loud, user-initiated, once-per-change prefix break. Unchanged state
hashes identically and stored bytes are reused verbatim (E2E-proven).

Validation: 9 probe unit tests incl. per-axis fingerprint changes;
E2E v3 against the real restore path (fresh build → verbatim reuse →
skill install → single refresh w/ new skill in index → verbatim reuse;
regular sessions dated, unstamped, never refreshed); tests/agent/
4647/4647.
2026-08-16 18:30:53 -07:00
Teknium 78a4693eef fix(agent): scope the Bot Mode protocol section to canonical Bot Chat sessions
Per review: the protocol belongs only in official Bot Mode interactions,
not every session on a managed install. The prompt builder now injects
the section only when the agent's session row is titled "Bot Chat"
(BOT_CHAT_TITLE, matching the desktop's createCanonicalChat pin and the
`hermes -p <bot> chat -c "Bot Chat"` resume target). Regular sessions
never carry it; the desktop composer middleware owns @mention sends.

Title is read once at first prompt build and the rendered prompt is
cached + DB-restored — cache-safe. E2E against the real AIAgent +
SessionDB: absent in an untitled session, present in Bot Chat,
byte-stable across rebuilds, absent after retitle, absent with the
flag off. Overhead unchanged (~916B, Bot Chat sessions only).
2026-08-16 18:30:53 -07:00
Teknium 2b39e92e9e feat(agent): core Bot Mode teammate protocol — stable-tier prompt section
Replaces the plugin-side SOUL.md protocol append: on Bot-Mode-managed
installs (any profile carrying ui_meta['hermes-bots']) the prompt builder
injects the "Messaging other agents" section into every session of every
profile — including headless `hermes -p <bot> chat` sessions a teammate
starts — so bot handoffs work without mutating user-authored SOUL files.

- tools/bot_mode_probe.py: silent-when-unmanaged probe, cached per
  (process, home), keyed off the agent's OWN home (not ambient
  HERMES_HOME); silent when SOUL.md already carries the legacy section
- agent/system_prompt.py + agent_init.py + config_defaults.py: wired as
  agent.bot_mode_protocol (default True), stable tier, byte-stable
  across rebuilds (E2E-verified against the real build_system_prompt)
- tui_gateway profiles.list gains bot_mode_protocol capability flag;
  the bundled plugin gates ALL SOUL protocol writes on it (backfill,
  composeSoul, Edit save) — older gateways keep the SOUL-append path
- overhead: ~916 bytes, only on Bot-Mode installs; zero elsewhere

Supersedes the SOUL backfill half of Hermes-Bot-Mode#99 (credit
@kaduxo — the handle fix, `hermes profile list` correction, and
idempotent-append guards from that PR ship in the bundled plugin).
2026-08-16 18:30:53 -07:00
Teknium 00c12dac61 fix(computer-use): auto-repair an installed driver that fails the runtime contract
A same-day version-floor bump (0.20 runtime contract) left every install
with an older cua-driver hard-failing on all computer_use calls: the
start() gate fails closed, while the `hermes update` refresh defers to the
driver's own check-update verb — whose ~20h cache routinely answers "no
update available" right after we raise the floor. Hermes knew it required
0.20+ but never acted on that knowledge.

Two changes:

- tools_config.install_cua_driver(): a contract-failed installed driver is
  repaired on the upgrade=True path too (previously only upgrade=False).
  The contract failure itself is the confirmation, so the
  require_confirmed_update gate and the check-update short-circuit are
  bypassed for repairs — an indeterminate or stale-cached check can no
  longer pin users on an unusable driver.

- cua_backend.CuaDriverBackend.start(): when the contract gate fails on an
  installed binary, attempt one automatic repair per process via the
  standard install path, then re-probe. HERMES_CUA_DRIVER_CMD overrides
  are never repaired (explicit override is authoritative even when broken)
  and a missing binary still just reports the install hint. A failing
  installer can't loop: the second start() surfaces the original error.

Tests: contract-repair coverage in test_computer_use.py (auto-repair
success, failed repair surfaces the original error, once-per-process
guard, override never repaired, missing binary never repaired) and
test_install_cua_driver.py (incompatible driver repairs despite an
indeterminate check-update, check-update not consulted). All new tests
verified to fail against the unfixed source (sabotage run).
2026-08-16 13:01:20 -07:00
f-trycua 12b1f0f83d fix(computer-use): align browser guidance and screenshots 2026-08-16 11:34:40 -07:00
Francesco Bonacci 3e0087abe2 fix(computer-use): warn when an approval bypass widens the driver mode
`--yolo` / `-z` read as "don't prompt me", but they also swap computer_use
onto a private `unrestricted` daemon, dropping the ceilings the configured
mode would have applied. Nothing said so. A script picks up `-z` for quiet
output and loses its limits as a side effect, and the only trace is a driver
process nobody inspects.

The mapping itself stays. It is deliberate, and `unrestricted` is reachable
no other way: it is intentionally not a config value so a stale config line
can never silently bypass approvals (see `_cua_configured_permission_mode`).
Removing the mapping would delete the capability rather than fix it, and
splitting it onto a second CLI flag was declined to avoid growing the
surface.

So the widening is now stated instead: one warning per session naming the
configured mode it left, what stopped applying, and the two ways to keep a
ceiling - drop the bypass flag, or declare a version-3 capability manifest,
which now rides along with unrestricted as of the previous commit.

Once per session, not per dispatch: the resolver runs on every tool call.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 11:34:40 -07:00
Francesco Bonacci cdbea83e1c fix(computer-use): keep a v3 capability manifest on approval-bypassed runs
`--yolo` / `-z` route the session onto a private embedded daemon in
`unrestricted` mode. That daemon was constructed without the configured
capability manifest, and the serve command only attached
`--capability-manifest` when the mode was exactly `bounded`. So the moment a
run was bypassed, the user's declared ceiling was dropped:

    without -z:  --permission-mode bounded --capability-manifest ...
    with -z:     --permission-mode unrestricted --dangerously-bypass-approvals

No manifest, no warning. The most carefully configured run - a reviewed
ceiling, written by hand - became the least constrained one, silently, and
it failed open.

That was never a driver limitation. cua-driver documents the manifest as a
ceiling across modes ("A manifest can narrow a profile but never widen it";
its own authorization table calls it `optional_capability_manifest_ceiling`),
and accepts it alongside `--permission-mode unrestricted`.

The forwarding is version-aware, because the two manifest schemas differ
(cua-driver session_manifest.rs):

* v1/v2 are legacy and must declare `mode: bounded`. Handing one to an
  unrestricted runtime aborts startup with "legacy capability manifest mode
  must be bounded", so a naive forward would turn a working session into a
  hard failure. These are forwarded for bounded only, and a warning names
  the migration when one cannot apply.
* v3 must not declare a mode. It is the mode-independent ceiling, and it now
  rides along with unrestricted.

Unreadable or unparseable manifests are not forwarded outside bounded, on
the same fail-safe reasoning; bounded still forwards unconditionally and
lets the driver be the authority there.

Verified against cua-driver 0.20.0 on Windows. Launch args now carry
`--permission-mode unrestricted --dangerously-bypass-approvals
--capability-manifest <v3> --approve-capability-manifest`, and the ceiling
is enforced in the bypassed run - a tool outside the manifest is refused
("outside the capability manifest for this session ... blocked as a
protected resource") where the same config previously ran unbounded. A
legacy manifest was confirmed to abort driver startup when forwarded, which
is what the version gate prevents.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 11:34:40 -07:00
Francesco Bonacci 5f049b517b fix(computer-use): make the typed-browser bind/snapshot split discoverable
`cua_browser_state` has two branches, chosen implicitly: any call carrying
pid or window_id is a *binding* (browser_route.py:252), anything else is a
*snapshot*. A binding clears session state, mints fresh tab_ids, returns
binding metadata with no page content, and sets verification_required.

Nothing in the response says that. A caller that keeps passing pid/window_id
- the natural reading of "bind to this window, then read it" - re-binds
forever: the tab_id it just received is unbound by the next bind, so every
cua_browser_navigate comes back browser_verification_required, and the
refusal ("take a fresh snapshot") points at the same call that just re-bound.
Observed live as 11 consecutive refused navigates before the model gave up
and fell back to foreground SendInput on the address bar.

The same confusion silently swallowed include_screenshot: both calls that
requested one were bindings, which carry no page content, so the flag had
nothing to attach to and was dropped without comment.

A binding response now reports snapshot_required, next_step
(fresh_browser_state, matching the existing token convention) and a hint
naming the exact next call; requesting a screenshot on a binding reports
screenshot_deferred instead of dropping it. The verification refusal now
says to call cua_browser_state WITHOUT pid/window_id and why re-sending them
does not help. The schema documents that include_screenshot applies to
snapshots.

Behavior of the bind and snapshot branches themselves is unchanged - this is
purely about making the split legible to the caller.

Unit-tested. Not verified end to end on the reporting host: the driver
refuses the bind upstream there (`browser_requires_setup: no owned DevTools
endpoint`, and it does not accept a user-launched --remote-debugging-port),
so the typed route never reaches this branch. That attach failure is a
separate cua-driver issue.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 11:34:40 -07:00
Francesco Bonacci 9a96fdc5b8 fix(computer-use): enforce existing-profile grant, unblock the opt-in
Live-testing the Cua Driver 0.20 convergence on Windows 11 (session 2,
cua-driver 0.20.0) surfaced three defects in the existing-profile browser
path and in install status.

1. The config grant was silently nullified by an approval bypass.

`--yolo` / `-z` map onto a private unrestricted daemon, which answers every
browser_prepare. Because the host delegated the entire existing-profile
decision to the driver, that bypass also nullified
`computer_use.grant_existing_profile: false`: a plain `hermes -z` attached
to the user's real Chrome profile and read live page content over CDP, with
the driver reporting it as "the approved existing Chromium profile". It was
never approved.

An approval bypass is consent to skip prompts, not consent to read an
existing profile's pages, cookies, and storage. CuaTypedBrowserRoute.prepare
now enforces the key itself, regardless of permission mode. bounded stays
exempt - its reviewed capability manifest is the authorization boundary.
The authorization inputs are resolved in the backend from config and the
backend's immutable mode, never from model-supplied kwargs.

2. The grant, once set, still could not be used.

With `grant_existing_profile: true` the runtime is launched
`--grant existing-profile` correctly, but cua_browser_prepare then hit a
runtime approval prompt anyway - re-asking the user to authorize what the
config already authorized, and making the documented opt-in unusable on any
non-interactive run, where the prompt has nobody to answer it and the call
dies on approval timeout. The durable, file-backed grant now stands in for
that prompt. Scope is narrow: only the existing-profile prepare, only when
the grant is present; isolated launches still prompt and any resolution
failure falls closed to prompting.

3. `computer-use status` hid a custom override and spliced its output.

With HERMES_CUA_DRIVER_CMD pointed at cmd.exe, status printed the child's
multi-line banner and prompt inside the one-line version field, never
mentioned the override, and advised `hermes computer-use install` - which
install itself (correctly) refuses to run against an overridden path. It now
names the override and mirrors install's update-or-unset guidance, and
version output is reduced to one bounded line.

Verified on the reported host: `-z` existing-profile attach now refuses and
names the key; `grant: true` no longer prompts (33s vs a 300s approval
timeout); status names the override and prints one line. No change to the
reconciliation path - driver SHA256 unchanged end to end.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 11:34:40 -07:00
Francesco Bonacci 81af2ef013 fix(computer-use): reconcile existing cua-driver installs 2026-08-16 11:34:40 -07:00
Francesco Bonacci a403fe6f92 feat(computer-use): support Cua Driver 0.20 runtime contracts 2026-08-16 11:34:40 -07:00
Teknium c257e9196b fix: make every tool interruptible — sequential executor abandons on user interrupt
The sequential tool path only noticed a user interrupt after the running
tool returned: with the deadline disabled it ran the tool inline (fully
blocking), and with a deadline it waited in 5s slices without ever
checking agent._interrupt_requested. Any tool without cooperative
is_interrupted() polling (image_generate, tts, transcription, skills
sync, ...) held the whole turn hostage — the reported symptom was a
redirect queued ~40s behind a FAL image generation + upscale pass.

Executor backstop (class fix, covers ALL tools):
- _run_sequential_tool_execution_middleware always dispatches on the
  daemon worker (timeout None no longer means inline blocking) and polls
  the interrupt flag every 1s.
- On interrupt: 3s cooperative grace (mirrors the concurrent path), then
  synthesize a cancelled tool result (_ToolCancelledResult), emit the
  terminal post_tool_call with status=cancelled, and abandon the worker.
- _ToolCancelledResult suppresses downstream post-hook double emission
  exactly like _ToolTimeoutResult, so an abandoned worker finishing late
  cannot report success for a cancelled call.
- clarify (interactive, _NEVER_PARALLEL_TOOLS) keeps the inline path —
  it owns its own human wait.

Cooperative layer in the reported offender:
- image_generation_tool: blind handler.get() (generation + Clarity
  upscale) replaced with _wait_fal_result(), which polls is_interrupted()
  in 0.5s slices and raises ImageGenerationInterrupted immediately.
- _upscale_image propagates the interrupt instead of swallowing it into
  the "upscale failed, use original" fallback.

Message alternation is preserved: the cancelled result is a normal tool
result for the call_id. Sabotage-verified: with the old wait loop
restored, the new tests fail (tool blocks full runtime); with the fix
they pass in ~4s.
2026-08-16 11:32:02 -07:00
Teknium f06c41522e fix(image_gen): disable default-on upscaling everywhere — opt-in only
The Aug 8 default-on upscaling policy (66ea4e686) chained the Clarity
Upscaler after every sub-2MP generation. Clarity is an SD1.5 creative
tile-diffusion enhancer (creativity 0.35, "masterpiece" prompt prefix) —
it redraws content, which degraded output on 100% of generations for
models like GPT Image 2 and Ideogram whose value is precise text
rendering, CJK, and photorealistic detail.

Policy now: no model upscales by default, on FAL or Krea. The `upscale`
tool param remains as a per-call opt-in (`upscale: true`); explicit
requests still chain Clarity (FAL) / Krea Enhance as before.

- FAL catalog: all 17 default-on entries flipped to upscale=False
- Krea plugin: medium + medium-turbo per-model defaults flipped off
- Tool schema: upscale param described as opt-in with a fidelity warning
- Tests updated: catalog invariant now pins all-off; default-on cases
  now assert no upscaler call
- Docs (en + zh) updated to the opt-in policy
2026-08-16 11:12:05 -07:00