Phase 2 of the MCP 2026-07-28 migration (#69931), on top of the SDK 2.x
migration (#88180):
- Protocol-era negotiation (_negotiate_session): per-server `protocol`
config key — auto (default, handshake-first with server/discover
fallback on -32022/-32601), stateless (discover-first), legacy
(handshake only). Auto is handshake-first deliberately: zero extra
round-trips and zero behavior change for the entire existing server
fleet, while 2026-07-28-only servers now connect via the fallback.
All four transport call sites (stdio, SSE, new HTTP, legacy HTTP)
route through the one choke point, so the CLI/desktop probe path
inherits it too.
- SEP-2549 list caching: tools/list ttlMs/cacheScope hints are captured
during discovery and bound to the lazy-startup schema cache — TTL'd
entries expire and force a live re-probe; hint-less (pre-2026)
servers keep the never-expires behavior. Pagination continuation now
speaks both SDK generations (params= vs cursor=).
- SEP-837: OAuth client metadata declares application_type=native
(config-overridable), with a fallback for 1.x-era metadata models.
(RFC 9207 iss validation and SEP-2352 issuer-keyed credentials are
native to SDK 2.0's OAuthClientProvider — verified, no client-side
gap.)
- SEP-2577 deprecation posture: SamplingHandler docstring marks the
Sampling feature as upstream-deprecated (12-month window) — kept
fully functional, closed to new capability.
- Docs: `protocol` key in the MCP config reference.
Shorter, single-source ownership explanation for
_strip_hermes_owned_pythonpath (the code-level Check comments already
carry the per-branch detail; the docstring only needs the contract).
Same behavior, same coverage, less boilerplate (test file 1691 -> 1512
lines; PR diff unchanged in semantics).
Production (mechanical only):
- Extract _strip_hermes_owned_pythonpath_and_runtime_markers(): the three
builders (_make_run_env, _sanitize_subprocess_env, hermes_subprocess_env)
ran the identical strip-then-pop-markers sequence in the same order
(ordering is load-bearing for VIRTUAL_ENV validation); the helper makes
that explicit once instead of three times.
Tests:
- Non-owned preservation: 11 single-shape tests -> one parametrized matrix
(user/Nix/other-version/python2.7/pythonX.Y-contained/raw spelling/empty
component/empty PYTHONPATH) + one runtime-shaped matrix (other-version SP,
venv-SP descendant, repo direct child, repo deep child).
- Owned stripping: venv SP, repo root (independent parents[2] computation),
duplicates, all-owned key removal, mixed ordering -> one matrix.
- Builder integration: _make_run_env/_sanitize_subprocess_env/
hermes_subprocess_env venv-SP stripping -> one parametrized test;
same for the four PYTHONHOME builders (incl. build_subprocess_env).
- Junction: same-named non-owned negative control now covers both the
configured-root location and an unrelated location; shared
_physical_repo_root helper; profile resolution matrix (root->named,
profile-shaped->named no nesting, profile-shaped->default, custom root).
- Every independent proof preserved: home-level junction, repo-level
junction, profile interaction, negative identity control, uv-base lexical
VIRTUAL_ENV, validated/unrelated VIRTUAL_ENV, no-scrub escape hatch,
#84500 same-env/external-env composition, PYTHONHOME removal, real
Windows-only semantics, POSIX fail-closed backslash paths.
Second real-world topology reported and confirmed on native Windows 11:
the repository itself is a cross-drive junction (D:\hermes\hermes-agent ->
C:\...\hermes-agent) under a real HERMES_HOME directory. The editable
import spelling resolves to the physical location, so _hermes_repo_root is
physical while the launcher writes the lexical spelling into PYTHONPATH.
The home-relative mapping cannot express a cross-drive link (commonpath
raises on different drives), so the lexical repo root survives stripping;
and with the repo alias missing, a lexical VIRTUAL_ENV
(D:\hermes\hermes-agent\venv) also fails _validated_runtime_venv, so the
venv site-packages survives too (uv-base gateway: both entries survive).
Fix: after the existing home/profile-root mapping, try the single
deterministic candidate <lexical root>/<repo dirname> for every trusted
home candidate (configured home, plus the profile root when the configured
home is a profile path) and accept it only when strict resolve proves it is
the exact physical repo root (fail-closed: missing paths, real directories
that are not the known repo, and unrelated spellings are never aliased).
This also re-enables the VIRTUAL_ENV validation for lexical venv spellings,
so uv-base gateway site-packages cleanup follows the repo alias.
Tests: repo-level junction positive + negative control (same-named real
directory preserved), profile-home + repo-level junction combination,
lexical VIRTUAL_ENV validation after recovery (root + site-packages
stripped, user entries kept), and a no-provenance lookalike preserved.
The execute_code composition test now compares composed paths with
os.path.normcase so a Windows case-only spelling difference (resolve() vs
abspath() casing) can never fail the composition contract.
Confirmed on native Windows 11 with a real junction and the real startup
chain: when the desktop/CLI spawns the backend with HERMES_HOME in the
configured (lexical) spelling and --profile / sticky active_profile is in
play, _apply_profile_override() re-homes HERMES_HOME through
resolve_profile_env(), which resolves the junction under the platform
default and returns the PHYSICAL spelling. tools.environments.local is
imported after that mutation, so _hermes_repo_root_aliases is built from
the physical home, the lexical repo-root spelling written into PYTHONPATH
by the launcher (D:\hermes\hermes-agent) is not derivable, and the entry
survives stripping (reproduced: cases --profile default / named / sticky
active_profile / cross-drive junction all leave it in place; no-profile
strips it).
Two narrow changes, no heuristics, no new env vars:
- hermes_cli/profiles.py::resolve_profile_env: when HERMES_HOME is set,
the configured spelling IS the launch root (junction-transparent,
physically identical dirs); keep it instead of re-deriving the native
default. This is the same producer contract _preserve_hermes_home_path
already follows.
- tools/environments/local.py::_build_hermes_repo_root_aliases: when the
configured home is a profile home (<root>/profiles/<name>), also derive
the root spelling lexically (parent of the profiles component, same
rule get_default_hermes_root uses) and run the exact-ownership mapping
against it, so the launcher's lexical root is recovered after re-home
without ever matching arbitrary descendants of HERMES_HOME.
Regression test test_profile_rehome_keeps_junction_lexical_alias covers
junction + profile re-home + inherited lexical PYTHONPATH end to end.
The PYTHONPATH/PATH sanitization suite was written POSIX-centric and
failed on real Windows 11 (reproduced natively: 4 failures before this
change). Fix the tests to express the true per-platform contract:
- test_other_major_version_site_packages_preserved /
test_make_run_env_injects_hermes_bin_dir: build inputs with
os.pathsep instead of hardcoded ':'.
- test_make_run_env_appends_homebrew_on_minimal_path: split on
os.pathsep, neutralise Git Bash dir prepending, and assert the
documented Windows passthrough (_append_missing_sane_path_entries is
a no-op off POSIX) instead of the Homebrew append.
- test_make_run_env_real_launchd_path_gains_homebrew: mark
macos_only per repo OS-marker policy (the regression is the macOS
launchd PATH; the merge is a passthrough on Windows).
- test_configured_home_alias_matches_launcher_output: create the
configured-home link via a helper that falls back to an unprivileged
directory junction (cmd /c mklink /J) when symlink creation raises
WinError 1314, and skips with a clear reason if no mechanism exists.
Also correct a stale comment in execute_code: the child is not always
the same Python as Hermes (project mode can select an external venv),
so the strip is about compatibility, not redundancy.
Adversarial review of the previous two commits (and #78917 itself)
found three ownership-boundary issues; this commit addresses them:
1. Repo direct-child over-strip (Finding A)
No launcher injects <repo>/tools or another direct child as an
independent PYTHONPATH entry - audited all four producers (Electron
electron-main.mjs, gateway/run.py::_ensure_windows_gateway_venv_imports,
cron/scheduler.py::_windows_cron_python_invocation,
tui_gateway/host_supervisor.py). The depth<=1 rule deleted user paths
that merely live under the repo directory; only the EXACT repo root is
now stripped.
2. Windows junction/symlink alias (Finding B)
The gateway launcher renders Hermes-owned paths under the configured
HERMES_HOME spelling (gateway_windows.py::_preserve_hermes_home_path),
which may be a junction to another drive, so it differs lexically from
the resolved repo root. _hermes_repo_root_aliases now carries both the
resolved and unresolved spellings; both are recognized as Hermes-owned.
3. Stale abstraction rename (Phase 4)
_strip_mismatched_site_packages -> _strip_hermes_owned_pythonpath:
the cross-version heuristic is gone, so the old name misdescribes the
behavior (ownership-based, not version-based).
Tests: direct-child now preserved; junction alias stripped (lexical pair
monkeypatched); Windows-only real-semantics test added (POSIX test remains
a safety test); mixed-ordering, duplicate-Hermes, and no-scrub PYTHONHOME
contract tests added. Full file: 52 passed / 16 failed (identical failure
set to base, all isolation-venv environment issues).
The gateway runs inside its own venv; if its PYTHONHOME leaks into
subprocesses (terminal commands, cron no_agent scripts, TTS providers),
any child interpreter redirects its stdlib search to the Hermes venv and
crashes with version-mismatch errors before importing anything.
PYTHONHOME is now part of _ACTIVE_VENV_MARKER_VARS so all env builders
(_make_run_env, _sanitize_subprocess_env, hermes_subprocess_env, and
build_subprocess_env used by cron) drop it, consistent with Hermes'
existing PYTHONHOME handling in managed_uv.py and sqlite_runtime.py.
execute_code already scrubbed it via _SAFE_ENV_PREFIXES.
Tests cover all four builders plus the marker constant.
Remove the cross-version heuristic from _strip_mismatched_site_packages:
the subprocess env builder cannot know which Python version a child will
run, so judging user PYTHONPATH entries against the backend interpreter's
version deletes legitimate paths meant for a different child Python
(e.g. /custom/lib/python3.13/site-packages while Hermes runs 3.11).
Also fix over-strip: entries merely containing a pythonX.Y path component
(e.g. /opt/tools/python3.13/bin) were stripped even though they are not
site-packages. Hermes-owned entries (repo root, own venv site-packages)
are now identified by path ownership, not by version.
Regression tests cover both cases; user paths with any pythonX.Y
component are preserved.
`ElicitationHandler` read `params.requested_schema`, but on the pinned
`mcp==1.28.1` the model field is spelled `requestedSchema`. The getattr
always missed and returned its `{}` default, so
`_format_elicitation_schema_summary` took its no-properties branch and the
approval prompt collapsed to the generic
Approval requested by MCP server '<name>'.
for every request. The field names, types, and descriptions the summary
exists to surface never reached the user, so an elicitation asking for a
card number rendered identically to one asking for a nickname — consent
without the substance of what was being consented to.
Read both spellings rather than just correcting to the 1.x name: mcp 2.0
renames this field to `requested_schema` (it renamed every model field to
snake_case and kept camelCase only as a serialization alias, which
pydantic does not expose to attribute access), so a dual read is correct
on either SDK generation and does not go wrong again on the next bump.
Verified against real 1.28.1 and 2.0.0 installs.
Every existing test in tests/tools/test_mcp_elicitation.py builds a
duck-typed `SimpleNamespace` stand-in, which carries whatever field name
the test wrote and therefore cannot detect a mismatch with the real model.
Add one test that constructs the actual `ElicitRequestFormParams` and
asserts the requested field name reaches the consent description; it fails
on the unfixed tree. The cheap stand-ins are left alone elsewhere.
Found while porting the tree to the mcp 2.x SDK in #76736, but independent
of it: this reproduces on the current pin with no other changes, #76736
does not touch this line, and the two branches merge cleanly in either
order.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The teammate-messaging protocol told bots to send with
`chat -c "Bot Chat" -Q -q` and, on 'No session found', fall back to a
manual two-step (send without -c, then sessions rename) — a dance the
CLI already made unnecessary when --create-if-missing landed (#86794).
Profiles that never went through the Bots-panel birth flow (CLI-created,
pre-Bot-Mode, remote-source) hit that miss on every first contact, and
background sends swallowed the error entirely (the original silent-drop
in Hermes-Bot-Mode#48 / #88059).
- tools/bot_mode_probe.py: protocol command gains --create-if-missing
- bundled plugin: Bot Chat prompt section + @mention handoff note use
the flag; rename-dance instructions deleted
The capability epoch hashes the protocol section, so existing eternal
Bot Chat sessions pick the new instructions up on their next message via
the established once-per-change rebuild — no per-turn cache drift.
Live-verified: fresh profile with zero sessions, protocol command
created 'Bot Chat' and delivered (PONG round-trip); second send resolved
the same session by title (no duplicate); missing-title send WITHOUT the
flag still errors loudly on stderr.
The HTTP transport seeded `MCP-Protocol-Version` from LATEST_PROTOCOL_VERSION,
which on mcp 2.x is 2026-07-28 — a revision that replaced the `initialize`
handshake with a per-request envelope. But this transport connects through
`ClientSession.initialize()`, which sends LATEST_HANDSHAKE_VERSION (2025-11-25)
in the body. Header and body therefore disagreed by construction, and a
conforming 2.x server honours the header: it routed the request onto its
per-request-envelope ladder and rejected the legacy body with
params._meta is missing the required envelope key(s):
io.modelcontextprotocol/protocolVersion,
io.modelcontextprotocol/clientCapabilities
Observed against a live MCP endpoint, and confirmed by probing the same
endpoint three ways: the header at 2026-07-28 is rejected, at 2025-11-25 it
succeeds, and with no header at all it succeeds.
Third defect in this migration from one cause: the 2.x bump changed what an
existing constant *means* without revisiting its uses. The header seed was
written when LATEST_PROTOCOL_VERSION was 2025-03-26 and was correct then.
Seeded from LATEST_HANDSHAKE_VERSION, imported with a fallback to
LATEST_PROTOCOL_VERSION for SDKs predating the split, where the two are the
same thing and header and body agree either way. An explicitly configured
header still wins — that override is why servers demanding a specific revision
can have one, and a test pins it.
`streamable_http_client` yields `(read, write, get_session_id)` on mcp 1.x and
`(read, write)` on 2.x. `_run_http` unpacked a fixed 3-tuple, so on 2.x every
HTTP and SSE MCP server failed its handshake with `ValueError: not enough
values to unpack (expected 3, got 2)` and parked after exhausting its retry
ladder. Only stdio servers kept working.
This is the same defect as the import gating fixed earlier in this branch, one
layer further in. That fix's own comment claimed reaching
`streamable_http_client` was "the path that does work" — reaching it was
necessary and not sufficient, and the comment asserted the half that was never
exercised. Corrected along with the code.
Unpacked positionally rather than by arity, since this file deliberately
supports both SDK generations and `get_session_id` was never used here.
The reason this survived review is worth the test it now has: the existing
coverage in test_mcp_client_cert.py fakes the transport with a 3-tuple, so it
encoded 1.x's shape into the assertion and passed on 2.x regardless. The new
test drives `_run_http` once per arity the supported SDK range actually yields,
and asserts the streams handed to ClientSession are the first two — positional,
because 1.x's third element is not a stream. Verified it fails on the 2.x case
without this change.
Found while pointing a real HTTP MCP server at a live deployment running this
branch: the server parked at startup and no tool from it ever registered.
mcp 2.0.0 implements MCP revision 2026-07-28 and makes three breaking
changes Hermes sits on top of: `mcp.server.fastmcp` is gone, every model
field is renamed to snake_case (camelCase survives only as a
serialization alias, which pydantic does not expose to attribute
access), and the SDK's own HTTP stack moved from `httpx` to `httpx2`.
Bump the pin across the dev/mcp/computer-use extras and port the tree:
- `mcp_serve.py` and `agent/transports/hermes_tools_mcp_server.py` move
from `FastMCP` to `mcp.server.MCPServer`, which has the same
decorator/add_tool surface. The hermes-tools server already
synthesised `__signature__` from Hermes' JSON Schema, which is exactly
what 2.0's `add_tool` reads.
- SDK model reads go through `mcp_field(obj, snake, camel)`, which reads
both spellings. A single-spelling read fails *silently* on the other
generation — empty tool schemas, dropped structured content, tool
results vanishing from sampling conversations — and `mcp` is an
optional extra users install at their own version.
- `sdk_httpx()` resolves the httpx flavour from the SDK's own transport
module, so objects handed to `streamable_http_client`, the `sse_client`
factory, and the OAuth metadata helpers come from the module the
installed SDK actually imports.
- HTTP support is gated on either streamable-HTTP entry point, not just
the deprecated alias 2.0 removed.
- OAuth: `OAuthClientProvider` lost its `timeout` argument (the
configured `oauth.timeout` now bounds the callback waiter's own poll
loop, where the browser round-trip was always awaited), and
`callback_handler` must return `AuthorizationCodeResult` rather than a
tuple. 2.0 also validates the RFC 9207 `iss` parameter, so the
callback handler and paste fallback capture it.
`mcp`/`mcp-types` 2.0.0 are inside the 14-day `exclude-newer` window, so
two narrow `exclude-newer-package` entries unblock `uv lock`, annotated
for removal on or after 2026-08-11. `httpx2` needs no exemption: 2.7.0 is
already outside the window and satisfies mcp's floor.
Refs #69931
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
_hermes_repo_root used parents[1] which resolves to tools/ instead of
the repository root. The file lives at tools/environments/local.py so it
needs parents[2] to reach the actual repo root that Electron injects
into PYTHONPATH.
The test test_repo_root_stripped reused the module constant under test
as its input. This made it pass regardless of what the constant pointed
at. The test now computes the real repo root independently from the
source file location. It fails with the old parents[1] code and passes
with the fix.
Reported by spfcraze in PR #78917 review.
The Desktop Electron process injects the Hermes venv's site-packages path
(e.g. .../python3.11/site-packages) into PYTHONPATH so the Python 3.11
backend can import its packages. When this PYTHONPATH leaks into terminal
subprocesses running a different Python version (e.g. Python 3.13), 3.11
C extension modules appear on sys.path ahead of the correct 3.13 versions
and crash with ImportError (PIL _imaging, cryptography, etc.).
Replace the existing blunt pop of PYTHONPATH from _ACTIVE_VENV_MARKER_VARS
with a surgical Hermes-venv-aware filter:
- Parse each PYTHONPATH entry by path
- Strip only paths under ~/.hermes/hermes-agent/venv/.../site-packages
- Preserve the Hermes source root (needed for import hermes_cli)
- Preserve all user-set PYTHONPATH entries
The same filter is applied in all three env builders:
- _make_run_env (foreground terminal commands)
- _sanitize_subprocess_env (background/PTY spawns)
- PTY env builder
This preserves env_passthrough semantics and never silently discards the
user's own PYTHONPATH configuration.
command_allowlist glob rules (e.g. 'cargo *') rejected any command whose
quoted arguments contained shell metacharacters — a cargo benchmark
regex filter like '^layer3/write/(a|b)$' disqualified the whole command
even though those characters are literal to the shell.
_has_allowlist_shell_operator is now quote-aware:
- metacharacters inside single/double quotes or behind a backslash are
treated as literal arguments;
- $ and backtick inside DOUBLE quotes still disqualify (expansion is
active there);
- quoted/escaped control characters still disqualify when the command
carries a -c/-e/--command/--eval-style option that hands the payload
to another interpreter (sh -c '...', git -c alias.x='!...' x);
- unterminated quotes disqualify (shape can't be reasoned about).
Compound commands (unquoted ; & | < > backtick $( newline) are rejected
exactly as before. hermes_cli/approvals_suggest.derive_glob picks up the
same semantics via its existing import.
Factory Droid v0.175.0 made TaskOutput/TaskStop accept task-ID prefixes so
background tasks can be referenced without pasting the full ID. Hermes'
process tool had the same friction: every action required the exact
proc_<12-hex> session ID.
ProcessRegistry.get() now falls back to unique-prefix resolution when the
exact lookup misses: 'proc_4dae' or bare '4dae' resolves to
proc_4dae56ca81f6 when exactly one running/finished session matches.
Ambiguous or too-short (<4 suffix chars) prefixes still return None, so
callers keep their existing 'No process with ID ...' error and nothing is
ever picked arbitrarily. Exact IDs never pay the scan, and a full ID that
happens to prefix another always wins.
All process actions (poll/log/wait/kill/write/submit/close) route through
get(), so they all gain prefix support from the single change.
CLI parity for the continuity toggle:
- subcommands/cron.py: --continuity on create; --continuity / --no-continuity
tri-state pair on edit (same store_const pattern as --no-agent/--agent)
- cron.py: forwarded to the cronjob tool; created/edited job summaries print
a "Continuity: on" line
- cronjob_tools._format_job: reports continuity as an explicit boolean and
strips the reserved 'self' entry from the reported context_from list
- cron-job.ts: form reader accepts both shapes (raw store record with 'self'
inside context_from, or formatted record with the explicit flag)
- docs: CLI flag examples in the continuity section
E2E (real argparse -> cron_create/cron_edit -> jobs.json in temp HERMES_HOME):
create --continuity stores ['self']; edit --no-continuity clears; edit
--continuity restores; default-off unchanged. 91 cron/tool tests + 16 CLI
cron tests + vitest 10/10 pass.
Per review: expose run-to-run continuity as a boolean `continuity` flag on
cronjob create/update instead of asking users to know the reserved
context_from='self' value. The flag translates to the 'self' entry in
context_from internally (create: appends/omits; update: adds or removes
'self' while preserving other upstream refs). Schema documents the flag and
steers context_from back to job-id chaining only. Docs updated; 7 new tests.
Amp's 'Right on Schedule' (Jul 21 2026) lets scheduled agents wake up with
their saved context and continue where they left off. Hermes cron jobs run
in isolated sessions with per-run amnesia; the existing context_from chain
mechanism only referenced OTHER jobs. This adds the special value 'self'
(and treats a job's own literal id the same way): the job's most recent
output is injected with continuity framing so recurring scouts/monitors
dedupe against what they already reported and continue where they left off.
- cron/scheduler.py: resolve 'self'/own-id in _build_job_prompt with
continuity framing instead of upstream-job framing
- tools/cronjob_tools.py: allow 'self' through create/update validation
(can't be validated against the store — the job doesn't exist yet at
create time); schema description documents the value
- tests: 6 new tests incl. sabotage-verified failures without the fix
- docs: self-context section in cron.md
Copilot CLI 1.0.78 reworked /rewind to restore only the files the agent
changed, 'skipping any file whose contents no longer match what Copilot
last wrote'. This ports that protection to Hermes checkpoints:
- tools/checkpoint_manager.py: per-project agent-write ledger
(sha256 of every landed write_file/patch), safe_restore_plan()
classifier, and restore(safe=True) that reverts only agent-authored
changes, deletes agent-created files, and preserves user hand-edits.
Empty ledger (pre-existing stores) falls back to the classic full
restore.
- run_agent.py: feed the ledger from _record_file_mutation_result on
every landed mutation (zero new hooks; rides the existing verifier).
- CLI + gateway /rollback: safe mode is the default; --all/--force
restores everything; skipped files are reported with a hint.
- 17 locales: new gateway.rollback.kept_user_edits key.
- Docs: checkpoints-and-rollback.md updated.
- Tests: 7 new cases incl. user-edit preservation, post-agent user
tweaks, agent-created file removal, empty-ledger fallback.
Claude Cowork (Aug 6, 2026) added skill & plugin security scanning:
third-party skills and plugins are automatically checked for malicious
content on upload/edit, returning pass/warn/fail. Hermes already scans
hub-installed skills (tools/skills_guard.py), but `hermes plugins
install` cloned and activated arbitrary Git repos completely unscanned —
and plugins run Python in-process, making them the more dangerous
surface.
- tools/plugin_guard.py: plugin-adapted scanner reusing the skills_guard
pattern engine. Exempts the documented provider-plugin patterns (own
requires_env API-key reads, HTTP calls with keys) on code files while
keeping true threat signals (foreign credential-store access, reverse
shells, destructive/persistence/obfuscation patterns, prompt injection
in docs). Plugin-sized structural limits; VCS/venv dirs excluded.
- hermes_cli/plugins_cmd.py: scan the temp clone before it is moved into
~/.hermes/plugins/. safe=install, caution=confirm (interactive prompt
or --force), dangerous=blocked (--force does NOT override). Re-scan on
`hermes plugins update`; a dangerous updated tree is deactivated until
the user reviews the findings. Dashboard install path returns
structured scan_blocked/scan_findings.
- Config gate: plugins.scan_on_install (default true) in config.yaml.
- Validated against all 60 bundled plugins: 57 safe, 3 caution (real
sudo / curl|sh content in their docs), 0 false-positive blocks.
- 15 new tests incl. E2E through _install_plugin_core with real git
clones.
UTF-16 text files (Windows Notepad .txt, PowerShell > redirects) were
refused as binary: the terminal env decodes stdout as UTF-8 with
errors=replace, so their content arrived mangled with U+FFFD and
tripped the binary guard.
ShellFileOperations.read_file now probes the raw bytes via the
backend's Python when the binary guard fires: a BOM or the zero-byte
parity heuristic (derived from VS Code's encoding sniffer, tolerant of
mixed Latin/CJK content) identifies UTF-16 LE/BE, and the file is
transcoded to UTF-8 with CRLF normalized and the BOM stripped. Real
binaries (zeros at both parities), binary extensions, files over
10 MiB, and legacy 8-bit encodings (GBK, Big5) still refuse — a wrong
silent guess is worse than a clear refusal. Works on every shell
backend (local/docker/ssh) since the probe runs via python3 -c.
Tests run against a real LocalEnvironment (E2E, no mocks); sabotage
run confirmed 6/9 fail without the fix.
MCP tool results carry a server _meta mapping (exposed as .meta by the
Python SDK) alongside structuredContent. Servers return namespaced
machine-readable contracts there (validated payloads, browser-handoff
URLs); Hermes previously dropped the field entirely, so that data was
invisible to the agent.
Now _meta is included in the JSON tool output, after filtering
protocol-reserved keys per the MCP spec's key-name rules: a prefix is
reserved when a modelcontextprotocol or mcp label is followed by at
least one more label (modelcontextprotocol.io/..., tools.mcp.com/...).
Vendor namespaces with a trailing reserved word (com.example.mcp/...)
and unprefixed keys pass through. Non-serializable metadata drops the
extras rather than failing the call.
Unicode TAG characters (U+E0000-U+E007F) render as nothing in terminals
and chat UIs but are fully visible to LLM tokenizers, making them an
ASCII-smuggling prompt-injection channel for untrusted MCP servers.
- tools/ansi_strip.py: new strip_unicode_tags() with fast path; unlike
goose we preserve valid emoji tag sequences (U+1F3F4 base + tag spec +
U+E007F cancel), so regional flags survive.
- tools/mcp_tool.py: applied at every MCP text ingestion point — tool
result text blocks, embedded resource text, read_resource contents,
get_prompt message content, and tool descriptions entering the schema.
- tests/tools/test_unicode_tag_strip.py: smuggled-instruction vectors,
goose's test vector, emoji-tag-sequence preservation, ZWJ untouched.
Port from anomalyco/opencode#40869: parallel tool calls hitting the same
dangerous-command gate each enqueued their own _ApprovalEntry and fired
their own notify_cb — the user got N identical prompts and had to
/approve N times while the agent sat wedged.
_await_gateway_decision now detects an already-pending identical
approval (same command text + pattern-key set) in the session queue and
waits on the leader's event via _await_coalesced_leader instead of
re-prompting. Followers adopt session/always (persistence would auto-pass
a re-check anyway) and deny/timeout (re-asking a just-declined command is
prompt spam); a single-use 'once' makes the follower issue a fresh
prompt. Pre/post approval hooks fire with coalesced=True for followers.
The casing/hidden/literal probes already ran the widened search to produce
their counts, then threw away the paths and returned a hint-only warning.
Strong models pivot in one turn; weak models spiral — the A/B eval measured
qwen3-coder-30b going 3.3 -> 9.3 turns on err_case_search, retrying casing
variants the probe had already resolved.
All three probes (case-insensitive, hidden/gitignored, literal-vs-regex) now
include up to 5 matched paths (+N more) in the warning via a shared tally
helper. Fixes the class, not the site.
Closes#80522
Tracker #79686 P3. Every skill mutation — curator, agent, or user — now
appends one entry to the append-only JSONL ledger at
~/.hermes/skills/.curator_ledger.jsonl, with per-file before/after
manifests whose contents are stored content-addressed (sha256-deduped)
under ~/.hermes/.curator_backups/blobs/.
- tools/skill_ledger.py: append/list/get, blob store, actor derivation
(curator|agent|user), single-entry rollback that takes a pre-rollback
safety entry first and FAILS CLOSED when that capture fails (consistent
with the whole-run tarball rollback hardening from #63366). Path
containment check so a hand-edited ledger can't write outside
HERMES_HOME.
- Hooked all three choke points: skill_manage() dispatch (all actors,
delete intent recorded via absorbed_into/archived evidence),
archive_skill()/restore_skill(), and curator auto-transitions (tagged
actor=curator via a ContextVar override).
- Ledger failures never block the mutation — telemetry, not a gate.
Config gate skills.ledger (default true).
- hermes curator ledger [--skill NAME] [--limit N] and
hermes curator rollback <entry-id> (whole-tree snapshot rollback
unchanged).
- Optional TTL purge of skills/.archive/: curator.archive_ttl_days
(default 0 = never) + explicit hermes curator purge, recorded in the
ledger with before-blobs so purges stay recoverable.
- Docs: curator.md sections on the ledger, single-edit rollback, and
archive TTL purge.
Curator invariants unchanged: only created_by:agent skills auto-transition,
never hard-delete autonomously, pinned exempt; foreground user deletes stay
hard-delete (and are now recoverable via the ledger).
Closes#45778, #50875. Tests adapted from #50261 by @yu-xin-c.
When delegation.provider/model is explicitly pinned, the child no longer
inherits the parent's fallback chain: a mid-run auth/429 failure on the
pin previously rerouted the quiet-mode child onto parent fallback models
with no surfaced signal. Same treatment as the existing override_provider
OpenRouter filter-clearing — explicit pins are honored or fail loudly.
Also upgrades the pinned delegation.command-missing-from-PATH case from
warning + silent transport fallback to a loud spawn refusal, both at
credential preflight and in _build_child_agent.
Fixes#80450 (tracker #79686 audit item).
Port from Kilo-Org/kilocode#12698: report signal-terminated commands with
a human-readable note instead of a bare numeric exit code.
Kilo's fix settles a signal-killed process as the conventional 128+signum
exit code so its bash tool stops hanging. Hermes already produces numeric
codes for signal deaths (subprocess -signum, or the shell's 128+signum),
but the model saw a bare exit_code=-9 or 137 and burned turns
mis-diagnosing (137 = OOM kill being the most common). This adapts the
idea to Hermes' existing exit-code semantics tier:
- _interpret_signal_exit(): maps negative codes (definite signal death)
and the 128+signum band (hedged with 'usually') to a note naming the
signal and its likely cause, wired into _interpret_exit_code() ahead of
the per-command semantics table.
- Curated signal table (SIGKILL/SIGSEGV/SIGTERM/SIGABRT/...) so ambiguous
application exit codes are never mislabeled; uncurated 128+N codes stay
silent, SIGINT is excluded (executor's interrupt-marker path owns
rc=130).
- Notes surface via the existing exit_code_meaning result field.
E2E verified against real SIGSEGV/SIGKILL processes.
Rework of the #88049 inline early-return per review:
- resolve_xai_http_credentials gains an opt-in prefer_api_key flag that
checks the explicit XAI_API_KEY first and falls back to OAuth. The key
is read through tools.tool_backend_helpers.resolve_provider_secret
(config -> profile secret scope -> env/.env -> credential pool) so the
preferred path enforces the same scope policy as the existing fallback
branch, including failing closed under a multiplexed gateway turn.
- The preferred path's base URL honors HERMES_XAI_BASE_URL then
XAI_BASE_URL behind hermes_cli.auth._xai_validate_inference_base_url,
mirroring the OAuth branch (a foreign origin can't exfiltrate the key).
- x_search's _resolve_xai_bearer now calls the shared resolver with
prefer_api_key=True instead of re-implementing precedence inline (#88040).
- tools/tts_tool.py _generate_xai_tts converted to the same flag — same
root cause for /v1/tts 403s (#87045, supersedes the inline shape in
#87081 by @enwaiax).
- Regression tests retargeted at the tools.xai_http.get_env_value seam and
the shared resolver; added coverage for the flag's OAuth fallback,
HERMES_XAI_BASE_URL + origin validation, default-order stability, and a
profile-scope-only key on the preferred path.
- Docs: x-search authentication section now states the explicit API key
wins (metered billing implication).
When a paid XAI_API_KEY is configured alongside xAI SuperGrok OAuth,
_resolve_xai_bearer() took the OAuth path unconditionally. The OAuth
credential authorizes /v1/responses but answers in a degraded Grok
explanatory mode with no citations, while the API key returns real
posts - so every x_search query silently degraded (#88040).
Prefer the explicit API key when set (same shape as the TTS fix for
#87045 in #87081), keeping OAuth as the fallback when no API key is
configured. The shared resolver and every other xAI call site are
unchanged.
The config-read guard (test_config_read_guard) correctly flagged the
probe's raw yaml.safe_load of config.yaml — raw reads miss the managed
overlay, env expansion, and normalization. Use load_config_readonly()
under a scoped HERMES_HOME override instead. E2E v3/v3b and the guard
both green.
Bot Chats created before the epoch mechanism persisted prompts with no
protocol section and no stamp — the staleness check only fires on
stamped prompts, so pre-existing bots would never learn to message
teammates. stored_bot_chat_prompt_needs_upgrade() migrates them: one
rebuild, title-gated to Bot Chat, only when the probe would actually
emit a section (SOUL-append legacies and unmanaged installs are left
alone — rebuilding those would loop). The rebuilt prompt carries the
stamp, so the upgrade can never re-fire.
E2E v3b through the real restore path: legacy Bot Chat upgraded once
then verbatim-reused; legacy regular sessions byte-untouched.
tests/agent/ 4648/4648.
Bot Chats break the "new sessions come often" assumption behind
build-once system prompts: capability edits used to sit invisible until
/new or compression, and the frozen birth date became misinformation.
- tools/bot_mode_probe.py: capability_fingerprint() hashes the profile's
capability surface (disabled skills, toolset pins, MCP config, SOUL.md,
installed skills, Bot-Mode roster); Bot Chat prompts embed the 12-hex
epoch stamp
- agent/conversation_loop.py restore path: stored Bot Chat prompt whose
epoch mismatches disk → ONE rebuild (through a cleared skills-prompt
cache so new installs appear), persisted so the next turn reuses the
new bytes verbatim. Prompts without a stamp — every non-Bot-Chat
session — never take the branch; probe failure fails closed to reuse
- agent/system_prompt.py: Bot Chat prompts are timeless — the
"Conversation started:" date is dropped (timezone kept); no ticking
fields in an eternal session
- tui_gateway: _sync_bot_capabilities at turn start rebuilds the live
agent (tool definitions are construction-baked) when the fingerprint
moves, same session id/history, with a user-visible notice
Cache stance: this is the /model exception applied to capabilities — a
loud, user-initiated, once-per-change prefix break. Unchanged state
hashes identically and stored bytes are reused verbatim (E2E-proven).
Validation: 9 probe unit tests incl. per-axis fingerprint changes;
E2E v3 against the real restore path (fresh build → verbatim reuse →
skill install → single refresh w/ new skill in index → verbatim reuse;
regular sessions dated, unstamped, never refreshed); tests/agent/
4647/4647.
Per review: the protocol belongs only in official Bot Mode interactions,
not every session on a managed install. The prompt builder now injects
the section only when the agent's session row is titled "Bot Chat"
(BOT_CHAT_TITLE, matching the desktop's createCanonicalChat pin and the
`hermes -p <bot> chat -c "Bot Chat"` resume target). Regular sessions
never carry it; the desktop composer middleware owns @mention sends.
Title is read once at first prompt build and the rendered prompt is
cached + DB-restored — cache-safe. E2E against the real AIAgent +
SessionDB: absent in an untitled session, present in Bot Chat,
byte-stable across rebuilds, absent after retitle, absent with the
flag off. Overhead unchanged (~916B, Bot Chat sessions only).
Replaces the plugin-side SOUL.md protocol append: on Bot-Mode-managed
installs (any profile carrying ui_meta['hermes-bots']) the prompt builder
injects the "Messaging other agents" section into every session of every
profile — including headless `hermes -p <bot> chat` sessions a teammate
starts — so bot handoffs work without mutating user-authored SOUL files.
- tools/bot_mode_probe.py: silent-when-unmanaged probe, cached per
(process, home), keyed off the agent's OWN home (not ambient
HERMES_HOME); silent when SOUL.md already carries the legacy section
- agent/system_prompt.py + agent_init.py + config_defaults.py: wired as
agent.bot_mode_protocol (default True), stable tier, byte-stable
across rebuilds (E2E-verified against the real build_system_prompt)
- tui_gateway profiles.list gains bot_mode_protocol capability flag;
the bundled plugin gates ALL SOUL protocol writes on it (backfill,
composeSoul, Edit save) — older gateways keep the SOUL-append path
- overhead: ~916 bytes, only on Bot-Mode installs; zero elsewhere
Supersedes the SOUL backfill half of Hermes-Bot-Mode#99 (credit
@kaduxo — the handle fix, `hermes profile list` correction, and
idempotent-append guards from that PR ship in the bundled plugin).
A same-day version-floor bump (0.20 runtime contract) left every install
with an older cua-driver hard-failing on all computer_use calls: the
start() gate fails closed, while the `hermes update` refresh defers to the
driver's own check-update verb — whose ~20h cache routinely answers "no
update available" right after we raise the floor. Hermes knew it required
0.20+ but never acted on that knowledge.
Two changes:
- tools_config.install_cua_driver(): a contract-failed installed driver is
repaired on the upgrade=True path too (previously only upgrade=False).
The contract failure itself is the confirmation, so the
require_confirmed_update gate and the check-update short-circuit are
bypassed for repairs — an indeterminate or stale-cached check can no
longer pin users on an unusable driver.
- cua_backend.CuaDriverBackend.start(): when the contract gate fails on an
installed binary, attempt one automatic repair per process via the
standard install path, then re-probe. HERMES_CUA_DRIVER_CMD overrides
are never repaired (explicit override is authoritative even when broken)
and a missing binary still just reports the install hint. A failing
installer can't loop: the second start() surfaces the original error.
Tests: contract-repair coverage in test_computer_use.py (auto-repair
success, failed repair surfaces the original error, once-per-process
guard, override never repaired, missing binary never repaired) and
test_install_cua_driver.py (incompatible driver repairs despite an
indeterminate check-update, check-update not consulted). All new tests
verified to fail against the unfixed source (sabotage run).
`--yolo` / `-z` read as "don't prompt me", but they also swap computer_use
onto a private `unrestricted` daemon, dropping the ceilings the configured
mode would have applied. Nothing said so. A script picks up `-z` for quiet
output and loses its limits as a side effect, and the only trace is a driver
process nobody inspects.
The mapping itself stays. It is deliberate, and `unrestricted` is reachable
no other way: it is intentionally not a config value so a stale config line
can never silently bypass approvals (see `_cua_configured_permission_mode`).
Removing the mapping would delete the capability rather than fix it, and
splitting it onto a second CLI flag was declined to avoid growing the
surface.
So the widening is now stated instead: one warning per session naming the
configured mode it left, what stopped applying, and the two ways to keep a
ceiling - drop the bypass flag, or declare a version-3 capability manifest,
which now rides along with unrestricted as of the previous commit.
Once per session, not per dispatch: the resolver runs on every tool call.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`--yolo` / `-z` route the session onto a private embedded daemon in
`unrestricted` mode. That daemon was constructed without the configured
capability manifest, and the serve command only attached
`--capability-manifest` when the mode was exactly `bounded`. So the moment a
run was bypassed, the user's declared ceiling was dropped:
without -z: --permission-mode bounded --capability-manifest ...
with -z: --permission-mode unrestricted --dangerously-bypass-approvals
No manifest, no warning. The most carefully configured run - a reviewed
ceiling, written by hand - became the least constrained one, silently, and
it failed open.
That was never a driver limitation. cua-driver documents the manifest as a
ceiling across modes ("A manifest can narrow a profile but never widen it";
its own authorization table calls it `optional_capability_manifest_ceiling`),
and accepts it alongside `--permission-mode unrestricted`.
The forwarding is version-aware, because the two manifest schemas differ
(cua-driver session_manifest.rs):
* v1/v2 are legacy and must declare `mode: bounded`. Handing one to an
unrestricted runtime aborts startup with "legacy capability manifest mode
must be bounded", so a naive forward would turn a working session into a
hard failure. These are forwarded for bounded only, and a warning names
the migration when one cannot apply.
* v3 must not declare a mode. It is the mode-independent ceiling, and it now
rides along with unrestricted.
Unreadable or unparseable manifests are not forwarded outside bounded, on
the same fail-safe reasoning; bounded still forwards unconditionally and
lets the driver be the authority there.
Verified against cua-driver 0.20.0 on Windows. Launch args now carry
`--permission-mode unrestricted --dangerously-bypass-approvals
--capability-manifest <v3> --approve-capability-manifest`, and the ceiling
is enforced in the bypassed run - a tool outside the manifest is refused
("outside the capability manifest for this session ... blocked as a
protected resource") where the same config previously ran unbounded. A
legacy manifest was confirmed to abort driver startup when forwarded, which
is what the version gate prevents.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`cua_browser_state` has two branches, chosen implicitly: any call carrying
pid or window_id is a *binding* (browser_route.py:252), anything else is a
*snapshot*. A binding clears session state, mints fresh tab_ids, returns
binding metadata with no page content, and sets verification_required.
Nothing in the response says that. A caller that keeps passing pid/window_id
- the natural reading of "bind to this window, then read it" - re-binds
forever: the tab_id it just received is unbound by the next bind, so every
cua_browser_navigate comes back browser_verification_required, and the
refusal ("take a fresh snapshot") points at the same call that just re-bound.
Observed live as 11 consecutive refused navigates before the model gave up
and fell back to foreground SendInput on the address bar.
The same confusion silently swallowed include_screenshot: both calls that
requested one were bindings, which carry no page content, so the flag had
nothing to attach to and was dropped without comment.
A binding response now reports snapshot_required, next_step
(fresh_browser_state, matching the existing token convention) and a hint
naming the exact next call; requesting a screenshot on a binding reports
screenshot_deferred instead of dropping it. The verification refusal now
says to call cua_browser_state WITHOUT pid/window_id and why re-sending them
does not help. The schema documents that include_screenshot applies to
snapshots.
Behavior of the bind and snapshot branches themselves is unchanged - this is
purely about making the split legible to the caller.
Unit-tested. Not verified end to end on the reporting host: the driver
refuses the bind upstream there (`browser_requires_setup: no owned DevTools
endpoint`, and it does not accept a user-launched --remote-debugging-port),
so the typed route never reaches this branch. That attach failure is a
separate cua-driver issue.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Live-testing the Cua Driver 0.20 convergence on Windows 11 (session 2,
cua-driver 0.20.0) surfaced three defects in the existing-profile browser
path and in install status.
1. The config grant was silently nullified by an approval bypass.
`--yolo` / `-z` map onto a private unrestricted daemon, which answers every
browser_prepare. Because the host delegated the entire existing-profile
decision to the driver, that bypass also nullified
`computer_use.grant_existing_profile: false`: a plain `hermes -z` attached
to the user's real Chrome profile and read live page content over CDP, with
the driver reporting it as "the approved existing Chromium profile". It was
never approved.
An approval bypass is consent to skip prompts, not consent to read an
existing profile's pages, cookies, and storage. CuaTypedBrowserRoute.prepare
now enforces the key itself, regardless of permission mode. bounded stays
exempt - its reviewed capability manifest is the authorization boundary.
The authorization inputs are resolved in the backend from config and the
backend's immutable mode, never from model-supplied kwargs.
2. The grant, once set, still could not be used.
With `grant_existing_profile: true` the runtime is launched
`--grant existing-profile` correctly, but cua_browser_prepare then hit a
runtime approval prompt anyway - re-asking the user to authorize what the
config already authorized, and making the documented opt-in unusable on any
non-interactive run, where the prompt has nobody to answer it and the call
dies on approval timeout. The durable, file-backed grant now stands in for
that prompt. Scope is narrow: only the existing-profile prepare, only when
the grant is present; isolated launches still prompt and any resolution
failure falls closed to prompting.
3. `computer-use status` hid a custom override and spliced its output.
With HERMES_CUA_DRIVER_CMD pointed at cmd.exe, status printed the child's
multi-line banner and prompt inside the one-line version field, never
mentioned the override, and advised `hermes computer-use install` - which
install itself (correctly) refuses to run against an overridden path. It now
names the override and mirrors install's update-or-unset guidance, and
version output is reduced to one bounded line.
Verified on the reported host: `-z` existing-profile attach now refuses and
names the key; `grant: true` no longer prompts (33s vs a 300s approval
timeout); status names the override and prints one line. No change to the
reconciliation path - driver SHA256 unchanged end to end.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The sequential tool path only noticed a user interrupt after the running
tool returned: with the deadline disabled it ran the tool inline (fully
blocking), and with a deadline it waited in 5s slices without ever
checking agent._interrupt_requested. Any tool without cooperative
is_interrupted() polling (image_generate, tts, transcription, skills
sync, ...) held the whole turn hostage — the reported symptom was a
redirect queued ~40s behind a FAL image generation + upscale pass.
Executor backstop (class fix, covers ALL tools):
- _run_sequential_tool_execution_middleware always dispatches on the
daemon worker (timeout None no longer means inline blocking) and polls
the interrupt flag every 1s.
- On interrupt: 3s cooperative grace (mirrors the concurrent path), then
synthesize a cancelled tool result (_ToolCancelledResult), emit the
terminal post_tool_call with status=cancelled, and abandon the worker.
- _ToolCancelledResult suppresses downstream post-hook double emission
exactly like _ToolTimeoutResult, so an abandoned worker finishing late
cannot report success for a cancelled call.
- clarify (interactive, _NEVER_PARALLEL_TOOLS) keeps the inline path —
it owns its own human wait.
Cooperative layer in the reported offender:
- image_generation_tool: blind handler.get() (generation + Clarity
upscale) replaced with _wait_fal_result(), which polls is_interrupted()
in 0.5s slices and raises ImageGenerationInterrupted immediately.
- _upscale_image propagates the interrupt instead of swallowing it into
the "upscale failed, use original" fallback.
Message alternation is preserved: the cancelled result is a normal tool
result for the call_id. Sabotage-verified: with the old wait loop
restored, the new tests fail (tool blocks full runtime); with the fix
they pass in ~4s.
The Aug 8 default-on upscaling policy (66ea4e686) chained the Clarity
Upscaler after every sub-2MP generation. Clarity is an SD1.5 creative
tile-diffusion enhancer (creativity 0.35, "masterpiece" prompt prefix) —
it redraws content, which degraded output on 100% of generations for
models like GPT Image 2 and Ideogram whose value is precise text
rendering, CJK, and photorealistic detail.
Policy now: no model upscales by default, on FAL or Krea. The `upscale`
tool param remains as a per-call opt-in (`upscale: true`); explicit
requests still chain Clarity (FAL) / Krea Enhance as before.
- FAL catalog: all 17 default-on entries flipped to upscale=False
- Krea plugin: medium + medium-turbo per-model defaults flipped off
- Tool schema: upscale param described as opt-in with a fidelity warning
- Tests updated: catalog invariant now pins all-off; default-on cases
now assert no upscaler call
- Docs (en + zh) updated to the opt-in policy