Composio-style MCP servers return un-paginated 22-47K-char payloads that
sail under the generic 100K per-result spillover threshold, bloating
context and ballooning per-turn reasoning time on long conversations.
Competitors cap harder (OpenCode/pi 50KB, Claude Code 30K, Codex ~10K
tokens). Three changes:
- mcp_* tools spill at a tighter 50K default (BudgetConfig.mcp_result_size,
config-overridable via tool_budget.mcp_result_size_chars; pinned and
per-tool overrides still win; capped by the context-scaled default).
- The persisted-output preview now teaches recovery: page the saved file
with read_file or process with execute_code instead of re-requesting the
same data from the remote API.
- Untrusted/MCP string results are scanned (bounded, first 64KB) for
provider-side elision markers ('...N more items', "has_more": true,
'saved to sandbox', data_preview) and get ONE cache-safe incompleteness
notice appended at result-construction time, before untrusted wrapping —
so the model stops treating provider-elided enumerations as complete.
- Hard 2M-char allocation cap in mcp_tool.py (text, error, and
structuredContent paths) so a pathological multi-MB server payload is
bounded before it propagates, while ordinary large results reach
spillover intact. Distilled from #56060/#56072/#56511 (issue #56059);
supersedes their 50K lossy truncation with spillover-friendly semantics.
Docs: configuration.md spillover-budget section + cli-config.yaml.example.
Co-authored-by: Stoltemberg <215755014+Stoltemberg@users.noreply.github.com>
Co-authored-by: AlexFucuson9 <295703459+AlexFucuson9@users.noreply.github.com>
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
Unpinned zero-credential installs now pick Exa or Parallel by the
parity of the per-process random session id (stable within a process,
even split fleet-wide) instead of always favoring Parallel. An explicit
hermes tools selection (web.backend / per-capability keys) bypasses the
split entirely; the runner-up vendor stays in the walk as fallback.
Live E2E: 6 fresh processes split 3/3 between vendors, each performed
a real keyless search via its picked endpoint; explicit pin verified.
Un-fences OPENAI_MODEL_EXECUTION_GUIDANCE from the gpt/codex/grok substring
check and gives it its own injection gate, independent of
tool_use_enforcement, controlled by config.yaml `agent.execution_guidance`
(auto/true/false/list — same semantics as tool_use_enforcement). The "auto"
list (EXECUTION_GUIDANCE_MODELS) now also covers deepseek, kimi, qwen, glm,
minimax, mimo, and mistral.
Composio agentic-eval traces showed Hermes+DeepSeek/Kimi failing where
competitors passed: financial math done in prose, no read-back after
external writes, malformed identifiers "repaired", completeness claimed
despite count mismatches. The discipline block existed but those models
never received it.
The block is extended with compact clauses distilled from that analysis:
- external-write read-back (tool-call success is not task success; internal
file edits already confirmed by the tool are not re-verified)
- count reconciliation (declared totals/has_more are hard assertions)
- literal preservation (never normalize identifiers that fail a stated
format; lookup success does not validate a malformed token)
- retry-differently (empty/partial/suspiciously narrow results get a
broader retry before concluding)
- completion gated on verification (done = every named acceptance
criterion verified, never a plausible subset)
The todo tool description now encourages enumeration-as-checklist for
"all N items" tasks and gates completed status on verified work, never
intent.
Guidance is chosen once at session start keyed on model name, so the
system prompt stays byte-stable for the life of a conversation.
Supersedes/absorbs prior contributor proposals: #20588, #35087, #41874
(MiMo), #53847 (GLM tool-calls-as-text stall).
Co-authored-by: Mat-London <56627804+Mat-London@users.noreply.github.com>
Co-authored-by: intelac <8803887+intelac@users.noreply.github.com>
Co-authored-by: 6ylqq <51219463+6ylqq@users.noreply.github.com>
Co-authored-by: tauros1983 <267660491+tauros1983@users.noreply.github.com>
Two real gaps the CI-red sibling tests exposed:
- read_selection() treated EVERY raw stt.provider: local as the legacy
DEFAULT_CONFIG seed and reported no-selection — but the seed never
reached config.yaml (save_config strips schema defaults), so a
picker- or hand-written local pick was silently discarded and the
autodetect ladder could route an explicit local user to cloud STT.
A raw 'local' is now a genuine selection; the merged-view ambiguity
note replaces the over-broad shim (mirror comment updated in
nous_subscription._selected_provider and _get_provider).
- _reconfigure_provider was half-migrated: the tts/stt/browser/web
branches and the managed-category fallthrough still wrote
use_gateway flags and vendor names for managed rows. They now write
the single provider string ('nous' for managed rows) and pop the
legacy key, matching _write_provider_config.
An explicitly stored browser.cloud_provider that names no registered
plugin now raises the honest selection-naming error instead of warning
and silently auto-detecting; the auto-detect walk (including the managed
gateway entitlement probe) runs only when no cloud_provider key was ever
written. The 'nous' selection routes to the Browser Use provider, whose
config resolver is now a strict switch: 'nous' => managed only, stored
vendor => direct BROWSER_USE_API_KEY only with a selection-naming error
when missing. Camofox is selected via browser.cloud_provider: camofox;
CAMOFOX_URL stays the server ADDRESS only and can no longer override an
explicit different selection (never-configured installs keep the legacy
env-var activation).
Both _resolve_openai_audio_client_config resolvers now switch on the
stored provider string: 'nous' (or legacy use_gateway: true) => managed
openai-audio gateway only, erroring by selection name when unentitled —
the STT twin previously never read the stored gateway intent at all, so
a direct OPENAI_API_KEY silently overrode the Nous Subscription pick;
stored vendor => direct credentials only with a selection-naming error
on missing keys (no silent managed fallback); never-configured keeps the
legacy ladder. DEFAULT_CONFIG stops seeding stt.provider: local, and the
seeded value on existing configs is treated as no-selection so autodetect
keeps working for that installed base.
_get_backend returns the stored web.backend verbatim (mapping the managed
'nous' selection to the firecrawl provider) — unknown names surface the
honest selection-naming error at dispatch instead of silently rerouting
through the credential ladder, which now runs only on never-configured
installs. _get_capability_backend no longer discards an explicit
search/extract backend when its availability probe fails. The firecrawl
client resolves strictly: 'nous' => managed gateway only (unavailable =>
selection-naming error), stored vendor => direct only (no FIRECRAWL key
=> error, never a silent managed fallback billed to Nous).
Add read_selection()/selection_exists()/selection_error() to
tool_backend_helpers: one provider string per category ('nous' = managed
Nous Tool Gateway, vendor name = direct with the user's own credentials,
no key ever written = legacy credential autodetect). Legacy configs are
interpreted at read time only (use_gateway: true => nous); nothing is
migrated on disk, and the DEFAULT_CONFIG-seeded stt.provider: local is
treated as never-configured.
_resolve_managed_fal_gateway / _resolve_managed_fal_video_gateway now
switch on that string: 'nous' routes managed only (unentitled => error
naming the selection), a stored vendor routes direct only (missing
FAL_KEY => error naming FAL_KEY and the selection, no silent managed
reroute), and FAL_KEY presence no longer selects the route. Krea's
model-driven managed interception now requires no stored provider (or
the managed selection) instead of merely provider != krea, and the
image/video registries map the 'nous' selection to the FAL plugin.
With zero web credentials configured, web_search/web_extract previously
resolved to the nonfunctional firecrawl sentinel and errored. Now the
backend resolution walks a strictly-last keyless tier: Parallel's and
Exa's public anonymous MCP endpoints (the same free tiers opencode ships
as its default search path).
- plugins/web/keyless_mcp.py: minimal JSON-RPC tools/call client for
mcp.exa.ai + search.parallel.ai (SSE + plain JSON parsing, typed
errors, per-process random session id, no user identifiers)
- WebSearchProvider.is_keyless_available(): separate weaker tier that
never leaks into is_available(), so keyed setups are never pre-empted
- Exa/Parallel providers: route to keyless endpoints when their key is
absent; keyed SDK path unchanged
- registry + _get_backend(): keyless walk (parallel -> exa) strictly
after every keyed/importable candidate; check_web_api_key() lights
the tools up on zero-credential installs
- web.keyless_fallback config key (default true) to disable the tier
- docs: web-search.md + configuration.md
E2E-verified against both live endpoints from an isolated HERMES_HOME
(search + extract via the real dispatchers, disable-flag negative path).
open_preview and read_preview could drive the page, but nothing could close
the pane. Same desktop_ui / session-source gate as the rest of the GUI tools.
Follow-up to the salvaged fix. Three parity gaps in the new CLI branch:
- Timeout arm dropped the denial-breaker addendum that the same function's
gateway arm and check_all_command_guards' CLI tail both append, so a
tripped breaker went unreported on a timeout.
- Human deny called _record_denial(), advancing a tally scoped to guardian
LLM DENY verdicts. Neither sibling CLI tail does this, so three
deliberate user denials escalated to breaker hard-stop text.
- The platform-marker half of the leak (HERMES_SESSION_PLATFORM set, no
HERMES_EXEC_ASK) was unpinned; it reaches the same branch.
Adds a platform-marker regression test plus two breaker-parity guards, and
clears the process-global _denial_tally in the shared fixture so a leaked
tally can't bleed the escalated addendum into unrelated assertions.
e37a0321eb fixed _run_approval_gate and check_all_command_guards: when
HERMES_EXEC_ASK (or a session platform marker) leaks into an interactive CLI
process with no gateway notify callback registered, those two functions now
prefer the registered CLI Dangerous Command panel over a silent
pending_approval nobody can see.
check_execute_code_guard — the whole-script gate for execute_code, a
separate function with its own copy of the same notify_cb-less
short-circuit — never got the same treatment. It doesn't even accept an
approval_callback parameter. In the same leaked-ask-mode-into-CLI scenario,
execute_code calls still silently drop into pending_approval with the panel
never shown, even though a CLI callback is registered.
Compute is_cli/approval_callback the same way the two fixed functions do,
and when _should_fall_through_to_cli_approval() says yes, run the same
hook-fire -> prompt_dangerous_approval -> hook-fire -> choice-branch
sequence _run_approval_gate's tail already uses, adapted to this function's
own message/persistence conventions (smart-denied session/permanent
suppression, denial-breaker addendum). Falls back to the existing
pending_approval behavior when no CLI callback is available.
Tests: 4 new cases in tests/tools/test_cli_approval_exec_ask_leak.py
mirroring the existing check_all_command_guards pair (approve/deny/timeout/
session-persistence). Mutation-verified: all 4 fail against the pre-fix code
and pass with it restored.
Neighbor suites: tests/tools/*approval* (291+ tests) and
tests/gateway/{test_approval_prompt_redaction,test_tui_approval_redaction,
test_plaintext_approval_routing,test_discord_exec_approval_content} +
tests/cli/test_cli_approval_ui.py all green. The 7 test_approval_mode_parity
/ test_nonrecursive_verification_artifact_cleanup failures seen in one full
batch run are pre-existing and independent of this change — confirmed by
re-running the identical batch with tools/approval.py stashed back to
pre-fix: the same 7 fail for the same reasons either way.
Note on an adjacent open PR: #65592 also touches check_execute_code_guard,
but an earlier, unrelated region of the function (adding an AST dangerous-
operation scanner to the "not is_gateway and not is_ask" auto-approve
branch). No semantic overlap with this fix's notify_cb-less branch; a small
rebase may be needed depending on merge order.
ceabb030f added the Grok Imagine Image 2.0 catalog entry with upscale=True,
violating the Aug 2026 opt-in-only upscaling policy (f06c41522) and breaking
test_upscale_defaults_are_all_off on main, which reddened every PR's slice
12/12.
Tool Search treated desktop_ui and project as plugin catalog, so
read_window_below dropped out of the model-facing array and the HUD
note never fired. Session-gated GUI tools stay off the core list but
remain direct once a session enables them.
Co-authored-by: fangliquan <fangliquan@qq.com>
One generic tool in the desktop_ui toolset: discover what is on screen,
highlight an element with narration, or hand the user a paged tour. No tour
content lives in the code — the agent authors each one live, which is what
makes 'how does this work?' answerable as a walkthrough instead of a wall
of text.
Rides the existing blocking-prompt bridge (tour.request/.respond) like
read_preview, so it works on every connection topology.
The Bot Mode teammate-DM protocol told agents to inline the message into a
double-quoted shell argument: quotes truncated the body and $(...)/backticks
executed on the sender's machine. The protocol now writes the message to a
temp file and delivers it via a new 'hermes chat --query-file' flag (or '-'
for stdin); 'hermes peer dm' already accepted stdin and the peer recipe now
uses it. No shell pass touches the body at any point.
Supersedes the tool-based approach in #89077 — same bug, fixed with a CLI
flag + protocol rewrite instead of a new model tool.
Co-authored-by: mehmetkr-31 <mehmetkr-31@users.noreply.github.com>
Some authorization servers and WAFs reject httpx's default User-Agent on
the OAuth token endpoint. mcp_servers.<name>.oauth.user_agent now stamps a
custom User-Agent onto the two token-endpoint requests (authorization-code
exchange and refresh) on both provider construction paths. Opt-in,
per-server, token requests only — never MCP traffic or discovery, and no
other headers are configurable. Empty/null/non-string values are ignored.
Completes the second half of #75576 (the CIMD half landed via #89566).
The _MAX_RESERVED_SOCKETS cap applied to pinned CIMD sockets too, so under
heavy concurrency an ephemeral-reservation churn could close a parked pinned
socket before _wait_for_callback adopted it, silently reopening the
port-stealing window the pin exists to prevent (#22161). Eviction now skips
the pinned range; it is already bounded by _CIMD_PORTS.
Follow-up to the #84050 salvage.
The agent could reveal single panes (focus_pane) but had no way to arrange
the workspace as one act. apply_layout closes that gap: a desktop_ui tool
that emits layout.apply over the existing bridge, resolved in the renderer
against the layouts contribution registry — the same list the layout picker
reads — so core presets (default/focus/terminal-deck/quad), plugin presets,
and user-saved presets are all addressable by id. Active session only, same
as pane.reveal: a background turn never rearranges the user's desktop.
The `questions` parameter had a full description, but the top-level
tool description still described three single-question modes and never
mentioned batching. The model decides how to call a tool from that
description, so it kept asking one question per call.
The description now states that 2-5 independent questions can go in
one call and that one batched call is preferred over a chain of
single-question calls. The parameter description also tells the model
to put a short batch title in the still-required top-level question.
Two schema tests pin the contract: the description names the batch
capability, and the questions parameter stays optional with the
MAX_QUESTIONS cap.
The clarify tool gets an optional questions parameter (2-5 independent
questions, issue #18450). Batch-capable platform callbacks receive the
normalized list in one call and reply with per-question answers. Legacy
callbacks are looped one question at a time. The loop stops on timeout
so the user is not asked the remaining questions after they walk away.
Locked answers survive a timeout: the result carries them plus a
timed_out flag, and unanswered entries have an empty user_response.
The single-question path is byte-identical to the previous behavior.
tools/skills_sync.py bound HERMES_HOME / SKILLS_DIR / MANIFEST_FILE at
import time — the third module in the same lineage as skills_tool
(f8723c478) and skill_manager_tool (c6a3d412d). In a long-lived
dashboard/TUI process, console skills commands (reset, diff,
list-modified, opt-in/out, repair-official) dispatched in-process under
_profile_scope's set_hermes_home_override(), but skills_sync's frozen
constants kept resolving against whichever profile was live at import.
Sharpest edge: reset_bundled_skill()'s #48200 rmtree strict-child guard
was computed against the WRONG skills root.
Fix: same call-time accessor pattern as the two prior fixes —
_hermes_home()/_skills_dir()/_manifest_file() honor an explicitly
patched module global (tests, retargeting) and otherwise re-resolve
from the live profile-scoped get_hermes_home() on every call. All 37
call sites migrated; module constants kept for compat.
Also documents in _profile_scope() that skills_sync needs no module
retargeting since the contextvar override now reaches it.
Regression tests (sabotage-verified: all 3 fail on the old binding):
- accessors follow set_hermes_home_override at call time
- explicit module patch still wins over the override
- rmtree guard anchors on the overridden profile's skills root
Fixes#65828
_write_manifest still used a hand-rolled mkstemp + atomic_replace,
so every sync reset .bundled_manifest to mkstemp's 0600, dropping a
group-readable or shared mode the operator had set. Replace the block
with utils.atomic_write_text(preserve_mode=True) — the same shared
writer and mode-preservation contract PR #86255 applied to the skill
manager's document writes.
This is the remaining half of PR #14410 by @sgaofen, who reported the
manifest mode reset first; the skill-manager half of that PR was
superseded by the atomic_write_text refactor and #86255.
_fetch_owner_handle and the browse catalog walk both inlined the
same owner-handle extraction logic that _owner_from_payload was
extracted to centralize. Replace both with calls to the helper.
ClawHub treated the last path segment as a slug, so a GitHub-style
id like owner/repo/skills/skillopt fetched a different author's
skillopt. Pair metadata and files from the same source so inspect
cannot show one registry's header and another's SKILL.md.
Per review: even with an active sandbox env, spilled tool results
belong in $HERMES_HOME/cache/spillover with the other Hermes-owned
caches — not the sandbox temp dir as primary storage.
- Host-side write happens first on every backend; local/no-env
sessions reference the host path directly (unchanged).
- cache/spillover joins the auto-mount/sync cache-dir list
(credential_files._CACHE_DIRS), so docker bind-mounts it and
modal/ssh/daytona file-sync it. Remote references use the
translated in-sandbox path after a readability probe.
- Probe failure (persistent containers created before spillover
joined the mount list, translation failures) falls back to the
previous in-sandbox temp-dir copy, so nothing regresses.
Sessions that never ran a terminal command (MCP-only, cron, gateway)
have no active sandbox environment, so maybe_persist_tool_result()
got env=None and fell through to the inline-truncate fallback --
a 467K MCP result was cut to a ~1.3K preview with no file written
('Full output could not be saved to sandbox').
Now the host-side cases (env=None or the local backend) write the
spill file directly to $HERMES_HOME/cache/spillover/<id>.txt,
alongside the other Hermes-owned caches instead of littering /tmp.
Remote backends (docker/ssh/modal/daytona) keep the in-sandbox
env.execute() write since read_file resolves in-sandbox there.
Cleanup: the gateway housekeeping loop prunes spillover hourly with
the other media caches, and a once-per-process best-effort prune on
first spill covers CLI-only installs.
Control path: delegate_task(action=list/steer/stop) resolved ownership
purely through the _delegate_parent_ref weakref identity chain. The CLI
rebuilds its AIAgent mid-session (self.agent = None on route-signature
change, credential refresh, /model, MoA one-shots), so a running child's
chain pointed at a dead object and the child went invisible/unsteerable
while completion delivery (durable session-id routed) still worked.
Observed live 2026-08-17: deleg_88454b70 / sa-0-dc0100f4.
Fix: register each child with the owning conversation's durable session
id (owner_agent_session_id, the same spine delivery routes by) and add a
second ownership tier that matches it against the calling parent's
session_id with compression-lineage resolution on both sides. Foreign
sessions still fail closed.
Presentation path: background processes started BY a subagent (task_id ==
subagent_id) route their notify_on_complete notifications to the parent
conversation by design, but arrived as anonymous raw output walls. The
formatter now resolves the task_id against the live + recently-finished
subagent registry (bounded retention survives child completion) and adds
a provenance line (subagent id, delegation id, goal snippet), trimming
the output tail for subagent-owned processes. Parent-owned process
notifications are byte-identical to before.
Review feedback from NVIDIA (Nir Paz), minus the LLM items (declined
on the thread: cost-by-default + prompt-injection surface; static-only
also keeps the timeout moot at ~1.5s vs the 120s ceiling):
- Incomplete-validator findings are now PRESERVED as partial evidence;
only the validator's pass/fail verdict is excluded from the advisory
verdict. A report with findings from an incomplete check no longer
reads as clean.
- Clean-report wording is now "no findings from completed checks"
whenever any validator was incomplete.
- Pinned both scanner binaries to known releases in code comments,
config guidance, and docs: SkillEvaluator v0.1.0, SkillSpector v2.9.5.
- Tests: 29 (was 28) — partial-evidence preservation flips the old
discard-pinning test, plus the completed-checks wording case.
Review feedback from NVIDIA (Nir Paz): run the full deterministic
Tier 1 surface, not just pii,unicode,lint.
- TIER1_CHECKS now pii,unicode,lint,license,security. License is pure
static (no measurable cost); security invokes NVIDIA SkillSpector in
its keyless static-rules mode (~+1.2s per install). schema/quality
stay excluded: hygiene signal ("author not specified" is
high-severity upstream), wrong noise for an install prompt.
- SkillSpector is a second optional binary, pinned separately. Absent
or failing, the security check reports status="incomplete" and the
adapter treats it as "no opinion" — surfaced as a dim "(not run: ...)"
note, never as a failure.
- _parse_report derives the verdict from COMPLETED validators only.
This also absorbs a live upstream inconsistency: SkillEvaluator's
anti-tamper cross-check on SkillSpector's risk score currently trips
on moderate-finding skills (fail verdict with zero findings, e.g.
github-pr-workflow at 15 MEDIUM issues / score 35). Reported to
NVIDIA separately; either way an evidence-free fail must not render
as an unexplained failure at install time.
- Dashboard tier1 block gains incomplete_checks.
- Docs: SkillSpector install command + not-run semantics.
- Tests: 28 (was 24) — incomplete-status exclusion, verdict derivation,
not-run formatting.
E2E against real binaries: clean skill (no findings), skill tripping
the upstream consistency check (passed, "(not run: Security Scan)"),
seeded dirty skill (2 findings, SECRETS row). Full scan cost measured
at ~1.4-1.5s per skill, install-time only.
Adds an optional, advisory second-opinion scan to the skills hub install
path using NVIDIA SkillEvaluator's deterministic, keyless Tier 1 checks
(PII, unicode smuggling, script lint).
- tools/skillevaluator_scan.py: subprocess adapter — runs the scanner
over the quarantined bundle, parses the JSON report, classifies
secrets-class findings (private keys, tokens, credentialed connection
strings) apart from advisory PII findings. Every failure mode
(binary missing, timeout, crash, bad JSON) degrades to a no-op.
- hermes_cli/skills_hub.py: prints the advisory panel after the built-in
guard's policy decision and before the install confirmation. Findings
are shown with file:line; secrets-class findings render red with a
loud warning. Warn-and-continue by design — the built-in skills guard
remains the only enforcement layer, because the upstream PII scanner
has known false-positive classes (git@github.com, docs example
emails, op:// references).
- hermes_cli/web_routers/skills.py: the dashboard Browse-hub scan
endpoint returns the same advisory data in a new `tier1` field.
- config: skills.tier1_advisory (default true; no-op without the
optional scanner binary on PATH).
- docs: user-guide/features/skills.md section with install command and
config toggle.
Scanner install (optional):
uv tool install --python 3.13 \
"skillevaluator @ git+https://github.com/NVIDIA/SkillEvaluator.git"
E2E-validated against the real scanner binary: clean bundled skill (no
findings, "no findings" line), seeded dirty skill (email + credentialed
connection string -> yellow/red panel, install continues), config
disable via real config.yaml (silence). Real scan cost: ~0.2s per skill.
Bots could message teammates on their own machine (hermes -p <bot> chat) and
the desktop could relay user mentions over Connections, but a bot had NO
transport to a bot on another gateway. This adds one, with zero new server
surface: the peer's existing api_server platform is the wire.
- hermes_cli/subcommands/peer.py: `hermes peer add/list/remove/dm`.
`dm <peer>[/<agent>]` resolves the remote agent's canonical "Bot Chat"
(list by title, create when missing), runs one synchronous agent turn via
POST /api/sessions/{id}/chat, and prints the reply on stdout — the exact
cross-machine twin of the local bot-messaging command, so the Bot Mode
protocol composes over it unchanged. Named profiles route via the peer's
/p/<profile>/ multiplex mirror. Peer URLs live in config.yaml
(`bot_peers`); the peer's API_SERVER_KEY is a credential and lives in
~/.hermes/.env as HERMES_PEER_<NAME>_KEY.
- hermes_cli/main.py: parser wiring + fast-path/session-flag command sets.
- tools/bot_mode_probe.py: when peers are registered, the injected Bot Chat
messaging protocol gains a cross-machine paragraph (peer roster +
`hermes peer dm` pattern) so agents discover remote teammates on their
own; peers join the capability fingerprint so registering/removing one
refreshes eternal Bot Chat prompts on the next message (loud, one-time,
user-initiated — no per-turn cache drift).
- Docs: Bot Mode guide (bot-initiated DMs across machines) + cli-commands
reference (`hermes peer` section + summary row).
Tests: tests/hermes_cli/test_peer_cmd.py (target parsing, /p/ scoping,
registry round-trip in isolated config, real-loopback-HTTP dm flow incl.
Bot Chat create-vs-reuse and bearer auth), bot_mode_probe peer-paragraph +
epoch tests. E2E: real `python -m hermes_cli.main peer ...` against a live
fake peer over HTTP with isolated HERMES_HOME (config/.env persistence,
bare + /p/<profile> routing, stdin, --json). 23 passed; ruff clean.
Completes the project-local skills epic's remaining skill items (#48974,
#48975) on top of the discovery/trust work in #88566.
Quarantine (#48974): trust is a repo-level decision made once, but repo
skill content changes with every pull — the hub install path scans, a
checkout didn't. Every project SKILL.md dir now runs through the same
skills_guard scanner as hub installs (content-hash cached under
~/.hermes/cache/project_skill_scans/, never inside the repo). Verdict
'dangerous' quarantines the skill: excluded from the index, skills_list,
and slash commands via the single iteration chokepoint
iter_project_skill_files(), and skill_view refuses by name with an
explanatory error. Scanner failure fails closed. Verified against a real
injection fixture (6 findings: prompt_injection_ignore, deception_hide,
invisible_unicode, credential exfil patterns).
Non-interactive inheritance (#48975): find_project_root() now resolves
from TERMINAL_CWD (the per-surface workdir cron jobs and the terminal
tool already use) before falling back to process cwd. Cron/API/ACP
surfaces inherit a prior interactive trust decision by project identity:
job workdir inside a trusted repo => project skills load; untrusted or
no workdir => nothing loads; no surface ever prompts.
Tests: +10 cases in tests/agent/test_project_skills.py (real malicious
fixture, fail-closed, rescan-on-change, cache location, TERMINAL_CWD
inheritance matrix). Docs: quarantine + non-interactive sections in
skills.md.
Sessions started inside a git checkout now source skills from
<root>/.hermes/skills/ and <root>/.agents/skills/ (the cross-tool
convention shared with other agent harnesses) as the highest-precedence
skill tier: project > local > external_dirs.
Loading is trust-gated per repo (skills.trusted_project_dirs, managed by
'hermes skills trust'/'untrust') because skills are executable procedure
documents — auto-sourcing them from any cloned repo is a prompt-injection
vector. Untrusted repos with skills get a one-line banner notice instead.
- agent/skill_utils.py: find_project_root, get_project_skills_dirs,
get_untrusted_project_skills_root, get_scan_ordered_skills_dirs;
project dirs join the curator read-only ownership boundary
- agent/prompt_builder.py: project tier scanned first, entries tagged
[project], same-named local entries shadowed; cache key extended
- tools/skills_tool.py: skills_list scans project dirs first (first-wins);
skill_view resolves cross-tier collisions in favor of the project tier
(same-tier ambiguity still refuses); security warning recognizes the tier
- agent/skill_commands.py + hermes_cli/commands.py: /skill-name slash
commands and gateway slash menus include project skills
- tools/credential_files.py: project dirs mounted into remote backends
- cli.py: banner notice (loaded count / trust hint)
- hermes_cli/main.py + subcommands/skills.py: hermes skills trust/untrust
- config: skills.project_discovery (default on), skills.trusted_project_dirs
- docs: Project-Local Skills section in skills.md
- tests: tests/agent/test_project_skills.py (18 cases)
Session cwd is fixed at agent build time, so the resolved tier is stable
for the conversation and the system prompt stays byte-stable (cache-safe).
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).
Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
last_fire_error ({at, detail}) on the job record; mark_job_run clears
it on the next successful run so it always describes current
auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
job on the gateway-unreachable path, best-effort (never disturbs the
503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
"Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
provider is active but the api_server adapter is not running (the
fire path is dead-on-arrival; most common cause is API_SERVER_KEY
missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
The live-checkout git mutation guard blocked history-rewriting git ops
(checkout, reset --hard, rebase, cherry-pick, ...) in the running source
checkout and its worktrees on every platform. The hazard it protects
against is only real on Windows, where NTFS locks loaded module files and
an in-place rewrite can corrupt the running process. On POSIX, open file
handles pin the old inodes, so a checkout swap under a running process is
safe, and the guard mostly taxed normal dev/salvage workflows with clone
workarounds.
- tools/self_repo_guard.py: add guard_active() -> os.name == "nt"
- tools/terminal_tool.py: consult guard_active() before running the
detector; detector logic and block message unchanged for Windows
- tests: wiring tests force the guard on; new tests cover the POSIX
pass-through and the platform predicate
/simplify-code residual. The note hard-coded "'commits' and 'dirty' are
UNKNOWN", but the two probes fail independently: a bad base_commit fails
rev-list while `git status` still succeeds, so `dirty` is a REAL measurement
being reported as unknown. Safety was never affected (the worktree is preserved
either way), but telling the parent a measured value is untrustworthy is its own
kind of misreport — and it would push a human toward re-inspecting something
already proven.
`mark_worktree_payload_unproven()` now takes an `unmeasured` argument, and
finalize tracks which probe actually failed. The raising path still disclaims
both, because which probe raised is unknowable there.
Validation: 22/22 tests/tools/test_subagent_worktree.py; ruff + ty clean. New
guard mutation-checked (hard-coding "commits/dirty" back fails it).
Phase 2c fold. The schema guard added in the previous commit read and
AST-parsed delegate_tool's source, which AGENTS.md:1514 bans outright ("Never
read source code in tests" -- it passes when the implementation is subtly
broken and fails on a correct refactor). Extracting the shared factory the rule
prescribes removes the duplication the AST test was invented to police, so one
change resolves both.
- subagent_worktree: new module-level `mark_worktree_payload_unproven()` +
`unproven_worktree_payload()`. Both producers of this schema now call them,
so the payload cannot drift and the note string exists once.
- delegate_tool: the finalize-raised fallback calls the factory instead of
hand-building the dict (-16 lines). The re-import is guarded: the outer
`except` can be entered because the `from tools import subagent_worktree`
itself failed, in which case the name is unbound -- an inline fallback keeps
the flag rather than raising NameError and losing it.
- Test replaced with a BEHAVIORAL equivalent: it calls the real factory and
compares its key set against live `finalize_subagent_worktree()` output. Same
contract, no source reading, refactor-proof, and it actually executes the
code.
Also folded from the same review:
- Fail-closed on an unmeasurable commit count. With no `base_commit` the
rev-list probe never ran, `commits` kept its unproven 0 default, and a clean
tree still reached `git worktree remove --force` + `git branch -D` -- the
exact bug class #88113 is about, on a public function that takes a
caller-supplied dict. Now returns un-inspected instead, with a test driving a
real child commit.
- Per-probe diagnostics: the note said only "rev-list/status non-zero". It now
names WHICH probe failed, its exit code, and a bounded git stderr tail, so
the parent (and the human) can act on first read.
- Dropped the redundant `inspection_ok` bool for a `failed: list` of reasons;
removed the duplicated index-corruption block in favor of the existing
`_break_git_index()` helper.
Validation: 21/21 tests/tools/test_subagent_worktree.py; ruff clean; ty clean
on subagent_worktree.py and 64-vs-64 unchanged on delegate_tool.py (all
pre-existing, verified against the base commit). All 6 guards mutation-checked
twice -- neutering the flag fails 6, reverting production to pre-fix main fails
the same 6. E2E on real git: clean still prunes; corrupt index keeps the work
and reports the real stderr; empty base_commit keeps a committed child.
Review fold on the #88113 follow-up. The new guards asserted implementation
details that a strictly-better future change would break, and the second
producer of the payload schema had no coverage at all.
- The distinguishability test asserted the failure payload was byte-identical
to the genuinely-clean one (`for key in commits/dirty/pruned: assertEqual`).
That freezes the AMBIGUITY as a required property: emitting `commits: None`
for "unknown" would improve exactly what #88113 is about and fail the test.
Now asserts what the parent actually depends on -- both keep the worktree,
and only the flag separates them.
- `assertNotIn("inspection_failed", ok_payload)` pinned key ABSENCE on the
happy path, forbidding an always-present-but-False flag (a legitimately
better JSON contract: stable key set for serializers). Now
`assertFalse(...get("inspection_failed", False))` -- same coverage, tolerant
of that refactor.
- `assertIn("UNKNOWN", note)` coupled tests to one word of English prose, and
was not even a cross-producer contract: delegate_tool's note said "state
unknown" (lowercase), so a copy-edit broke the implied convention. Tests now
assert the note names the worktree AND branch -- the actionable part for a
human -- and both producers' notes were aligned to read as one contract.
- The raises test never proved its patched seam ran (a future short-circuit
before any git call would keep it green while proving nothing). Now checks
`call_count` and mirrors the branch-survival + note-names-path legs its
sibling had.
- NEW `WorktreePayloadSchemaTests`: commit 2's whole point is the schema the
parent reads, but delegate_tool's fallback -- the second producer -- was
verified only by reading. It now AST-parses the real fallback dict literal
and compares against live `finalize_subagent_worktree()` output, so the two
producers cannot drift and the pre-fix leak (repo_root/base_commit, missing
commits/dirty/pruned) cannot come back.
- Docs/docstring drift: the flag has a second trigger (finalization itself
raising, handled in delegate_tool), and the module docstring listed
`inspection_failed` without `note`. Both corrected.
- Extracted the duplicated 5-line "corrupt the index" setup into
`_break_git_index()` beside the file's other module-level helpers.
Validation: 19/19 tests/tools/test_subagent_worktree.py; ruff clean. New
schema guard mutation-checked -- reverting delegate_tool's fallback to the
pre-fix `dict(_worktree_info)` shape fails it. Restores checksum-verified.
The preserved worktree is invisible to the only consumer that can act on it.
Completes the #88113 fix. That change correctly stops the destructive prune
when a git probe fails, but still returns commits=0 / dirty=False -- values
that were never measured. Those are the defaults the prune used to delete on,
so the failure payload is byte-identical to "inspected fine, child left
nothing":
inspection FAILED, uncommitted work kept -> {commits: 0, dirty: False, pruned: False}
inspected OK, child produced nothing -> {commits: 0, dirty: False, pruned: False}
The only failure signal was a logger.warning, and the sole consumer of this
payload is the parent agent reading the serialized delegate_task entry -- it
cannot read logs (no in-repo code reads the key back). So the parent's rational
reading of the failure case is "the child produced no work", which is the exact
wrong conclusion: a worktree possibly full of uncommitted work is preserved and
then never looked at. The data survives but nobody is told to recover it.
Changes:
- subagent_worktree: one _unproven() helper stamps inspection_failed + a note
naming the worktree/branch, warns, and returns the payload. Both unproven
exits route through it, so they cannot drift apart again.
- subagent_worktree: the pre-existing exception path (timeout, OSError, a
non-numeric rev-list stdout) produced the same unproven payload but logged at
DEBUG -- effectively silent. It now takes the same flagged path as a non-zero
exit; identical outcomes get identical reporting.
- delegate_tool: the caller's finalize-raised fallback assigned the
creation-side metadata dict (path/branch/repo_root/base_commit) -- a disjoint
schema missing commits/dirty/pruned. It now emits the same flagged shape, and
logs at WARNING.
- Docs + docstring + module contract now state that pruning requires
affirmative proof, so a future cleanup doesn't "fix" the preserved worktree
by restoring the unconditional prune and reintroducing this P1.
Purely additive: the happy-path payload shape is unchanged, so no existing
reader can break.
Validation:
- 18/18 tests/tools/test_subagent_worktree.py; 127 passed across the delegation
suites (test_delegate, batch_validation, control_actions, timeout_diagnostic).
- 3 new guards mutation-checked: neutering the flag fails all three; reverting
the production file to pre-fix main fails all three. Restores checksum-verified.
- E2E on real git: inspection-failure now returns inspection_failed=true with
work intact on disk; proven-clean still prunes (pruned=true).
finalize_subagent_worktree() treated a non-zero exit from its rev-list
or status probes as proof of the payload defaults (commits=0, clean),
then pruned on them: git worktree remove --force plus branch -D
permanently deleted a child's uncommitted work whenever git could not
inspect the tree (e.g. a corrupted index) (#88113).
A destructive cleanup now requires affirmative proof of zero commits
plus a clean tree. Any non-zero inspection result keeps the worktree
and branch for manual review, with a warning naming both.
Phase 2 of the MCP 2026-07-28 migration (#69931), on top of the SDK 2.x
migration (#88180):
- Protocol-era negotiation (_negotiate_session): per-server `protocol`
config key — auto (default, handshake-first with server/discover
fallback on -32022/-32601), stateless (discover-first), legacy
(handshake only). Auto is handshake-first deliberately: zero extra
round-trips and zero behavior change for the entire existing server
fleet, while 2026-07-28-only servers now connect via the fallback.
All four transport call sites (stdio, SSE, new HTTP, legacy HTTP)
route through the one choke point, so the CLI/desktop probe path
inherits it too.
- SEP-2549 list caching: tools/list ttlMs/cacheScope hints are captured
during discovery and bound to the lazy-startup schema cache — TTL'd
entries expire and force a live re-probe; hint-less (pre-2026)
servers keep the never-expires behavior. Pagination continuation now
speaks both SDK generations (params= vs cursor=).
- SEP-837: OAuth client metadata declares application_type=native
(config-overridable), with a fallback for 1.x-era metadata models.
(RFC 9207 iss validation and SEP-2352 issuer-keyed credentials are
native to SDK 2.0's OAuthClientProvider — verified, no client-side
gap.)
- SEP-2577 deprecation posture: SamplingHandler docstring marks the
Sampling feature as upstream-deprecated (12-month window) — kept
fully functional, closed to new capability.
- Docs: `protocol` key in the MCP config reference.
Shorter, single-source ownership explanation for
_strip_hermes_owned_pythonpath (the code-level Check comments already
carry the per-branch detail; the docstring only needs the contract).
Same behavior, same coverage, less boilerplate (test file 1691 -> 1512
lines; PR diff unchanged in semantics).
Production (mechanical only):
- Extract _strip_hermes_owned_pythonpath_and_runtime_markers(): the three
builders (_make_run_env, _sanitize_subprocess_env, hermes_subprocess_env)
ran the identical strip-then-pop-markers sequence in the same order
(ordering is load-bearing for VIRTUAL_ENV validation); the helper makes
that explicit once instead of three times.
Tests:
- Non-owned preservation: 11 single-shape tests -> one parametrized matrix
(user/Nix/other-version/python2.7/pythonX.Y-contained/raw spelling/empty
component/empty PYTHONPATH) + one runtime-shaped matrix (other-version SP,
venv-SP descendant, repo direct child, repo deep child).
- Owned stripping: venv SP, repo root (independent parents[2] computation),
duplicates, all-owned key removal, mixed ordering -> one matrix.
- Builder integration: _make_run_env/_sanitize_subprocess_env/
hermes_subprocess_env venv-SP stripping -> one parametrized test;
same for the four PYTHONHOME builders (incl. build_subprocess_env).
- Junction: same-named non-owned negative control now covers both the
configured-root location and an unrelated location; shared
_physical_repo_root helper; profile resolution matrix (root->named,
profile-shaped->named no nesting, profile-shaped->default, custom root).
- Every independent proof preserved: home-level junction, repo-level
junction, profile interaction, negative identity control, uv-base lexical
VIRTUAL_ENV, validated/unrelated VIRTUAL_ENV, no-scrub escape hatch,
#84500 same-env/external-env composition, PYTHONHOME removal, real
Windows-only semantics, POSIX fail-closed backslash paths.
Second real-world topology reported and confirmed on native Windows 11:
the repository itself is a cross-drive junction (D:\hermes\hermes-agent ->
C:\...\hermes-agent) under a real HERMES_HOME directory. The editable
import spelling resolves to the physical location, so _hermes_repo_root is
physical while the launcher writes the lexical spelling into PYTHONPATH.
The home-relative mapping cannot express a cross-drive link (commonpath
raises on different drives), so the lexical repo root survives stripping;
and with the repo alias missing, a lexical VIRTUAL_ENV
(D:\hermes\hermes-agent\venv) also fails _validated_runtime_venv, so the
venv site-packages survives too (uv-base gateway: both entries survive).
Fix: after the existing home/profile-root mapping, try the single
deterministic candidate <lexical root>/<repo dirname> for every trusted
home candidate (configured home, plus the profile root when the configured
home is a profile path) and accept it only when strict resolve proves it is
the exact physical repo root (fail-closed: missing paths, real directories
that are not the known repo, and unrelated spellings are never aliased).
This also re-enables the VIRTUAL_ENV validation for lexical venv spellings,
so uv-base gateway site-packages cleanup follows the repo alias.
Tests: repo-level junction positive + negative control (same-named real
directory preserved), profile-home + repo-level junction combination,
lexical VIRTUAL_ENV validation after recovery (root + site-packages
stripped, user entries kept), and a no-provenance lookalike preserved.
The execute_code composition test now compares composed paths with
os.path.normcase so a Windows case-only spelling difference (resolve() vs
abspath() casing) can never fail the composition contract.