Commit Graph

231 Commits

Author SHA1 Message Date
Teknium a0f15d1f2a refactor(mcp): _running_loop helper for the loop-up check 2026-09-02 16:58:35 -07:00
Teknium 6517b2d1c4 refactor(mcp): rewrap re-export blocks (formatting only) 2026-09-02 16:18:43 -07:00
Teknium b9ec602392 refactor(mcp): update mcp_tool module map for the new siblings 2026-09-02 16:15:11 -07:00
Teknium a09512c3ad refactor(mcp): move loop start/stop + reconnect signalling into mcp_tool_loop.py (writes via origin module) 2026-09-02 16:14:37 -07:00
Teknium 7be1f2abe0 refactor(mcp): table-driven SDK symbol binding in _ensure_mcp_sdk 2026-09-02 16:01:09 -07:00
Teknium dc1401425d refactor(mcp): drop 33 round-1 re-export shims nothing outside the group imports 2026-09-02 15:58:35 -07:00
Teknium d4602539d8 refactor(mcp): move discovery lock + loop scheduling into mcp_tool_loop.py 2026-09-02 15:52:38 -07:00
Teknium 46697da751 refactor(mcp): move connect/lazy-start/discovery/public API into mcp_tool_discovery.py 2026-09-02 15:50:54 -07:00
Teknium cbecd7d0c6 refactor(mcp): lift MCPServerTask lifecycle into MCPServerRunMixin; split run() into branch helpers 2026-09-02 15:46:06 -07:00
Teknium c1f8af1e86 refactor(tools/mcp): split mcp_tool.py into transport/lifecycle/schema/handlers/... sibling modules; compact watchdog and schema cache 2026-09-02 14:44:15 -07:00
Teknium 8e4366d358 fix(tools): freeze tools[] across agent-cache eviction; make /reload-mcp the re-probe hatch
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:

* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
  rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
  agent.tools from live probes with no predecessor to preserve. Persist
  the session's resolved tool-name order in a new `sessions.tool_names`
  JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
  alongside the system prompt and re-pinned on every published refresh
  (so /reload-mcp and compaction naturally reset it; /new mints a new
  row). On restore-for-existing-session the fresh definitions are folded
  onto the saved order via the SAME `_merge_preserving_prefix` helper —
  a probe-flipped tool is carried forward from the registry schema, a
  deregistered one dropped, new tools appended at the tail.

* /reload-mcp (CLI, gateway, TUI RPC) now also calls
  `reprobe_tool_availability()` — drops the check_fn verdict cache and the
  get_tool_definitions memo — so a user can consciously pick up a
  credential/daemon that appeared mid-session. Docs updated.
2026-09-02 07:22:59 -07:00
joaomarcos 65b0f00002 fix(agent): stop the between-turns tool refresh from forking the cached prefix
The per-turn MCP refresh re-derives `agent.tools` from live availability and
publishes the result wholesale. Two kinds of bytes move as a result:

* a tool whose `check_fn` merely flapped (headless browser probe, expired
  credential, docker blip) disappears from the array, and
* a late-landing MCP tool splices into sorted position, which can be index 0.

Providers that render `tools` ahead of the messages re-prefill the entire
history behind any moved byte, so either case costs a full re-prefill of the
session — the measured 2% cache hit in #100336. The caller's own comment
claimed the refresh "only ever extends a fresh request prefix"; it did not.

`refresh_agent_mcp_tools(..., preserve_prefix=True)` makes that claim true.
The live order becomes authoritative: existing tools keep their slot (fresh
schemas still land), a tool that is still registered but momentarily
unavailable is carried forward, a tool that genuinely left the registry is
still dropped, and new tools are appended at the tail. Explicit `/reload-mcp`
and the compaction boundary keep the plain rebuild.

Refs #100336
2026-09-02 07:22:59 -07:00
Teknium ee0e234a2c fix(gateway): discover and reload MCP servers per profile under multiplex
A multiplexed gateway ran `discover_mcp_tools()` once, unscoped, at boot
and again on `/reload-mcp`, so only the launch profile's `mcp_servers`
ever connected; secondary profiles' servers never registered, and a
`/reload-mcp` from any profile tore down every profile's connections.

- `_discover_gateway_mcp_tools()`: under multiplex, run discovery once per
  served profile inside `_profile_runtime_scope`, carried into the
  executor via `copy_context()` (same shape as
  `_run_in_executor_with_context`). Single-profile path unchanged.
- `_execute_mcp_reload()`: enter the requesting profile's scope when the
  caller (e.g. button-confirm callback) did not; shut down / rediscover /
  report only that profile's servers; refresh only that profile's cached
  agents.
- `shutdown_mcp_servers(scope=)`: scoped teardown keyed by the new
  `_server_scope_keys` ownership map; leaves the shared MCP loop running
  while other profiles' servers are live. Unscoped call keeps the full
  historical behavior.
- MCP tools register into the owning profile's registry overlay
  (`registry.register(scope=...)`), and `registry.deregister()` gains a
  matching `scope=` kwarg. Plugin callers still cannot name another
  profile's scope; the plugin-vs-global guard is unchanged for them.

Fixes #95518

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Kong <mgongzai@gmail.com>
Co-authored-by: roraag <232666910+roraag@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
codexbt 86fa1fcd4f fix(mcp): treat silent ping drop as unsupported rather than dead transport (Closes #97245)
A stdio server that never answers the optional ping (no -32601, no
response at all) produced a bare TimeoutError that _keepalive_probe
classified as a dead transport, tearing down and respawning a healthy
subprocess on every keepalive tick. On a first ping timeout, confirm with
list_tools before declaring death; if it answers, latch _ping_unsupported
and use list_tools from then on. If both fail, propagate as before.
2026-09-01 23:56:41 -07:00
NATHAN Menkin b828624479 fix(mcp): respawn and retry once when a stdio child died
A gateway restart kills every MCP stdio subprocess. An agent session that
outlives the restart still holds a handle to the dead child, so its next
tool call fails in 0.00s -- before anything reaches the network -- while
the subprocess is respawned seconds later. Cron runs spanning a restart
lose tool calls silently.

The #81995/#95626 machinery already detects the dead child and signals a
reconnect; it just never waits for it, so the caller eats the failure.
Both fast-fail sites now raise _StdioChildExited, and the handler respawns
the transport and retries the call once before any error reaches the model.

Retrying here cannot hot-cycle respawns: the handler never spawns anything.
It sets _reconnect_event (one signal per call, as before) and waits for the
server task to publish a fresh session, so spawn frequency stays governed by
run()'s rapid-drop budget (#62212). The retry is single-shot -- a child that
dies again immediately reports and stops, and a genuinely broken server
still parks with its tools deregistered.

The error text no longer claims a timeout. "failing the call fast instead of
waiting 300s" described a healthy remote backend as a timing problem and
sent an afternoon's investigation into the wrong system.

Verified on macOS against a real stdio subprocess, not only unit tests:
- SIGKILL the child of a live session (what a restart does to it), then
  call again: 0.00s error before, 0.51s success after.
- Child that exits on every tool call: 6 spawns across 8 calls, budget
  exhausted, parked, tools deregistered -- no respawn loop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 23:28:39 -07:00
Teknium 696d854b97 fix(mcp): stdio child PID snapshot reads every thread's /proc children
/proc/<pid>/task/<tid>/children is per-thread. stdio_client() spawns the
MCP subprocess from the background loop thread, so the main-thread-only
read returned an empty set on Linux and _stdio_child_pids/_stdio_pids
never tracked the child: the #81995 dead-child fast-fail, the #96452
respawn signal and the killpg shutdown sweep were all no-ops. Union the
children of every task instead.
2026-09-01 23:28:39 -07:00
kshitijk4poor 83b81fc6db test(mcp): pin the non-calling watcher probe + tighten to iscoroutinefunction
Follow-up to the salvaged #96044 hunk: drop the 'or callable(...)' arm —
callable(MagicMock) is True, which would have flipped stubbed sessions
into the fast-fail race the surrounding comment explicitly routes to the
plain-await path. inspect.iscoroutinefunction alone reproduces the old
isawaitable(call) split exactly (real async def / AsyncMock -> race,
MagicMock -> plain await) without creating the leaked coroutine.
2026-09-01 22:11:01 -07:00
loulanyue fe3e5dcb9e fix(mcp): stop leaking an unawaited watcher coroutine in the _watch_ok probe
The fast-fail gate probed the stdio child watcher by CALLING it —
inspect.isawaitable(_watch_children()) — creating a fresh coroutine on
every stdio MCP tool call that was never awaited (RuntimeWarning spam +
gc churn). Inspect the function instead of invoking it.

Salvaged (unique hunk only) from PR #96044; the bundled
_stdio_children_dead polarity fix was already on main via #94339.
2026-09-01 22:11:01 -07:00
Teknium c30ac90a92 feat(compaction): rebuild dynamic tool schemas at the compaction commit boundary — forever-sessions finally pick up config changes (#97073) 2026-08-28 04:01:05 -07:00
kshitij 8966b0a700 review: tighten pre-call gate comment; drop redundant _ReadyAdapter test stub
Both from the simplify pass: the comment kept only the ownership-relevant
rationale (incl. the no-double-bump note); the test's _ReadyAdapter was a
verbatim delegate around threading.Event — the exercised paths only call
is_set/clear/set, so the bare Event is behaviorally identical.
2026-08-27 21:36:35 +05:30
kshitij 8aae2ea539 fix(mcp): signal reconnect from the mid-call fast-fail site too + regression tests
Widen the contributor's pre-call reconnect signal to the sibling site:
when the stdio subprocess dies mid-RPC the watcher race fast-fails, but
nothing cleared server.session, so the server stayed dead until the idle
keepalive probe noticed. Signal the reconnect there as well.

Also drop the explicit _bump_server_error at the pre-call gate: the
returned error payload already flows through the handler's JSON parse,
which bumps the breaker once — the explicit bump would double-count.

Two regression tests pin both sites (reconnect signaled exactly once,
no RPC attempted on a dead transport, single breaker bump).
2026-08-27 21:36:35 +05:30
deadczarvc 2663117f72 fix(mcp): correct inverted liveness check in _stdio_children_dead
_stdio_children_dead() returned True ('all children dead') when a child
process was ALIVE — the liveness predicate was inverted:

    for pid in pids:
        if not psutil.pid_exists(pid):
            continue   # dead — skip
        return True     # BUG: an ALIVE child reported as 'all dead'

Consequence: every tools/call on a stdio MCP server with a healthy
subprocess failed instantly (~0.01-0.3s) with
'MCP stdio subprocess has exited; failing the call fast' (#81995
fast-fail path), while hermes mcp test kept passing (it never reaches
tools/call). Servers appeared dead regardless of restarts.

Fix: return False as soon as one tracked child is alive; True only when
every child has exited:

    for pid in pids:
        if psutil.pid_exists(pid):
            return False  # at least one child alive
    return True

Also: on a genuinely-dead stdio (session object still present), signal
reconnect instead of a bare fast-fail TimeoutError, so the manager
respawns the subprocess instead of stranding the call slot.

Verified: alive child -> False, all exited -> True, no tracked pids ->
False (unknown, don't fail fast).
2026-08-27 21:36:35 +05:30
liuhao1024 ef46ec03e1 test(mcp): absorb watcher-consumer and fail-open liveness cases from #94521/#94661
Apply-ready delta distilled by @andrexibiza: deterministic watcher-consumer
tests (watcher times out while a child is alive, resolves when all are dead),
psutil-unavailable fail-open pin, and probe-failure fail-open handling in
_stdio_children_dead (unknown is never proof that every child exited).
Local: 8 passed on tests/tools/test_mcp_stdio_children_dead.py
2026-08-27 14:08:20 +05:30
liuhao1024 98fce8e52d fix(mcp): un-invert the stdio children liveness check (#94335)
_stdio_children_dead returned True ('all children dead') on the first LIVE
pid — the intended False was dead code right below it. Every spawn path
that captures child PIDs (observed in hermes -z oneshots) then failed the
#81995 pre-call fast-fail with 'TimeoutError: MCP stdio subprocess ... has
exited' on every tools/call while the subprocess was demonstrably alive.
Long-lived gateway/dashboard sessions were unaffected only when
_stdio_child_pids was empty (the not-pids short-circuit).

Return False on the first live pid and drop the unreachable line.
2026-08-27 14:08:20 +05:30
Teknium e1e72f109c fix(mcp): register stdio MCP helper children in the spawn ledger and reap orphans (#61514)
Stdio MCP helper subprocesses (npx/binary servers) never import Hermes
code, so they could not self-register in the machine spawn ledger and an
unclean parent exit left them running invisibly forever.

- process_identity.register_child(pid, purpose): ledger mirror of
  register_self for spawned children — records the CHILD (pid,
  create_time) with this process as spawner. Refuses pid-only entries a
  PID reuse could forge. Writes go through the single _append_entry
  path under _LEDGER_LOCK (prune + atomic tmp/replace unchanged).
- 'mcp-helper' added to REAPABLE_PURPOSES so the updater's
  _ledger_reapable_backend_pids rung flows helpers through its existing
  spawner_is_dead gate (live spawner => never reaped).
- tools/mcp_tool.py: best-effort register_child(pid, 'mcp-helper') at
  the post-spawn PID capture; never breaks MCP startup.
- reap_orphaned_mcp_helpers(): startup sweep mirroring
  _reap_orphaned_desktop_local_serves but ledger-driven — kills only
  helpers whose recorded spawner is PROVABLY dead, with a create_time
  re-check at kill time. Wired next to the desktop serve reap in
  web_server.py.
2026-08-26 17:42:58 -07:00
chelsealong 8f517f5ca6 chore: address AI-review nits on _ever_connected fix
Drop the try/except AttributeError guard in the new regression test
now that the slot is always defined, and note in the run() comment
that _ever_connected is set once and never cleared.
2026-08-26 08:40:07 -07:00
chelsealong c7673f322b fix(tools): stop treating a post-registration reconnect drop as an initial-connect failure
MCPServerTask.run() used `_ready.is_set()` to tell a genuine first
connection attempt from a later reconnect. `_ready` is cleared on every
reconnect cycle, so once a server has already registered its tools and
then drops (keepalive failure, transient TaskGroup exit, etc.), the next
failed reconnect attempt is misclassified as "never connected" and burns
the 3-attempt initial-connect ladder instead of the 5-attempt reconnect
budget, parking the server much sooner and logging "failed initial
connection after 3 attempts" even though tools were already registered.

Add a sticky `_ever_connected` flag, set once alongside `_ready.set()`
right after a successful `_discover_tools()` call and never cleared, and
gate the initial-vs-reconnect branch on it instead.

Fixes #94654
2026-08-26 08:40:07 -07:00
Jan-Stefan Janetzky fb1ec36a4b fix(mcp): treat tools.include: [] as an explicit empty whitelist
_normalize_name_filter([]) returns an empty set, which is falsy, so
_should_register fell through to "no filter" and registered every tool
— the exact opposite of what _apply_tool_selection wrote when the user
unchecked everything in the install checklist ("contributes nothing
until reconfigured"). Whitelist mode is now keyed on the include key
holding a valid filter shape (str/list/tuple/set) rather than on set
truthiness, at both the live-discovery and cached-manifest sites.
Invalid include values keep the old warn-and-ignore behaviour.
2026-08-25 04:21:37 -07:00
kshitijk4poor 786f37071a fix(mcp): psutil.pid_exists for stdio children liveness — Windows footgun (#85125 CI) 2026-08-25 03:30:11 +05:30
kshitijk4poor 2f33833de8 fix(mcp): recover poisoned connections + fail fast on dead stdio transports (#85125 3b)
Fixes the four poisoned-connection classes (#81051, #77765, #84132,
#81995) with the SuspectableBackend cheap-mark/lazy-verify contract:

- mark_suspect/ensure_healthy protocol (agent/deadline.py): noticing a
  poisoned state never does I/O; the NEXT caller pays once for a health
  probe that clears the suspicion or forces a reconnect. A single
  teardown-vs-keepalive race or auth-lock corruption can no longer park
  a connection permanently — park stays reserved for genuinely
  exhausted reconnect budgets.
- keepalive failure marks the connection suspect before requesting
  reconnect; the next tool call probes and recycles if unhealthy.
- auth-classified permanent failures on a previously-proven session get
  a suspect+reconnect path instead of an immediate park.
- fast-fail (#81995): stdio child pids are tracked at spawn and an
  in-flight RPC races a child-watcher task, so a dead subprocess fails
  the call immediately with a retryable timeout instead of riding out
  the full 300s. Deliberate teardown/reconnect also fails in-flight
  calls now instead of leaving them attached to a dying transport.

Dispatch-boundary hardening for test doubles: stubbed sessions
(MagicMock/non-awaitable call_tool, absent child-watcher) fall back to
the exact pre-change inline-await semantics, so only real transports
gain the race guard.

Salvage credit: in-flight approach from #73377 (@luijoc, wedged
transport recovery) and #48069 (@arminanton, keepalive/in-flight
interaction); both PRs' bases predate main's current park/reconnect
architecture, so this is a fresh implementation of their contracts.

Tests: tests/tools/ -k mcp = 639 passed (was 22 new failures during
development; final tree zero).
2026-08-25 01:34:55 +05:30
kshitijk4poor 7dde1b8b0b fix(mcp): resolve tool-call timeouts via the unified deadline layer (#85125 2g)
Both readers of the per-server MCP tool timeout (the connection's run()
and the cache-path registration) read config.get("timeout", 300) as
their own private resolution. Route them through _resolve_tool_timeout:
per-server mcp_servers.<name>.timeout still ALWAYS wins (most specific),
then timeouts.mcp.tool_call from the unified timeouts: section, then
the unchanged 300s default. Values pass through resolve_timeout's
platform clamp; resolution failure falls back to the historical default.

Default-behavior invariance pinned by contract tests (nothing
configured -> exactly 300, per-server beats section, section beats
default, invalid/failed resolution falls back).
2026-08-24 17:12:15 +05:30
Teknium 09e657793e feat: MCP tool results spill at 50K and carry upstream-elision warnings
Composio-style MCP servers return un-paginated 22-47K-char payloads that
sail under the generic 100K per-result spillover threshold, bloating
context and ballooning per-turn reasoning time on long conversations.
Competitors cap harder (OpenCode/pi 50KB, Claude Code 30K, Codex ~10K
tokens). Three changes:

- mcp_* tools spill at a tighter 50K default (BudgetConfig.mcp_result_size,
  config-overridable via tool_budget.mcp_result_size_chars; pinned and
  per-tool overrides still win; capped by the context-scaled default).
- The persisted-output preview now teaches recovery: page the saved file
  with read_file or process with execute_code instead of re-requesting the
  same data from the remote API.
- Untrusted/MCP string results are scanned (bounded, first 64KB) for
  provider-side elision markers ('...N more items', "has_more": true,
  'saved to sandbox', data_preview) and get ONE cache-safe incompleteness
  notice appended at result-construction time, before untrusted wrapping —
  so the model stops treating provider-elided enumerations as complete.
- Hard 2M-char allocation cap in mcp_tool.py (text, error, and
  structuredContent paths) so a pathological multi-MB server payload is
  bounded before it propagates, while ordinary large results reach
  spillover intact. Distilled from #56060/#56072/#56511 (issue #56059);
  supersedes their 50K lossy truncation with spillover-friendly semantics.

Docs: configuration.md spillover-budget section + cli-config.yaml.example.

Co-authored-by: Stoltemberg <215755014+Stoltemberg@users.noreply.github.com>
Co-authored-by: AlexFucuson9 <295703459+AlexFucuson9@users.noreply.github.com>
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-08-19 16:31:16 -07:00
Teknium 382060f022 feat(mcp): speak the 2026-07-28 stateless protocol
Phase 2 of the MCP 2026-07-28 migration (#69931), on top of the SDK 2.x
migration (#88180):

- Protocol-era negotiation (_negotiate_session): per-server `protocol`
  config key — auto (default, handshake-first with server/discover
  fallback on -32022/-32601), stateless (discover-first), legacy
  (handshake only). Auto is handshake-first deliberately: zero extra
  round-trips and zero behavior change for the entire existing server
  fleet, while 2026-07-28-only servers now connect via the fallback.
  All four transport call sites (stdio, SSE, new HTTP, legacy HTTP)
  route through the one choke point, so the CLI/desktop probe path
  inherits it too.
- SEP-2549 list caching: tools/list ttlMs/cacheScope hints are captured
  during discovery and bound to the lazy-startup schema cache — TTL'd
  entries expire and force a live re-probe; hint-less (pre-2026)
  servers keep the never-expires behavior. Pagination continuation now
  speaks both SDK generations (params= vs cursor=).
- SEP-837: OAuth client metadata declares application_type=native
  (config-overridable), with a fallback for 1.x-era metadata models.
  (RFC 9207 iss validation and SEP-2352 issuer-keyed credentials are
  native to SDK 2.0's OAuthClientProvider — verified, no client-side
  gap.)
- SEP-2577 deprecation posture: SamplingHandler docstring marks the
  Sampling feature as upstream-deprecated (12-month window) — kept
  fully functional, closed to new capability.
- Docs: `protocol` key in the MCP config reference.
2026-08-17 03:03:35 -07:00
elphamale 23a86594cc fix(mcp): read the elicitation schema under the SDK's real field name
`ElicitationHandler` read `params.requested_schema`, but on the pinned
`mcp==1.28.1` the model field is spelled `requestedSchema`. The getattr
always missed and returned its `{}` default, so
`_format_elicitation_schema_summary` took its no-properties branch and the
approval prompt collapsed to the generic

    Approval requested by MCP server '<name>'.

for every request. The field names, types, and descriptions the summary
exists to surface never reached the user, so an elicitation asking for a
card number rendered identically to one asking for a nickname — consent
without the substance of what was being consented to.

Read both spellings rather than just correcting to the 1.x name: mcp 2.0
renames this field to `requested_schema` (it renamed every model field to
snake_case and kept camelCase only as a serialization alias, which
pydantic does not expose to attribute access), so a dual read is correct
on either SDK generation and does not go wrong again on the next bump.
Verified against real 1.28.1 and 2.0.0 installs.

Every existing test in tests/tools/test_mcp_elicitation.py builds a
duck-typed `SimpleNamespace` stand-in, which carries whatever field name
the test wrote and therefore cannot detect a mismatch with the real model.
Add one test that constructs the actual `ElicitRequestFormParams` and
asserts the requested field name reaches the consent description; it fails
on the unfixed tree. The cheap stand-ins are left alone elsewhere.

Found while porting the tree to the mcp 2.x SDK in #76736, but independent
of it: this reproduces on the current pin with no other changes, #76736
does not touch this line, and the two branches merge cleanly in either
order.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:56:50 -07:00
elphamale 77ed1bbf40 fix(mcp): seed MCP-Protocol-Version from the handshake version, not the latest
The HTTP transport seeded `MCP-Protocol-Version` from LATEST_PROTOCOL_VERSION,
which on mcp 2.x is 2026-07-28 — a revision that replaced the `initialize`
handshake with a per-request envelope. But this transport connects through
`ClientSession.initialize()`, which sends LATEST_HANDSHAKE_VERSION (2025-11-25)
in the body. Header and body therefore disagreed by construction, and a
conforming 2.x server honours the header: it routed the request onto its
per-request-envelope ladder and rejected the legacy body with

    params._meta is missing the required envelope key(s):
    io.modelcontextprotocol/protocolVersion,
    io.modelcontextprotocol/clientCapabilities

Observed against a live MCP endpoint, and confirmed by probing the same
endpoint three ways: the header at 2026-07-28 is rejected, at 2025-11-25 it
succeeds, and with no header at all it succeeds.

Third defect in this migration from one cause: the 2.x bump changed what an
existing constant *means* without revisiting its uses. The header seed was
written when LATEST_PROTOCOL_VERSION was 2025-03-26 and was correct then.

Seeded from LATEST_HANDSHAKE_VERSION, imported with a fallback to
LATEST_PROTOCOL_VERSION for SDKs predating the split, where the two are the
same thing and header and body agree either way. An explicitly configured
header still wins — that override is why servers demanding a specific revision
can have one, and a test pins it.
2026-08-16 23:26:10 -07:00
elphamale 2e1d724e3e fix(mcp): accept both SDK generations' streamable-HTTP transport arity
`streamable_http_client` yields `(read, write, get_session_id)` on mcp 1.x and
`(read, write)` on 2.x. `_run_http` unpacked a fixed 3-tuple, so on 2.x every
HTTP and SSE MCP server failed its handshake with `ValueError: not enough
values to unpack (expected 3, got 2)` and parked after exhausting its retry
ladder. Only stdio servers kept working.

This is the same defect as the import gating fixed earlier in this branch, one
layer further in. That fix's own comment claimed reaching
`streamable_http_client` was "the path that does work" — reaching it was
necessary and not sufficient, and the comment asserted the half that was never
exercised. Corrected along with the code.

Unpacked positionally rather than by arity, since this file deliberately
supports both SDK generations and `get_session_id` was never used here.

The reason this survived review is worth the test it now has: the existing
coverage in test_mcp_client_cert.py fakes the transport with a 3-tuple, so it
encoded 1.x's shape into the assertion and passed on 2.x regardless. The new
test drives `_run_http` once per arity the supported SDK range actually yields,
and asserts the streams handed to ClientSession are the first two — positional,
because 1.x's third element is not a stream. Verified it fails on the 2.x case
without this change.

Found while pointing a real HTTP MCP server at a live deployment running this
branch: the server parked at startup and no tool from it ever registered.
2026-08-16 23:26:10 -07:00
elphamale 11a9dcf567 feat(mcp): migrate to the mcp 2.x SDK
mcp 2.0.0 implements MCP revision 2026-07-28 and makes three breaking
changes Hermes sits on top of: `mcp.server.fastmcp` is gone, every model
field is renamed to snake_case (camelCase survives only as a
serialization alias, which pydantic does not expose to attribute
access), and the SDK's own HTTP stack moved from `httpx` to `httpx2`.

Bump the pin across the dev/mcp/computer-use extras and port the tree:

- `mcp_serve.py` and `agent/transports/hermes_tools_mcp_server.py` move
  from `FastMCP` to `mcp.server.MCPServer`, which has the same
  decorator/add_tool surface. The hermes-tools server already
  synthesised `__signature__` from Hermes' JSON Schema, which is exactly
  what 2.0's `add_tool` reads.
- SDK model reads go through `mcp_field(obj, snake, camel)`, which reads
  both spellings. A single-spelling read fails *silently* on the other
  generation — empty tool schemas, dropped structured content, tool
  results vanishing from sampling conversations — and `mcp` is an
  optional extra users install at their own version.
- `sdk_httpx()` resolves the httpx flavour from the SDK's own transport
  module, so objects handed to `streamable_http_client`, the `sse_client`
  factory, and the OAuth metadata helpers come from the module the
  installed SDK actually imports.
- HTTP support is gated on either streamable-HTTP entry point, not just
  the deprecated alias 2.0 removed.
- OAuth: `OAuthClientProvider` lost its `timeout` argument (the
  configured `oauth.timeout` now bounds the callback waiter's own poll
  loop, where the browser round-trip was always awaited), and
  `callback_handler` must return `AuthorizationCodeResult` rather than a
  tuple. 2.0 also validates the RFC 9207 `iss` parameter, so the
  callback handler and paste fallback capture it.

`mcp`/`mcp-types` 2.0.0 are inside the 14-day `exclude-newer` window, so
two narrow `exclude-newer-package` entries unblock `uv lock`, annotated
for removal on or after 2026-08-11. `httpx2` needs no exemption: 2.7.0 is
already outside the window and satisfies mcp's floor.

Refs #69931

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 23:26:10 -07:00
Teknium c031fec365 Port from MoonshotAI/kimi-code#2596/#2600: surface MCP tool-result _meta to the model, minus protocol-reserved keys
MCP tool results carry a server _meta mapping (exposed as .meta by the
Python SDK) alongside structuredContent. Servers return namespaced
machine-readable contracts there (validated payloads, browser-handoff
URLs); Hermes previously dropped the field entirely, so that data was
invisible to the agent.

Now _meta is included in the JSON tool output, after filtering
protocol-reserved keys per the MCP spec's key-name rules: a prefix is
reserved when a modelcontextprotocol or mcp label is followed by at
least one more label (modelcontextprotocol.io/..., tools.mcp.com/...).
Vendor namespaces with a trailing reserved word (com.example.mcp/...)
and unprefixed keys pass through. Non-serializable metadata drops the
extras rather than failing the call.
2026-08-16 22:08:18 -07:00
Teknium 8bbda8ff33 Port from block/goose#10746: strip invisible Unicode TAG chars from MCP content
Unicode TAG characters (U+E0000-U+E007F) render as nothing in terminals
and chat UIs but are fully visible to LLM tokenizers, making them an
ASCII-smuggling prompt-injection channel for untrusted MCP servers.

- tools/ansi_strip.py: new strip_unicode_tags() with fast path; unlike
  goose we preserve valid emoji tag sequences (U+1F3F4 base + tag spec +
  U+E007F cancel), so regional flags survive.
- tools/mcp_tool.py: applied at every MCP text ingestion point — tool
  result text blocks, embedded resource text, read_resource contents,
  get_prompt message content, and tool descriptions entering the schema.
- tests/tools/test_unicode_tag_strip.py: smuggled-instruction vectors,
  goose's test vector, emoji-tag-sequence preservation, ZWJ untouched.
2026-08-16 22:08:05 -07:00
PRATHAMESH75 d6f18cd7db fix(mcp): prefer server-native tool over generated utility on name collision (#87112)
An MCP server exposing a native tool named read_resource (or
list_resources/list_prompts/get_prompt) collided with the auto-generated
resource/prompt utility of the same name. The registration collision
handler flagged the pair as ambiguous and skipped BOTH entries, so the
server's own tool became silently unavailable on every gateway boot.

Resolve this specific native-vs-utility collision in favour of the native
tool: keep it and drop the shadowed utility, which is only convenience
sugar for servers that expose no such tool of their own. The conservative
skip-everything path still applies to genuinely ambiguous collisions (two
or more native tools normalizing to one name), which we cannot
disambiguate. Add a regression test covering the native-tool-wins path.

Fixes #87112
2026-08-16 01:52:26 -07:00
Teknium 4a2198bf51 fix: Windows MCP PATHEXT resolution + python3 -> python in cross-platform skills (#84429)
Two Windows agent-loop friction fixes:

1. tools/mcp_tool.py (#56536): shutil.which(cmd, path=env_path) reads
   executable extensions from the PARENT process PATHEXT, not the MCP
   subprocess env — a stdio MCP config supplying both PATH and PATHEXT
   could fail to resolve a command its own env can locate, and startup
   then got a bare command name. On Windows, when the first which() call
   misses and the config env carries PATHEXT (any key casing), retry the
   resolution with the config's PATHEXT temporarily applied.

2. skills/ + optional-skills/ (#50606): 42 SKILL.md files that declare
   platforms: [.., windows] used python3 in their command examples.
   python3 does not exist on native Windows (the toolchain probe in the
   system prompt reports python3=missing), so every copy-pasted example
   burned a failed agent turn before self-correction. Replaced the
   command word python3 -> python (python3-config / python3.x version
   strings untouched). python is the spelling that exists in every
   Hermes-managed environment (Windows native, uv-managed venvs on all
   three OSes); agents on POSIX hosts additionally see the probed
   toolchain line and adapt either way.
2026-08-12 02:43:28 -07:00
Teknium 55f9e472a0 perf(cli): sub-400ms warm startup — probe-mode check_fns, lazy MCP SDK, banner snapshot, parallel worktree add
Cold CLI time-to-banner was ~1.8s (hermes) / ~2.8s (hermes -w). The banner
path was paying for work the session doesn't need before first input:

- aux availability probes built REAL OpenAI/httpx clients (openai import
  ~0.3s + SSL context) just to answer check_fns. New aux_probe_mode()
  returns a cache-excluded stub; resolution policy unchanged.
- tools/mcp_tool imported the mcp SDK (~260ms, mcp.types pydantic model
  construction) at module import even with zero MCP servers configured.
  SDK import is now lazy behind _ensure_mcp_sdk(); _MCP_AVAILABLE is a
  find_spec probe so every existing gate/test keeps its semantics.
- banner blocked 500ms on the update-check prefetch; now waits 50ms and
  defers the warning line to a daemon thread (prints above the prompt).
- banner recomputed get_tool_definitions + skills scan + git state every
  launch; now snapshotted to ~/.hermes/cache/banner_snapshot.json keyed on
  (config.yaml, .env, checkout rev, toolsets) and replayed on warm launches
  with a background refresh. Agent tool list is still computed fresh.
- _resolve_active_context_length probed the Nous portal /models (~200ms
  network) per launch; the tool-search gate now prefers the on-disk
  context cache when present.
- schema reconciliation re-executed SCHEMA_SQL in a scratch SQLite DB
  (~85ms) per SessionDB(); the reference parse is now disk-memoized by
  DDL hash (live-DB diffing still runs every startup).
- bundled-skills sync (~120-170ms rglob/hash) moved off the startup path
  to a daemon thread; plugin discovery starts in the background and every
  synchronous consumer joins via discover_plugins().
- hermes_cli.auth imported httpx eagerly (~30ms); now a lazy proxy that
  test monkeypatching still reaches (setattr forwards to the real module).
- fast chat launch: unambiguous 'hermes'/'hermes chat' invocations skip
  building all ~40 subcommand parsers (bails to full dispatch on anything
  else, incl. container mode).
- -w path: git worktree add runs with checkout.workers=8 (0.6s→0.2s) and
  overlaps HermesCLI construction; --skills preload runs in the background
  and is folded in at agent init (finalize_preloaded_skills, same
  fail-loud contract for fully-unknown skill lists); stale-worktree prune
  moved off the banner path.

Warm results (PTY time-to-banner, 5-run): hermes 1.80s → 0.38-0.40s;
hermes -w -s hermes-agent-dev --yolo 2.82s → 0.57-0.69s.
2026-08-10 10:40:19 -07:00
Teknium 471baea520 feat(plugins): map portable Agent Plugins streamable-http entries into the native MCP runtime
Agent Plugins v1 packages with 'streamable-http' mcp.json entries now load
through Hermes' existing URL-based MCP client instead of being reported and
skipped. The stdio-only limitation was the agreed follow-up slice from
PR #81196.

Boundary rules from the v1 spec (§7.2.1) are enforced:
- URL must be absolute http(s), no user information, no fragment; plain
  HTTP only for localhost/loopback hosts.
- Configured package headers are never forwarded across a cross-origin
  redirect: translation marks entries strict_redirect_headers, and the
  redirect hook in the native runtime strips those headers (plus
  Authorization) whenever a redirect leaves the original origin. On mcp <
  1.24.0, where the client cannot hook redirects, such servers fail closed
  with an actionable upgrade message.
- Legacy 'sse' entries remain reported and skipped.

The redirect hook is extracted into a testable module-level factory
(_make_redirect_header_stripper); default behavior for native config
servers is unchanged (Authorization-only stripping).
2026-08-08 23:56:46 -07:00
Teknium a978f769b1 Inspired by Cursor: MCP config context variables (${userHome}, ${workspaceFolder}, ...) 2026-08-08 03:57:00 -07:00
Brooklyn Nicholson f99d291247 fix(mcp): let a server that 401s at startup come back after re-login
An auth failure on the very first connect returned out of the run loop
instead of parking. That ended the run task, and the task is the only
listener on _reconnect_event — so the server stayed dead for the life of
the process. `hermes mcp login`, a /mcp refresh, and the 300s self-probe
all had nothing left to wake, and the only cure was a full restart.

_classify_mcp_failure already calls 401/403 "permanent" and documents
that run() parks those immediately; the early return above it meant auth
was the one permanent failure that never got there. Park it with the
others and keep the tailored log line, now pointing at `hermes mcp
login <server>`.
2026-08-08 02:21:36 -05:00
GodsBoy ca78c6d7a6 feat(plugins): load portable agent components 2026-08-07 09:44:21 -07:00
Teknium c8369e37f4 feat(mcp): trust-tier gating for write-capable MCP tools via readOnlyHint
Adds a per-server `trust: full|untrusted` config key
(mcp_servers.<name>.trust). On an untrusted server, every write-capable
tool call — any tool whose discovery-time annotations do not carry
readOnlyHint=True — routes through the existing approval surface
(tools.approval.request_elicitation_consent, same lazy-import +
surface-routing pattern the MCP elicitation handler uses) before the RPC
fires. Denied/cancelled/errored approvals fail closed: the RPC never
runs, including the lazy first-use server spawn.

Design points:
- Classification happens at CALL TIME from metadata captured at
  DISCOVERY (_record_tool_trust_metadata in _register_server_tools and
  the lazy cache-registration path). No toolset/schema mutation, so the
  toolset stays byte-stable and prompt caching is preserved.
- readOnlyHint is a server-supplied HINT: on an untrusted server a lying
  server can at most skip approval for tools it claims read-only — it
  can never widen access. Trust tiering itself is operator config.
- Missing/malformed annotations => write-capable (fail closed).
- Unrecognized trust values => untrusted (fail closed); missing key =>
  full (backward compatible, documented in mcp-config-reference).
- The schema cache now persists readOnlyHint so lazy-registered servers
  gate identically on next startup without spawning.

Tests: tests/tools/test_mcp_trust_gating.py (11 tests, TDD red->green):
approval invoked + accept proceeds, deny/cancel blocks RPC, readOnlyHint
=true skips gate, trusted/unconfigured servers skip gate, explicit
readOnlyHint=false gated, approval exception fails closed, trust
normalization, discovery-time capture (SDK objects and cached dicts).

Ported from: cloudflare-os classifyTool() (Apache-2.0), corroborated by
Claude Cowork (idea-level).
2026-08-07 08:58:32 -07:00
Teknium 37cc999926 feat(mcp): collapse const-only anyOf/oneOf unions to property enums
MCP servers generated from Rust/TypeScript union types commonly emit
closed value sets as const unions:

    {"anyOf": [{"const": "red"}, {"const": "green"}, {"const": "blue"}]}

Strict tool-calling backends reject or mishandle these; the equivalent
property-level enum form is universally supported. Add
collapse_const_unions() to tools/schema_sanitizer.py and wire it into
the _normalize_mcp_input_schema discovery pipeline after the nullable
strip.

Rules:
- Collapse only when EVERY non-null branch is a pure const of the same
  primitive type (bool never merges with integer).
- Mixed unions, non-uniform const types, and mismatched declared types
  pass through untouched.
- A single {"type": "null"} branch is tolerated: consts -> enum,
  null -> nullable: true hint (matches strip_nullable_unions, which
  leaves null+multi-const unions alone by its one-non-null-branch rule).
- Outer title/description/default/examples carried onto the replacement.
- Deterministic, branch-order-preserving, non-mutating — applied at
  discovery only, so schemas stay byte-stable per conversation.

Ported from: block/goose tool_schema_normalize.rs (Apache-2.0)
2026-08-07 08:58:25 -07:00
Teknium 9fad45fcda feat(kanban,mcp): orphaned-card reconciliation + per-server MCP identity header
Two small config-gated features:

1. Kanban orphaned-card reconciliation (kanban.reconcile_orphans, default
   true, config.yaml): a running card with broken claim bookkeeping
   (claim_lock or claim_expires NULL — crash mid-claim, manual SQL, DB
   restore) is invisible to all existing recovery paths
   (release_stale_claims requires claim_expires NOT NULL,
   detect_crashed_workers requires host-local lock + pid,
   detect_stale_running is config-disabled by default) and shows Running
   forever. New reconcile_orphaned_running() pass in kanban_db.py runs
   each dispatch_once tick: requeues orphans to ready with an explanatory
   comment, closes any leaked run, emits a 'reconciled' event, and defers
   when the recorded PID is still alive on this host (never requeue
   beside a live worker). Surfaced via DispatchResult.reconciled_orphans.

2. Per-server MCP identity header (mcp_servers.<name>.identity_header,
   config.yaml): optional {name, value_from: static|profile, value}
   mapping; the header is attached to that server's HTTP/SSE transport
   requests. 'static' sends the config value; 'profile' resolves the
   active Hermes profile name once at connect time (no per-call
   mutation). Explicit per-server headers of the same name (any casing)
   win. Invalid blocks warn-and-ignore; stdio servers warn-and-ignore.

Tests: tests/gateway/test_kanban_reconcile_orphans.py (9),
tests/tools/test_mcp_identity_header.py (13), all written first (RED)
then implemented (GREEN). No new HERMES_* env vars.

Inspired by: openai/symphony tracker reconciliation (Apache-2.0) +
Poke per-user MCP identity (idea-level).
2026-08-07 08:58:20 -07:00
kshitij ebf967ff2c polish(mcp): simplify-pass folds on the lazy-startup salvage
Five review findings folded:
- schema cache writes via utils.atomic_json_write (fsync; was bare
  tmp+replace), file moved to cache/mcp_schema_cache.json with 0o600
  (sibling precedent: registry discovery cache)
- phantom-tool reconciliation: after a lazy server's first-use connect,
  cached tools the live server no longer offers are deregistered (were
  permanent registry ghosts burning circuit-breaker strikes on every
  'Unknown tool' round-trip); stale fingerprint logged
- cache-load path now runs _scan_mcp_description like the eager path
  (cache file is user-writable JSON; defense-in-depth)
- write-through skips the disk rewrite when the entry is unchanged
  (a flapping stdio server was rewriting byte-identical JSON per
  revival)
- _lazy_server_fingerprints no longer write-only dead state (consumed
  by the reconciliation logging)

444 mcp tests green (440 pre-fold + 4 new guards); phantom-dereg and
write-skip mutation-checked.
2026-08-03 14:24:37 +05:30