Commit Graph

35745 Commits

Author SHA1 Message Date
teknium1 16ec8a35eb test(api-server): keep the behavioural run-output test, drop the source-reading guard
The salvaged commit shipped two tests: a disconnect-mid-stream test that reads
the reply back through GET /v1/runs/{run_id}, and a structural guard that
opened api_server.py and searched its text for `output=`. Tests that read
source are change-detectors, not invariants; the behavioural test already
fails the moment `output` disappears from the terminal write, so the guard is
gone. The inline comment shrinks to the WHY (parity with /v1/runs, the
recovery path) and drops the narrative.
2026-09-15 19:24:21 -07:00
John Paul Soliva bb89074ef2 fix(api-server): the session stream records its reply text, like /v1/runs does
`POST /api/sessions/{id}/chat/stream` mints a run_id and writes run status,
but its terminal write omitted `output=` — the reply text went only onto the
SSE queue. `POST /v1/runs` has always recorded it.

The asymmetry costs a caller its answer. A client whose socket dies mid-turn —
a sleeping laptop, a dropped WiFi link, a peer DM over a flaky LAN — sees the
run reach "completed" through GET /v1/runs/{run_id} and has no way to learn
what the agent said. Worse, the disconnect path interrupts the agent, which
unwinds and *returns* a partial result, so that partial answer is recorded as
a clean "completed" and then discarded: indistinguishable from a run that
produced nothing, and equally unrecoverable.

Record `output` on this route too. Terminal statuses are already retained for
_RUN_STATUS_TTL (3600s), so the existing GET /v1/runs/{run_id} becomes a
recovery path for any client that loses its stream, without a new endpoint,
without touching the run registries, and with no change to the streaming
contract.

Deliberately nothing else: the detached turn stays out of _active_run_tasks
(it is already counted via _inflight_agent_runs, and a task entry would
double-count it in the shutdown drain), and the run stays out of
_run_streams_created, so the orphan sweeper's 300s reap still cannot see it.

Tests: a disconnect mid-stream now leaves the reply readable both in the run
record and through _handle_get_run; a structural guard asserts BOTH routes
still pass output=, anchored on the status write rather than the SSE payload's
"completed": True key, so the asymmetry cannot quietly return.
2026-09-15 19:24:21 -07:00
teknium1 51a2f4878f test: pin the non-SDK facade gate alongside the escape hatch
The MoA aggregator and test stand-ins never merge extra_body; the bypass
must hand them the kwargs untouched or the conversation would be sent
empty. Fold that control into the existing rail test (still two tests).
2026-09-15 19:23:53 -07:00
kshitijk4poor a12b3c7aa3 refactor(agent): import the transform bypass from its defining module, no re-export shim
Internal moves get no compat aliases (root AGENTS.md); codex_runtime and
auxiliary_client import bypass_sdk_request_transform from agent.sdk_transform_bypass.
2026-09-15 19:23:53 -07:00
kshitijk4poor af7b60e8ec fix(agent): keep moved chat fields as slot placeholders so the wire body is byte-identical; one shared escape hatch
The cherry-picked helper deleted 'tools' from the typed kwargs, so the SDK's
post-transform extra_body merge appended it after the caller's extra_body keys —
equal dict, different bytes (byte-keyed prompt caches would miss). keep_slots=True
leaves [] placeholders that the merge overwrites in place. Drop the invented
HERMES_CHAT_SDK_TRANSFORM env var; the pre-existing HERMES_CODEX_SDK_TRANSFORM
hatch from #93650 now disables both API families. Tests trimmed to the two
invariants (byte-identity incl. caller extra_body precedence; escape hatch).
2026-09-15 19:23:53 -07:00
John Paul Soliva 1e39c93710 perf(agent): keep bulk chat-completions payloads out of the SDK request transform
`chat.completions.create` re-walks the whole request body against the
`CompletionCreateParams` union graph client-side, with the GIL held, before
any byte leaves the process. #93650 documented that class of walk wedging
for 12+ hours on a ~1.4 MB conversation: no in-process watchdog can fire
while the GIL is held, and no socket kill helps a pre-network hang.
through `extra_body`, which the SDK merges into the JSON body after the
transform — but scoped it to `responses.create`. The default chat path,
which every OpenRouter / Nous / xAI / DeepSeek / Kimi / llama.cpp /
Ollama / LM Studio / LiteLLM request takes, still pays the full walk.

Measured against a real `openai.OpenAI` over an `httpx.MockTransport`
(canned SSE, no network), with the request body captured from the
transport on both sides:

    101 msgs /  76 KB   13.6 ms -> 1.2 ms
    401 msgs / 190 KB   48.6 ms -> 2.0 ms
   1601 msgs / 650 KB  188.7 ms -> 5.8 ms

and the bytes the server receives are IDENTICAL — literally equal, not
merely equivalent (194,894 == 194,894 at 401 messages). The cost is paid
per API call, so a tool-using turn multiplies it by its iteration count.

The three helpers move from agent/codex_runtime.py into a shared
agent/sdk_transform_bypass.py, re-exported under their original names so
agent/auxiliary_client.py and tests/run_agent/test_codex_sdk_transform_bypass.py
keep working untouched. The field tuple is now a parameter:
("input", "tools") for Responses, ("messages", "tools") for chat.

Two chat-specific details. `messages` is a @required_args parameter, so it
stays in the typed kwargs as an empty list and the extra_body copy
replaces it in the body — hence the new `required_empty` argument, which
Responses does not use. And the bypass is gated on the target actually
being the SDK's Completions: Hermes also drives chat-completions-shaped
facades that are NOT the SDK — the in-process MoA aggregator most
importantly — and those never merge extra_body, so handing them one would
silently send an empty message list. That guard is also why this needs no
edits to the 32 test files that assert on kwargs["messages"]: they mock
with stand-ins, not the SDK.

Every rail the merged PR was reviewed on is kept: the plain-JSON-only
guard so pydantic models and generators stay on the typed path, caller
`extra_body` precedence via setdefault (load-bearing here — the chat path
already populates extra_body from custom providers, reasoning config and
Nous Portal), and an env escape hatch, HERMES_CHAT_SDK_TRANSFORM=1,
mirroring HERMES_CODEX_SDK_TRANSFORM.

The summary/compression call sites at chat_completion_helpers.py:3449 and
:3514 carry the largest payloads in the process and are deliberately left
for a follow-up: they route through a lambda whose client is not in scope
at the call site, so they need a slightly different shape and a wider
test surface than this change.
2026-09-15 19:23:53 -07:00
teknium1 10652c9345 fix(desktop): classify any 3xx as a redirect and name its Location
fetchJson/fetchPublicJson only reached the redirect diagnostic when the 3xx
carried an HTML body or text/html content-type; an empty-body 302/307 (the
common reverse-proxy / forward-auth shape) still resolved null through the
earlier empty-body check. http.request never follows redirects, so classify
on status first and reject every 3xx with the redirect error regardless of
body.

Pass res.headers.location through so the message says where the request was
sent, and stop blaming credentials when the Location differs from the
requested URL only by scheme or trailing slash -- that is a saved-URL
mismatch, not an authentication proxy. Drop the 404 mention from the
docstring/test: both callers reject >= 400 before this branch, so only the
2xx leg carries the endpoint-missing capability wording.

Part of #112072
2026-09-15 19:10:53 -07:00
teknium1 c902f50efb fix(desktop): send the connection extra gateway headers on remote media streams
createMediaProtocolHandler() set only the session token / bearer on the
/api/files/stream request. The descriptor it resolves now carries the
connection's extra gateway headers, but the handler never read them, so
attachments in session history from a header-gated remote still bounced off
the access proxy even though every fetchJsonForBackend() call got through.

Merge connection.headers into the media request before auth, letting the
forwarded range/cache negotiation headers win, and pin it on both the
token and the OAuth cookie-session legs.

Part of #112072
2026-09-15 19:10:53 -07:00
teknium1 47c3683f23 fix(desktop): stop calling a 3xx HTML reply a missing endpoint
The JSON guard in fetchJson/fetchPublicJson blamed every HTML reply on a
missing backend endpoint. A 3xx HTML body is an access proxy redirecting
to its login page: the endpoint exists, the credentials never arrived.
Besides misleading the user, the wording is the capability signal that
isMissingHealthEndpointError and the renderer's gateway-rpc predicate key
on, so an auth redirect was silently classified as "endpoint missing" and
routed onto compatibility paths instead of surfacing as an error.

Build the error in api-transport.ts (where the sibling httpStatusError
lives) and pick the hint by status: 3xx names the redirect and points at
the saved token/extra headers; everything else keeps the endpoint-missing
wording the predicates rely on.

Part of #112072
2026-09-15 19:10:53 -07:00
teknium1 e65ddf1c7b fix(desktop): keep primary remote gateway extra headers on REST calls
createPrimaryRemoteConnection() rebuilt the primary remote descriptor
field by field and left out `headers`, so every REST call routed through
fetchJsonForBackend() (Settings profiles/config, session history) reached
the gateway without the configured extra headers while chat, the Test
button and the readiness probe -- which read the exact-URL WebSocket header
store or the resolved route directly -- kept working. Behind an access
proxy (e.g. a service-token gate) those REST calls came back as a 302 to
the login page.

Carry `headers` through the descriptor like every other rebuild site
(buildRemoteConnection, the registry pool path and both ensureRegistryBackend
reuse branches already spread it), and pin it with an invariant test on the
existing primary-descriptor seam.

Fixes #112072
Co-authored-by: plluviera <plluviera@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 19:10:53 -07:00
teknium1 66f9f3692a fix(desktop): share the profile-switch freshness latch across font settings and MCP tab
Extract useProfileSwitchLatch(query stamps) and reuse it in ChatFontSetting,
TerminalFontSetting and McpTab instead of three divergent dataUpdatedAt
copies. The MCP tab keeps its release-on-fresh-error behaviour via the
optional errorUpdatedAt stamp; the font settings keep data-only semantics.
2026-09-15 19:10:25 -07:00
teknium1 92a6639e84 refactor(desktop): fold the font reseed latch into one dataUpdatedAt stamp
The previous commit tracks "waiting for the next config fetch" with a
`profilePending` state, a `staleConfigStamp` ref and a second effect that
clears the flag once `dataUpdatedAt` moves. The same guarantee fits in the
seed effect itself: the profile-switch handler stores the query's current
`dataUpdatedAt` as `staleStamp` and the seed effect refuses to reseed while
`dataUpdatedAt === staleStamp`. One state cell instead of state + ref +
effect, no `no-restricted-syntax` ref write in an effect, and the stale
guard is unchanged: the previous profile's cached record carries the
recorded stamp, so it can never repaint the new profile; only a fetch that
lands after the switch bumps the stamp and seeds.

Co-authored-by: danrudy33 <danrudy33@users.noreply.github.com>
2026-09-15 19:10:25 -07:00
KoNit-K df832c86a2 fix(desktop): reseed font controls after config refetch 2026-09-15 19:10:25 -07:00
teknium1 1e9de95e86 docs(secrets): name OP_CONFIG_DIR among the env vars forwarded to the op child
The 1Password page describes the minimal allowlisted child environment;
list the config-location variables so a user moving the op config dir
(unwritable ~/.config in containers) knows the setting is honoured.
2026-09-15 19:10:00 -07:00
KoNit-K a457e91a50 fix(secrets): preserve OP_CONFIG_DIR for 1Password 2026-09-15 19:10:00 -07:00
teknium1 30b22b54ae fix(tools): bound execute_code's lifecycle probe and keep the terminal guard answerable to /stop
execute_code ran the same unbounded _is_supervised_gateway_process() probe
ahead of every cell, so the wedge #111922 bounds in terminal_tool still hung
an execute_code call (and its cron slot) forever: share the cell's deadline
and fail closed with a retryable error when the probe renders no verdict.

Moving the terminal pre-exec guard onto a deadline worker made it blind to
/stop, which keys on the tool thread's ident: record the acting-for tid in a
contextvar (copied into the worker by run_bounded_sync) so is_interrupted()
on the worker honours the tool thread's bit too.

Floor the guard's share of the deadline at 30s so a short command timeout
does not turn the guard's own cold-start cost (imports, git probes under
load) into a refusal — tests/tools/test_terminal_error_redaction.py was red
on the branch for exactly that.
2026-09-15 19:09:29 -07:00
teknium1 c832920275 fix(tools): pre-exec guard that misses the deadline refuses the command
The salvaged commit put `_pre_exec_block` behind the command's
`run_bounded_sync` deadline but let a timed-out guard fall through into
execution. The gateway-lifecycle, dangerous-workdir and self-repo checks
apply unconditionally (`force=True` cannot bypass them), so a guard that
never rendered a verdict must not let the command run unguarded: return
the terminal error envelope (`status: error`, "did not finish ... Retry
the call") instead, mirroring how the bounded `env.execute` path reports
its own expiry as a result rather than continuing.

Tests trimmed to the two invariants: a wedged guard returns a bounded
error without executing; a completed guard keeps its verdict (pass ->
execution, rejection -> its own blocked result).
2026-09-15 19:09:29 -07:00
KoNit-K c1e749d679 fix(tools): bound terminal pre-exec guards 2026-09-15 19:09:29 -07:00
teknium1 fff10484d6 fix: drop the dead disabled guard on the lazy MCP banner line
get_mcp_status reports status='disabled' (never 'lazy') for a disabled
server and derives the 'disabled' flag from that same status, so the
extra 'and not srv.get("disabled")' check could never change the branch.
2026-09-15 19:06:54 -07:00
teknium1 abdb402701 fix(mcp): carry the lazy status across the TUI wire, tests and docs
Follow-up to the ported status fix:

- `tui_gateway/contracts/tools_mcp_plugins.py::McpRuntimeStatus` is a
  closed wire enum; `mcp.servers.status` would raise `ContractViolation`
  on the new `lazy` value. Declare it and regenerate the TS/OpenRPC
  contract files.
- `ui-tui` session panel: an unknown status fell through to the red
  `failed` branch; render `lazy` with its cached tool count (inline
  branch, no component extraction).
- Two invariant tests, both red on origin/main: the real discovery path
  yields `status: lazy` with the cached tool count and a summary without
  `failed` (eager control stays `configured`, live control stays
  `connected`); a lazy-only run neither warns nor re-arms the startup
  retry, while a configured-only run still does.
- Document the per-server `lazy` key (undocumented until now) in
  `cli-config.yaml.example`, the MCP config reference and the MCP guide.
2026-09-15 19:06:54 -07:00
John Paul Soliva a17d0409be fix(mcp): report lazily registered servers as lazy, not configured or failed
A `lazy: true` MCP server registers its tools from the schema cache and
spawns on first use. Three consumers still equated "alive" with a live
session, so a healthy all-lazy startup was reported as a total failure:

- `get_mcp_status()` fell through to `status: configured, tools: 0` for a
  lazily registered server. It now reports `lazy` with the cached tool
  count (`connected: False`); an in-flight or failed first-use connect
  still outranks it because the error is the actionable part.
- `discover_mcp_tools()`'s summary counted every name absent from
  `_servers` as failed, logging `MCP: 0 tool(s) from 0 server(s) (2
  failed)` right after registering every cached tool, and re-announced
  the same "failure" on every repeat discovery. Lazy servers are now
  reported as `(N lazy, not spawned yet)` and an already-lazy server is
  not re-announced.
- `hermes_cli/mcp_startup.py` judged a discovery run by `connected` at
  two sites, so every startup logged `Background MCP discovery completed
  with zero connected servers` and every later call re-spawned the
  discovery thread as a retry. One predicate,
  `_discovery_registered_servers`, treats a lazy registration as a
  usable outcome at both sites.
- `hermes_cli/banner.py` rendered the unknown `lazy` status through the
  red "could not connect" line; it now shows the cached tool count with
  `(lazy, starts on first use)`.

Ported from #100648 (core hunks only; the toolsets-filter predicate
branch, the Ink TUI component extraction and 13 tests were not ported).

Fixes #111717
2026-09-15 19:06:54 -07:00
teknium1 05fac10a75 docs(api-server): MCP trust-gate consent surfaces as approval.request on /v1/runs
Document that an untrusted-server write-capable MCP tool now parks a run in
waiting_for_approval and is resolved through POST /v1/runs/{id}/approval,
the same bridge dangerous-command approvals already use.

Part of #111526
2026-09-15 19:06:27 -07:00
KoNit-K c001881d85 fix(mcp): route /v1/runs MCP trust-gate consent through the run's approval callback
A write-capable tool on a `trust: untrusted` MCP server was denied instantly
from POST /v1/runs: request_elicitation_consent only took the gateway path
when _is_gateway_approval_context() was true, and api_server sits in
_UNATTENDED_APPROVAL_PLATFORMS (webhook-style sessions have nobody to
answer). A live /v1/runs run is the exception: it registers a gateway notify
callback and answers via approval.request -> POST /v1/runs/{id}/approval —
the same bridge 04fcf9159 keeps alive for the dangerous-command gate. Treat
an api_server session that is neither cron nor single-query as
callback-backed; a run without a registered callback still fails closed.

Salvaged from #111529 with the redundant single-query re-gate on the
generic gateway branch dropped (no real surface binds a chat platform,
HERMES_SINGLE_QUERY_SESSION and an in-process callback together).

Part of #111526
2026-09-15 19:06:27 -07:00
teknium1 341f8b4d93 fix: cap, loopback-bypass and share the MCP proxy mounts
Review follow-up on the MCP HTTP proxy PR:
- Proxy mounts win over transport= for matching URLs, so a bare
  AsyncHTTPTransport mount bypassed the 10 MiB wire-body cap whenever a
  proxy applied. Each mount is now wrapped in _make_mcp_body_cap_transport.
- Loopback MCP servers (127.0.0.1 / ::1 / localhost) were dialed through
  HTTP_PROXY unless NO_PROXY covered them; _mcp_proxy_mounts now returns
  None for is_loopback_host (agent.proxy_bypass rule).
- Dropped the fail-open try/except around the proxy transport construction;
  a proxy httpx cannot build surfaces as the server's connect error.
- The content-type preflight client now takes an explicit transport plus
  the same proxy mounts as the SDK client, so probe and handshake take the
  same route (no httpx env auto-detection divergence).
2026-09-15 19:05:57 -07:00
teknium1 ee1bfef857 fix(mcp): NO_PROXY for MCP servers uses the repo matcher; trim tests to two invariants
Follow-up to the salvaged #111796 commit:

- NO_PROXY matching goes through `agent.proxy_bypass.should_bypass_proxy` (the one
  matcher the LLM transport and the gateway adapters already use), so CIDR ranges and
  `*.host` patterns bypass the proxy for MCP servers exactly as they do for the model
  endpoint. The stdlib `proxy_bypass` stays for the OS bypass list (Windows
  ProxyOverride / macOS exceptions). Live probe: NO_PROXY=10.255.255.0/24 still routed
  the MCP request through the proxy before this commit, direct after.
- Drop the try/except around `getproxies()` / `proxy_bypass()`: the stdlib guards its
  own registry/sysconf reads and httpx calls the same functions unguarded.
- Trim the six contributor tests to two invariants (mount + NO_PROXY incl. CIDR; both
  client builders carry mounts next to the body-cap transport). Fixture uses the stdlib
  `getproxies_environment` / `proxy_bypass_environment` instead of a hand-rolled copy and
  skips when the mcp SDK is absent.
- Docs: one sentence on the MCP page about proxy resolution for HTTP/SSE servers.
- contributors/emails mapping for the PR author.
2026-09-15 19:05:57 -07:00
VictorTran1023 ceb1aa19a0 fix(mcp): restore proxy support for HTTP/SSE MCP servers
httpx auto-detects proxies only when ``transport is None``
(``allow_env_proxies = trust_env and transport is None``). The wire-body cap hands
every MCP HTTP/SSE client a custom transport, so HTTP_PROXY / HTTPS_PROXY and the
Windows-registry / macOS system proxy were silently ignored: on a network that
reaches the MCP host only through a proxy, every connect failed with
"All connection attempts failed" and the server was parked (tools never appeared).

Rebuild httpx's own proxy resolution as explicit ``mounts`` — environment first,
then the OS proxy, NO_PROXY / platform bypass honoured, socks:// normalized, and
TLS settings identical to the transport they accompany.
2026-09-15 19:05:57 -07:00
teknium1 204f345816 refactor(codex): one shared constant for the hermes-tools MCP server name
The name of Hermes' MCP callback for the codex app-server runtime was spelled
as a string literal in five places (the server itself, the runtime migration
that writes `[mcp_servers.hermes-tools]`, the Kanban worker override launcher,
the elicitation auto-accept handler, the display-name stripper and the switch
report) and had already drifted once (#111707). Define it once in
agent/transports/hermes_tools_mcp_server.py — the module that IS the server and
whose module-level imports are stdlib only, so every higher layer (transports,
agent/codex_runtime, hermes_cli) can import it without a cycle — and read it
everywhere.

Two invariant tests in tests/agent/transports/: the worker's `-c
mcp_servers.<name>.env.*` overrides only ever target an entry the migration
really writes to config.toml (red on the pre-fix base: `{'hermes-mcp'}`), and
non-owned launches emit no override at all.

Refs #111707
2026-09-15 19:05:29 -07:00
Shenrui Ma 4521e04add fix(evals): align Kanban probe with managed MCP overrides
Use hermes-tools for the probe's managed server configuration and tool-call target so it exercises the endpoint receiving worker overrides. Leave intentional user-defined hermes-mcp fixtures unchanged.

Scope-risk: narrow
Tested: Three complete isolated probe runs on macOS with Codex CLI 0.137.0; related regression selection 202 passed, 2 existing skips; Ruff and repository static checks
Not-tested: Full repository suite, native Linux/Windows, hosted CI, cloud-model turns
2026-09-15 19:05:29 -07:00
Shenrui Ma 6973f2622d fix(codex): target hermes-tools in Kanban worker overrides
Dispatcher-owned Kanban workers on the codex app-server runtime injected their
HERMES_KANBAN_* scope into `mcp_servers.hermes-mcp.env.*`, but the runtime
migration registers Hermes' MCP callback as `[mcp_servers.hermes-tools]`. The
override therefore materialised a second, env-only server entry that codex
rejects at bootstrap ("invalid transport in `mcp_servers.hermes-mcp`"), so no
worker could initialize. Point the overrides at the entry that actually exists.

Salvaged from #107337 (its 147-line parametrized test file is replaced by two
invariant tests in a follow-up commit). #111711 proposed the identical two lines.

Fixes #111707

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 19:05:29 -07:00
teknium1 837acdc91d fix: share the control-frame opener list and cover [System:/[IMPORTANT:/[PRIOR CONTEXT/[CONTEXT SUMMARY]
Review finding on #112260: the hosted-room member relabel regex was a third
hand-copied opener list that missed the frames context_compressor and
title_generator already treat as harness input, so a member reply starting
with "[System: ..." or "[IMPORTANT: 2 background processes ..." reached peers
in its exact trusted shape. agent.prompt_builder.CONTROL_FRAME_OPENERS is now
the single source; the gateway regex is built from it and the desktop TS
literal mirrors it byte-for-byte.
2026-09-15 19:04:59 -07:00
teknium1 9988545a78 fix(gateway,desktop): one visible relabel for member-quoted control frames on both room surfaces
Follow-up to the two cherry-picked contributor commits (#111571 gateway, #111576 Desktop),
which neutralized the same class with two different mechanisms: an invisible U+200B after
the `[` on the gateway path and a phrase replacement ("RESERVED CONTROL MARKER NEUTRALIZED")
on the Desktop path.

Both room-transcript builders now share one frame set (the mid-turn steer marker open/close,
the compaction handoff, runtime/system notes, planning-state and async-delegation frames —
the same openers agent/title_generator and agent/context_compressor already classify as
harness-authored, case-insensitive) and one visible relabel: the opener `[` becomes
`[member-quoted `. The peer still reads what the member wrote, but the exact trusted shape
the system prompt tells the model to honour is gone and the label says who authored it.
A zero-width space is invisible to a human reading the transcript and easy for a model to
skip over; the visible label is not.

Genuine user lines, stored room events and the displayed message are untouched on both
surfaces (probed live: stored member event byte-identical, `User (user):` line verbatim).
Tests trimmed to one invariant per surface, built from agent.prompt_builder's real marker
constants on the Python side.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
2026-09-15 19:04:59 -07:00
Hukla 498633f4ec fix(desktop): neutralize member control markers 2026-09-15 19:04:59 -07:00
KoNit-K 925d3448ce fix(gateway): neutralize member control frames 2026-09-15 19:04:59 -07:00
teknium1 8e11666726 fix(mcp): resolve managed Windows Node launchers (npx.cmd/npm.cmd) for stdio MCP servers
On Windows a stdio MCP server configured with `command: npx|npm|node` failed
with WinError 2 whenever the desktop/gateway PATH lacked the managed Node dir:
`_node_fallback` probed only the POSIX shape `<HERMES_HOME>/node/bin/<cmd>`
with no extension, while `scripts/install.ps1` unpacks Node directly into
`<HERMES_HOME>\node` as `npx.cmd`/`npm.cmd`/`node.exe`. It also derived the
home from raw `os.getenv("HERMES_HOME")`, so a context-local profile home
(multiplexed gateway) was ignored.

Reuse the platform-aware helpers instead of a second hand-rolled layout:
`hermes_constants.iter_hermes_node_dirs(get_hermes_home())` supplies both
managed shapes in the right order, and the module's own `_npx_bin_candidates`
supplies the `.cmd` -> `.exe` precedence (same injectable `windows=` seam the
npx-cache shortcut already uses, so the branch is testable on Linux CI).
POSIX candidates (`node/bin`, `~/.local/bin`, `/usr/local/bin`) are unchanged.

Slimmer redo of #111941 by @KoNit-K, which re-derived the Windows shape
in-place and kept the raw env read.

Fixes #111937

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 19:04:29 -07:00
teknium1 a2837ec088 docs(docker): overriding entrypoint: drops the zombie reaper — document init: true; trim the PID-1 warning
Follow-up to the cherry-picked #111584 (@chelsealong):

- website/docs/user-guide/docker.md: new warning block next to the existing
  "do not override the entrypoint" note explaining WHY (with `/init` gone the
  hermes process is PID 1 and nothing reaps orphaned browser/MCP/shell
  children), the Compose `init: true` / `docker run --init` remedy, and that
  supervision is still lost on that path; plus a Troubleshooting entry for
  `<defunct>` processes under PID 1.
- hermes_cli/main.py: `_warn_if_unsupervised_pid1` keeps the `os.getpid() == 1`
  check and drops the `platform.system()` gate and the blanket
  `try/except Exception: pass` — a user process is never PID 1 on any host OS
  (PID 1 is init/launchd; Windows PIDs are multiples of 4), and nothing in the
  check can raise.
- tests trimmed to two invariants (warns at pid 1 / silent otherwise).

Not done, on purpose: a `prctl(PR_SET_CHILD_SUBREAPER)` + SIGCHLD reaper in
main-wrapper/hermes. As PID 1 hermes already receives the orphans; what is
missing is a `waitpid(-1)` loop, and a process-wide one races
`subprocess.Popen` for exit statuses. The maintainer decides whether that
runtime change is wanted; docs + the startup warning cover the reported
deployment.
2026-09-15 19:04:00 -07:00
chelsealong 8d5cce4d94 fix(cli): warn when hermes runs unsupervised as PID 1
A deployment that overrides the image's `entrypoint:` to invoke hermes
directly skips docker/entrypoint-dispatch.sh entirely, so hermes itself
becomes PID 1 with no s6-overlay /init (or any other init) above it.
Nothing then reaps orphaned grandchildren (browser tooling, MCP
subprocesses, shell-tool children) reparented to PID 1, and they
accumulate as zombies without bound.

entrypoint-dispatch.sh already warns on its own non-PID-1 fallback
path, but that script never runs in the entrypoint-override case, so
there was no signal at all. Add the same style of warning inside
hermes_cli.main, gated on being PID 1 on Linux, pointing users at the
image's default ENTRYPOINT or `docker run --init` / `init: true`.

Fixes #111577
2026-09-15 19:04:00 -07:00
teknium1 54ed7cbb7b fix: validate the skill name before opening its lock; key lock files on a digest
Review finding on #112218 (major): `_skill_lock_path` opened `<skills>/.locks/<name>.lock`
before the name was validated, so `skill_manage(action='create', name='a'*300)` raised
OSError (File name too long) and a NUL name raised ValueError instead of the handler's
JSON error, and every rejected name ('../../etc', '') left a residue lock file.

- tools/skill_manager_tool.py: lock filename is sha256(basename).lock (fixed width, no
  filesystem limit reachable; `foo` and `category/foo` still share one lock), the redundant
  `_find_skill` rglob is gone, and `skill_manage` runs `_validate_name` on the name
  (create) / basename (other actions) before the lock is opened.
- '.locks' joins the skills-dir exclusion sets (EXCLUDED_SKILL_DIRS, ledger
  _NON_PACKAGE_TOPS, learning-graph/skill-commands skip parts, curator backup excludes).
- tests: 2 invariants in TestSkillMutationLock (rejected names -> JSON + no .locks residue;
  digest-keyed lock shared across name forms), red on the old head.
2026-09-15 19:03:33 -07:00
teknium1 273986f88f fix(skills): route the skill_manage lock through the existing skill_usage lock helper
Slim follow-up to the cherry-picked #111585 (@KoNit-K):

- tools/skill_usage.py: generalize the usage ledger's `_usage_file_lock()` into
  `skill_file_lock(lock_path)` — same fcntl/msvcrt idiom, now thread-re-entrant
  via a per-thread held set (flock is not re-entrant across separate fds; a
  ContextVar would leak "held" into copy_context() timer threads).
- tools/skill_manager_tool.py: drop the third fcntl/msvcrt copy, hashlib and the
  ContextVar; the per-skill lock is `<skills>/.locks/<skill-dir-name>.lock`
  (readable, outside the skill dir so delete/recreate cannot unlink it under a
  waiting writer). Batch locks sort by lock PATH, not name, so two batches
  naming the same skills in different forms cannot deadlock.
- tools/skill_manager_batch.py: plain `with` around snapshot -> commit/rollback
  instead of manual __enter__/__exit__ bookkeeping.
- tests: trimmed to two invariants — the two-writer lost-update test on
  SKILL.md (from #111585) and a re-entrancy/exclusivity test on the helper.
  Dropped: the edit/write_file/remove_file parametrization (same dispatcher
  path as patch) and the category-dir cleanup test (lock files never lived in
  category dirs here).
2026-09-15 19:03:33 -07:00
KoNit-K 0b8b000ce9 fix(skills): serialize skill mutations 2026-09-15 19:03:33 -07:00
teknium1 173ccfe2dd fix: drop dead shutil.which patches in bot-chat delivery tests
Review finding (minor): after the module-first reorder in cron/scheduler_delivery.py the delivery.shutil.which -> /bin/hermes monkeypatches were unreachable; both tests already accept the module argv.
2026-09-15 19:03:08 -07:00
teknium1 da18c20226 fix: kanban dispatcher prefers module argv over PATH hermes
Review finding: hermes_cli/kanban_db_dispatch.py::_resolve_hermes_argv still resolved which('hermes') before sys.executable -m hermes_cli.main while claiming to mirror gateway.run._resolve_hermes_bin, which this PR made module-first (#111569). Keep the explicit HERMES_BIN override first, then the module argv whenever hermes_cli is importable, PATH only as fallback; docstring updated.
2026-09-15 19:03:08 -07:00
teknium1 b0bde32959 test(cron): bot-chat CLI-home test strips the launcher prefix without requiring -p
The fixture adaptation for the `python -m hermes_cli.main` launcher located the
CLI argv via `argv.index("-p")`, but profile homes (`profiles/<name>`) never get
a `-p` flag appended, so the `beta` parametrization raised ValueError inside the
fake subprocess.run and the delivery reported failure. Strip the launcher prefix
by shape instead (3 tokens for `python -m hermes_cli.main`, 1 for a binary).
2026-09-15 19:03:08 -07:00
teknium1 f336048b08 fix(cron): bot-chat delivery launches the running install, not whatever hermes PATH names
The cron scheduler runs inside the long-lived gateway process and spawned
`hermes ... chat` for Bot Chat delivery through `shutil.which("hermes")`
first, falling back to `sys.executable -m hermes_cli.main` only when PATH
had no `hermes`. That is the same resolution order gateway.run.
_resolve_hermes_bin just flipped for /update and /restart (#111569): the
running interpreter's module argv is exactly this install, PATH is not.

Delivery now resolves the running install first and uses PATH only as the
fallback. Tests that asserted the PATH argv shape or armed on the `which`
seam are moved to the module-argv shape / the `find_spec` seam.
2026-09-15 19:03:08 -07:00
liuhao1024 fea812824d fix(gateway): resolve the update/restart argv from the running install, not PATH
_resolve_hermes_bin preferred `shutil.which("hermes")` over the running
interpreter's module argv. On Windows a hermes.exe planted earlier in PATH is
therefore the argv /update and /restart re-exec, hijacking the update process
(#111569). Flip the order: python -m hermes_cli.main (exactly this install)
wins whenever hermes_cli is importable; PATH stays as the fallback, then None.
2026-09-15 19:03:08 -07:00
teknium1 4b1215be2f fix: map MrMongstad's GitHub noreply email for contributor credit
Review finding: the Co-authored-by trailer on 2751fb7 uses stephan@users.noreply.github.com, which resolves to a different GitHub account; the pushed commit cannot be rewritten, so record the correct mapping via scripts/add_contributor.py and co-credit in the PR body.
2026-09-15 19:02:39 -07:00
teknium1 55e2986dfd fix: walk a group's __cause__/__context__ in the exception node walker
Review finding: _exc_children returned only .exceptions for a group, so
_is_session_expired_error missed a session-expiry marker (or the
InterruptedError override) hanging off a group's __cause__/__context__
that main used to inspect. Groups now yield nested + chain like every
other node; _flatten_messages' "group str() is opaque" rule is unchanged.
2026-09-15 19:02:39 -07:00
teknium1 e1114bdcf9 refactor(mcp): one cycle-safe exception walker for every connect-error scan
The salvaged fix gave `_find_missing` and `_flatten_messages` each their own
visited-set loop, next to the one `_is_session_expired_error` already had —
three copies of the same idiom in one module. Collapse them into
`_iter_exception_nodes` (pre-order, left-to-right, each node once, bounded by
`_EXC_TRAVERSAL_MAX_NODES`) and read all three scans off that list. Acyclic
output is byte-identical: the missing-executable search keeps its depth-first
order and a message-less leaf still renders as its class name.

Tests move from the issue-numbered file into `tests/tools/test_mcp_tool_errors.py`
(mirror of the source module): a two-node cycle renders the real messages, and a
missing stdio binary wrapped deeper than the recursion limit with the chain
looping back to the top is still reported as the missing executable. Both are
red on origin/main (RecursionError).

Co-authored-by: Stephan Mongstad <stephan@users.noreply.github.com>
2026-09-15 19:02:39 -07:00
KoNit-K 030d4caa0e fix(mcp): bound nested connection error traversal 2026-09-15 19:02:39 -07:00
teknium1 2588c908e7 fix: recognise fences opened inside list items and blockquotes
Review finding on #112198: _mask_prose_link_destinations matched
_FENCE_LINE against the raw line, so a fence behind a CommonMark
container prefix (`- ```sh`, `1. ```sh`, `> ```sh`, nested) was not
seen and its body was scored as prose with link destinations masked.
Strip the container prefix before fence matching (open and close).
Bundled-skill rescan vs origin/main: 208 skills, 1447 findings on
both, no new/gone findings, no verdict changes.
2026-09-15 19:01:39 -07:00
teknium1 33292affd6 fix(skills): fenced blocks close only on a matching fence; temp-root rm covers //.. and ..;
Follow-up to the two cherry-picked contributor commits.

The picked fence tracker never checked for a closing fence once a block was
open (the closer test sat inside the not-in-code branch), so every prose link
after any code block was scanned verbatim again and the #111254 documentation
link exemption was lost; a fence line carrying an info string was also accepted
as a closer, which handed the scanner back to prose mode mid-block. Rewrite the
loop around CommonMark fence semantics: a block opens on 3+ backticks/tildes
indented at most 3 spaces (backtick info strings may not contain a backtick)
and closes only on a fence with the same marker, at least as long, and nothing
after it; tab- or 4-space-indented lines are code; an unclosed fence stays
code to EOF. plugin_guard inherits the behaviour through scan_file.

The temp-root exemption in destructive_root_rm now also refuses a parent
segment reached through an empty path segment or followed by a shell
separator, which the first cut let through.

Tests trimmed to one invariant per fix: the fence test covers the six code
shapes plus the prose-link-after-fence control that the picked version broke;
the rm test gains the two residual shapes.

Part of #111334
Fixes #112129
Fixes #111335
2026-09-15 19:01:39 -07:00