Commit Graph

3086 Commits

Author SHA1 Message Date
teknium1 d28938d3da fix(computer-use): screenshot dedup forgets its last frame at a compaction boundary
The "screen unchanged" result points the model at its previous capture. After
context compression that capture may be summarized away, so the note would refer
to pixels no longer in context. Mirror read_file's reset_file_dedup: the
compaction boundary (both the summary path and the codex app-server path) now
clears the session's screenshot digest, and the first capture afterwards delivers
the image again even when the screen is byte-identical.
2026-09-14 07:21:18 -07:00
teknium1 31964ff4c6 fix(computer-use): dedup keyed by the scoped session, cleared on release
Key the screenshot-dedup state by the same profile-scoped session id the
backend cache uses, so two multiplexed profiles sharing a session id (or a
DISPLAY) never dedup against each other's frames, and forget the state in
release_computer_use_session so a re-created session's first capture
always delivers pixels. Reword the unchanged note to cover the aux-vision
path (where the prior result was an analysis, not an image). Tests: the
dispatch path (explicit capture + capture_after) honours the streak cap;
release forgets state.
2026-09-14 07:21:18 -07:00
Teknium 682b973b32 feat(computer-use): stop resending unchanged screenshots
Port from openclaw/openclaw#129924: a capture whose pixels are
byte-identical to the previous capture of the same target in the same
session returns its full text metadata (element index included) plus an
explicit 'screen unchanged' note instead of the multimodal image block.

Adapted for Hermes: openclaw gates dedup on per-frame context-epoch
tracking; Hermes bounds staleness with a consecutive-omission streak cap
(2) so full pixels are re-delivered before compaction could evict the
referenced image. Dedup state is per-session (no cross-session leaks),
append-only (no history rewrites — prompt cache prefixes untouched),
and skipped entirely when no session_id is present.
2026-09-14 07:21:18 -07:00
teknium1 49c6d4a9e0 test(contracts): tests mirror tui_gateway/; the runtime-artifact spoof test asserts the new 4000
tests/contracts -> tests/tui_gateway/contracts (tree-layout rule: tests mirror a source
package). test_rpc_params_cannot_spoof_runtime_artifacts: forged owner_transport /
owner_session_record / owner_token keys are now refused at the wire (4000 + key path)
instead of silently dropped before the handler; the invariant (no steer reaches the
agent) is unchanged and asserted directly.
2026-09-14 06:12:19 -07:00
teknium1 01338e88ec test(video-gen): keep the i2v-only guard real once every family is dual-modality
The rebased synthetic guard built a bare dict that main's _build_payload
no longer accepts (it indexes aspect_ratios/resolutions directly) and never
exercised generate(); it now builds the family via _family() and asserts
generate() returns modality_unsupported without submitting. The surface
matrix's i2v-only parametrization collected zero cases after Gemini Omni
Flash 1.1 gained t2v, so it is removed along with the dead branch in the
t2v matrix.
2026-09-13 21:40:26 -07:00
teknium1 60559d4e0e fix(delegation): late-attached children take the parent's soft/hard stop kind; dedupe the fallback replay
_attach_child now mirrors a pending parent stop with the same split
interrupt() uses for its own fan-out (hard -> hard_interrupt, soft ->
interrupt), so a redirect is not turned into a cancel on a child that was
attached late. _restore_parent_cancellation collapses to re-attaching the
rejected unit's children: the replay is the attach step's job now.

Test fixture: _Batch gained origin_session_history_delivery on main after
the salvaged PR was written.

Co-authored-by: illidan <noequal666@gmail.com>
2026-09-13 21:31:57 -07:00
illidan 86258b9912 test(delegation): enable independent units in partial admission coverage 2026-09-13 21:31:57 -07:00
illidan 946111cb11 fix(delegation): restore parent cancellation for rejected async units 2026-09-13 21:31:57 -07:00
teknium1 8a5a66d0d8 test: trim the Aug-2026 conformance suites to behaviour invariants on main
Rebased onto main (~9700 commits): the branch predates the god-file
decomposition and the tests/state -> tests/hermes_state move.

Fixes for seams that moved:
- cron memory contract patches cron.scheduler_delivery._resolve_origin and
  hermes_state_registry.acquire (where run_job now reads them).
- state.db conformance imports _live_writer_holds_db from
  hermes_state_repair and drives the public create_quick_snapshot.
- update receipt: serve runtimes are reconciled in their own unit
  vocabulary since #100479, so the "full accounting" row names the serve
  unit instead of borrowing a gateway relaunch.

Deleted (change-detectors / source-greps / duplicates / dead code):
- source-text scan of hermes_state*.py for "maintenance-shaped" defs and
  the symbol REGISTRY it fed (renames are not regressions).
- per-op exact-outcome table under a live writer -> one invariant: refuse
  with a lock error or report zero work, DB stays intact.
- copy_db_and_verify pins: the symbol is a revert-scheduled plugin-compat
  pointer, not production code (check_compat_pointers.py).
- cron ON-direction tests already pinned by tests/cron/test_scheduler.py,
  plus a tautology that never called production code.
- exact warning-wording asserts in the env deprecation truth table.
- same-file duplicates: fixed-point (implied by idempotence), tripwire
  round-trips, absent-registry, sequential "race" re-enactment.
2026-09-13 21:13:23 -07:00
Teknium d15f4a7826 test: regression conformance suites for the 7 recurring Aug-2026 bug classes
Two weeks of closed issues/merged PRs show the same areas regenerating:
each salvage pinned its instance while the class invariant had no test.
These suites pin the invariants themselves:

- tests/conformance/test_profile_write_tripwire.py — no writes to the
  default profile tree while a profile is active (#88532 #92662 #89190
  #89625 #92156); reusable tripwire fixture, 4 surfaces
- tests/hermes_cli/test_env_deprecation_truthtable.py — 18-row truth
  table for the Deprecated-.env warning (#88829 #89016 #89389 #90299)
- tests/cron/test_cron_memory_contract.py — cron<->memory contract that
  flipped twice in Aug (#91269 -> #91384 -> #91447)
- tests/agent/test_injected_param_strip_retry_registry.py — every
  strippable injected param x real 400 shapes must strip-and-retry;
  unknown params must still fail (#90257 #89897 #91164 #89503)
- tests/agent/test_transcript_decoration_idempotence.py — f(f(x))==f(x)
  law + 4-breakpoint budget for apply_anthropic_cache_control (#90971)
- tests/state/test_state_db_maintenance_conformance.py — registry-
  enumerated maintenance ops refuse/degrade under a live writer; copies
  of corrupt DBs are refused or flagged (#91839 #90806 #90613 #88235)
- tests/tools/test_bot_mode_canonical_chat_resolution.py — canonical
  Bot Chat resolution is idempotent, never mints, unique per profile,
  race-safe (#92040 #90705 #92692 #90005 #90732, PR #92129)
- tests/hermes_cli/test_update_receipt_truthfulness.py — receipts:
  crash never claims success; success requires full fleet accounting;
  refusal != failure (#91283 #91439 #92902 #92780)

117 tests, all sabotage-verified (each suite proven to FAIL when its
bug class is reintroduced).
2026-09-13 21:13:23 -07:00
teknium1 230ca004a7 fix(delegate): forward inline data-URL images to vision children; trim tests to invariants
data:image/... entries were treated as local paths and silently skipped as
"unreadable". They now ride as image_url parts only (never pasted into the
text hint, never appended to a text-mode goal). Skips and forwarding
failures log at warning since the caller explicitly asked for the images;
decide_image_input_mode gets the child's requested_provider like the CLI
and gateway callers.

Tests collapse to five invariants, including one that drives _ChildRun
and asserts the multimodal content list reaches run_conversation as the
first user turn. Docs mention data: URLs and the read guard.
2026-09-13 21:05:42 -07:00
Teknium f3f5c4f7c7 Port from RooCodeInc/Roomote#1796: per-task image forwarding on delegate_task
Subagents can now SEE images. Each delegate_task task accepts an optional
images list (max 8; local paths or http(s) URLs). Vision-capable children
receive native image_url content parts on their goal turn (local files as
data URLs, remote URLs verbatim); non-vision children get
[Image attached at: ...] hints plus a vision_analyze pointer. Routing
reuses agent.image_routing (decide_image_input_mode /
build_native_content_parts), so agent.image_input_mode governs delegation
exactly like inbound gateway images.

Best-effort by contract: malformed images arrays fail the call loudly
before any child spawns; unreadable paths are skipped with a log line;
any exception in the forwarding path degrades to the text-only goal.

Adapted from RooCodeInc/Roomote#1796 / #1767 (Fast agent forwards bounded
current-turn attachments to delegated coding tasks).
2026-09-13 21:05:42 -07:00
teknium1 dcf17633de test(process_registry): trim handle-release tests to invariants, fix prune fixture
Drop the no-op and still-running change-detectors (4 invariant tests remain:
pipe closed on finish, PTY closed on finish, poll still serves buffered
output, prune releases handles). The accretion-caps fake session now
carries process/_pty like the real dataclass, since prune reads them.
2026-09-13 21:04:45 -07:00
teknium b3a8734d43 fix(process_registry): release handles on prune paths too (salvage follow-up)
Widen #75162: _prune_if_needed() drops finished sessions (TTL expiry and
oldest-finished eviction at MAX_PROCESSES) — release their Popen/PTY
handles there too, covering sessions inserted into _finished without
passing through _move_to_finished(). The release helper is idempotent,
so double-close on the normal path is a no-op. Adds two tests: prune
releases handles of dropped sessions, and a still-running session's
pipe stays open.
2026-09-13 21:04:45 -07:00
RGerrish 333733163e fix(process_registry): release Popen/PTY handles when a session finishes
Finished sessions retained their subprocess.Popen pipe objects (and PTY
masters) until the finished-process TTL (FINISHED_TTL_SECONDS, default 30
minutes) elapsed. Under heavy background churn — deployments, archivers,
watchers — finished-but-unpruned sessions accumulated one open pipe FD
each, exhausting the gateway process's file descriptor budget and
surfacing as a 'file descriptor limit' error on new background spawns.

The registry never rejects spawns (it prunes oldest-finished at
MAX_PROCESSES), so the real defect was the retained-handle leak, not a
registry-cap rejection. The fix closes each finished session's Popen
stdout/stderr/stdin streams and PTY master in _move_to_finished(), right
after the reader loop drains EOF. poll()/wait()/read_log() serve output
from the buffered output_buffer — never from the pipe — so the release is
lossless.

Tests: 4 new cases in TestFinishedHandleRelease — Popen pipes closed,
PTY closed, no-handle sessions safe, and poll() still serves buffered
output after the release. All 4 fail on main (reproduction) and pass
with the fix.
2026-09-13 21:04:45 -07:00
Teknium c63de5a231 feat(tool_search): long hunts for nonexistent tools now return no results instead of incidental matches
Port from nearai/ironclaw#7965: BM25 admits any document scoring above
zero, i.e. sharing ONE term with the query. A long descriptive search
for a capability that does not exist therefore returned a plausible-
looking ranked list, and the model read 'results exist' as 'it is in
here somewhere' and rephrased instead of stopping (IronClaw production
trace: 652 tool calls, 216 of them tool_search, hunting a 'data' tool
that did not exist).

A document must now match at least half the query's ANSWERABLE terms
(terms present anywhere in the index) before it is offered. Coverage
only engages from four answerable terms up, preserving recall on short
queries; exact tool-name matches remain authoritative; the substring
fallback is unchanged.

Docs: relevance-floor bullet added to tool-search.md implementation
details.
2026-09-13 21:04:08 -07:00
Teknium ae596d805b feat(image-gen): Kling Image v3 in the FAL catalog (t2i + singular-key i2i)
Adds fal-ai/kling-image/v3/text-to-image ($0.028/img, native 2K default,
8 aspect ratios) with its image-to-image edit endpoint. The i2i schema
takes a SINGULAR `image_url` string instead of the usual `image_urls`
list, so the catalog gains an `edit_image_param` knob that
_build_fal_edit_payload honors (first source image only); the
edit-contract test now validates whichever image key the entry declares.
2026-09-13 21:03:11 -07:00
Teknium f289aae1f5 Port from aaif-goose/goose#11466: recognize Windows package-runner shims in OSV malware preflight
uvx.exe (uv's actual Windows shim) and pipx.exe bypassed the MCP OSV
malware check entirely, and backslash-qualified commands only resolved
when running under ntpath. Basename now splits on both separators;
matching stays exact (npx.cmd / uvx.exe / uvx.cmd / pipx.exe) so
lookalikes like npx.exe or npx.cmd.bak remain fail-open.
2026-09-13 20:46:22 -07:00
teknium1 9f1ddd927c test: drop upstream product reference from body-cap test docstring 2026-09-13 20:45:25 -07:00
Teknium a6fdadfcee fix(mcp): cap HTTP/SSE response bodies before SDK parse
Port from openclaw/openclaw#123194: a hostile or misbehaving remote MCP
server could stream an unbounded HTTP catalog/tool-result body that the
MCP SDK buffers and JSON-parses before any of Hermes' post-parse limits
(resource cap, tool-result truncation) run.

New _make_mcp_body_cap_transport wraps the owned httpx AsyncClient's
transport on the Streamable HTTP (mcp >= 1.24) and SSE paths:
- finite HTTP bodies capped at 10 MiB (Content-Length rejected up front,
  streamed bodies capped chunk-by-chunk);
- each SSE event capped at 10 MiB, with accounting reset at completed
  event boundaries so long-lived streams/keepalives are unlimited;
- violations raise httpx.ReadError naming the byte cap, handled by the
  existing transport teardown/reconnect path (#66092).

verify/cert now live on the inner AsyncHTTPTransport (client-level TLS
kwargs are inert once a custom transport is passed); the SSE
httpx_client_factory is always injected so the cap applies with default
TLS too. Legacy mcp < 1.24 path (SDK-internal client, no hook) stays
uncapped — same degradation as strict_redirect_headers.
2026-09-13 20:45:25 -07:00
chelsealong ef0136385c fix(mcp): clamp generated MCP tool names to 64 chars
Portable Agent Plugin packages fold the plugin name into the MCP
registry name three times over (slug, digest, and again as the server
key), so mcp__<server>__<tool> routinely exceeds the 64-char function
name limit OpenAI-compatible providers enforce — while the same server
registered via `hermes mcp add` stays well under it. The oversized name
is never rejected loudly; the tool just becomes unreachable. Clamp
mcp_prefixed_tool_name() to 64 chars with a deterministic, collision-
safe hash suffix, mirroring the existing property-key clamp in
schema_sanitizer.py. Dispatch is unaffected since handlers already
close over the original unprefixed tool name.

Fixes #81331
2026-09-13 20:44:29 -07:00
Teknium d9e88e19e2 feat(mcp): bind stored OAuth refresh tokens to their issuer
Port from openai/codex#39615: the authorization server discovered for an
MCP server can change (protected-resource metadata edit, server
migration, DNS takeover). Without binding, Hermes would send the stored
refresh token to whatever issuer the server now advertises — handing a
long-lived credential to a different authorization server.

- HermesTokenStorage records hermes_issuer alongside cached tokens
  (stripped before OAuthToken.model_validate; never sent on the wire).
- Both provider classes (tools/mcp_oauth.py legacy path and
  tools/mcp_oauth_manager.py managed path) stamp the discovered issuer
  on every token save and enforce the binding on _initialize.
- On mismatch: refresh token is stripped from memory and disk; the
  unexpired access token keeps working; full re-auth happens at expiry.
- Legacy token files without an issuer adopt the current one once
  (no forced re-login for existing installs — deliberate divergence
  from Codex, which requires reauth).

Validated: 12 new tests + 163 existing MCP OAuth tests green; sabotage
run confirms the new tests fail without the enforcement; E2E against
the real manager provider class with a temp HERMES_HOME confirms
mismatch strips and match preserves.
2026-09-13 20:43:33 -07:00
memosr d3fc0cca0f fix(security): escape OAuth error parameter in callback HTML to prevent reflected XSS 2026-09-13 20:42:37 -07:00
teknium1 dd497c3d59 fix(delegate): grandchildren spawned after their orchestrator was stopped now die with it
AIAgent.interrupt() fans the stop out to a snapshot of _active_children. A child
that is attached after that snapshot — an orchestrator subagent still building
its fan-out siblings, or one that has not yet hit its next iteration check —
started with no signal and ran to completion as an orphan while its parent had
already reported `interrupted`. Live repro (mid orchestrator, stop delivered
between grandchild A and B builds): grandchild B kept its `sleep 20` alive and
stayed in the subagent registry after the mid returned.

_attach_child now mirrors a pending parent stop onto the newcomer, so the whole
spawn tree dies with the node that was stopped. Tests: unit (late attach gets
the stop, normal attach does not) + the real _build_child_agent path with a
stopped orchestrator.
2026-09-13 20:13:12 -07:00
teknium1 98a3324821 fix(approval): approvals.mode off bypasses the shared action gate (computer_use prompts)
The Desktop "Approvals: off" toggle persists approvals.mode: off. The shell
guards (check_all_command_guards / execute_code) honour it as a bypass, but
_run_approval_gate, the shared gate that computer_use, plugin approval rules,
SSH-config writes and the dangerous-pattern prompt all route through, only
checked _yolo_active() (process --yolo / session /yolo). So with approvals
off, every destructive computer_use action still prompted.

Regressed when 3e066dfedd moved computer_use onto the shared gate: its old
private gate never consulted mode at all, and the shared gate had never been
given the third bypass source. Gate now mirrors the shell guards:
yolo OR approvals.mode == "off" -> approved. Hardline blocks and deny rules
still run before it.
2026-09-13 19:58:33 -07:00
teknium1 9d39267def fix(mcp): supports_parallel_tool_calls is per profile, not per server name
_parallel_safe_servers was keyed by the raw server name while _servers
moved to (scope, name) connection keys (ceaf622c6d). With profile A's `x`
serial and profile B's same-named `x` opted into parallel calls, B's
discovery pass flipped A's tool to parallel-safe and the batch planner put
two A calls in one parallel segment against a server that never opted in
(review of #108352, finding B).

The opt-in is now recorded under the discovering profile's own key and
is_mcp_tool_parallel_safe() looks it up under the calling profile's key,
so one profile's policy never reaches another's same-named server.
Single-profile processes keep the bare-name key, byte for byte.
2026-09-13 15:41:01 -07:00
teknium1 e609efb06a fix(mcp): an adopting profile keeps its own trust policy for a shared MCP connection
Under gateway.multiplex_profiles a profile whose mcp_servers entry has the
same route and credentials as another profile's live connection adopts that
connection instead of opening its own. _same_server_route() compares only
the connection identity, so a `trust: untrusted` profile adopted a
`trust: full` profile's connection; _trust_gate_check() then resolved the
OWNER's connection key, read `full`, and let the untrusted profile run
write-capable tools without the approval prompt its config demands
(review of #108352, finding A; regression from ceaf622c6d, where the
name-keyed ledger let the last registrant's trust win instead).

`trust` is the consuming profile's policy, not a property of the
connection: _server_trust_levels is now keyed by the calling profile's own
key (recorded at its own registration and at adoption, dropped when its
overlay is removed), while readOnlyHint stays under the connection key
because it describes the server's tools. Sharing the connection is still
allowed — only the gate is per profile.
2026-09-13 15:41:01 -07:00
teknium1 1e76efbe28 fix: scan plugin test trees again, cap their criticals at caution
Skipping `tests/`, `spec/`, ... in EXCLUDED_DIRS made those trees
invisible to the guard, but `plugins_loader._load_directory_module`
sets `submodule_search_locations=[plugin_dir]`, so a plugin
`__init__.py` doing `from .tests import evil` imports and runs whatever
lives there: a `tests/evil.py` with a destructive root remove scanned
`dangerous` on main and `safe` on this branch. `_walk` also matched the
names at any depth, so `src/spec/handler.py` — plain runtime code — went
unscanned.

Keep scanning everything; instead cap a critical finding located under a
ROOT-level test dir at `high`, so the verdict is `caution` (confirmation
required, `--force` overridable) rather than the un-overridable
`dangerous`. Fixture strings still cannot brick an install, which was
the reported problem, while a critical in any runtime file (`setup.sh`,
`src/spec/...`) still yields `dangerous`. Trade-off stated in the PR
body: hostile code deliberately placed under `tests/` is now
force-installable rather than blocked outright.

Docs no longer claim test code never runs.
2026-09-13 14:43:04 -07:00
joaomarcos fe97c84c73 fix(plugin): identify critical findings in install blocks 2026-09-13 14:43:04 -07:00
liuhao1024 07b60880cf fix(plugins): stop the security scanner from reading test trees
plugin_guard walks the whole plugin clone, and EXCLUDED_DIRS skipped
caches and vendored dirs but not tests/. A security-conscious plugin's
test suite SHOULD contain adversarial fixtures — a test asserting the
trust boundary holds round-trips the injection string verbatim — and
any single critical finding makes the verdict dangerous, which --force
explicitly cannot override. Scanning tests therefore made exactly the
plugins that test their security unconditionally uninstallable, and the
only workaround was obfuscating the payload strings, weakening the
tests and inverting the incentive. Fixtures are never loaded into an
agent's context at runtime the way README/plugin.yaml are.

Add the conventional test/spec/fixture directory names to
EXCLUDED_DIRS, alongside the existing cache/vendored skips.
2026-09-13 14:43:04 -07:00
teknium1 71cebc6348 fix(tools): browser_exec and computer_use caches are namespaced by the served profile
Both process-global caches were keyed by the caller's session/task id alone, so under
gateway.multiplex_profiles two profiles using the same id — a shared `browser_exec session=`
name, or two Hermes sessions whose screens report the same DISPLAY — resolved to the FIRST
profile's cloud browser / cua-driver, and a command issued in one bot's chat could act on
another bot's screen.

The key now carries the routed profile's home key whenever a served-profile scope is active
(`get_hermes_home_override()` set), the same shape `tools/approval.py::_baseline_key` and the
camofox/cloud caches already use; outside a scope every key is byte-identical to before. The
computer_use lookup, install and release paths all go through one `_scoped_sid`, so a release
under profile B never stops profile A's driver; approval-bypass state keeps the bare session id.

Fixes #110032 (report by @wolfyy970, from @vandaimer's manual test on #108914).
2026-09-13 14:41:26 -07:00
webtecnica 9e6a8645ea fix(browser): resolve the Nous gateway from the picker selection, not only use_gateway
browser_exec with browser.cloud_provider: nous (the hermes tools picker row) fell into
the direct-API Browser Use branch and reported chrome-not-running, because
_resolve_backend_cdp gated on _use_gateway(), which only read the pre-picker
use_gateway: true flag. Recognize the picker selection too.
2026-09-13 14:41:26 -07:00
teknium1 50e5ffa346 fix(voice): retry a timed-out audio input stream start once
On WSL2 the only input device is the ALSA->PulseAudio bridge; with the
WSLg RDP source SUSPENDED the first InputStream.start() can exceed
PortAudio's 1 s thread-start window and fail with paTimedOut (-9987).
The failed open itself wakes the bridge, which is why the user's second
key press always worked. Retry the open exactly once when the error is a
timeout, on every platform: no WSL detection, no external parecord
warm-up. Any other error, or a second timeout, raises the same
RuntimeError as before.

Generic slim redo of #109313 by @liuhao1024 (WSL-gated parecord warm-up
and retry); diagnosis by @rugscan2021 in #109303.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-13 14:39:29 -07:00
kshitijk4poor 527bfe3cc0 test(mcp): drive OAuth profile isolation through register_mcp_servers()
The regression only exercised register_connected_into_current_scope() and
_select_new_servers() directly. Closed #109430 covered the user-visible
path: profile B, driven through the public register_mcp_servers() entry
point, must open its own connection (its own OAuth token) instead of
adopting A's session. Fold that drive into the existing test rather than
adding a second one.

The direct _select_new_servers() assertion is dropped because it marks
B's key as connecting as a side effect, which would make the subsequent
entry-point drive skip the server; the end-to-end drive subsumes it.
2026-09-13 14:38:20 -07:00
KoNit-K bbe4089405 fix(mcp): isolate mTLS connection identities 2026-09-13 14:38:20 -07:00
Amandeep Khurana 399238f2c2 fix(mcp): isolate OAuth connections by profile 2026-09-13 14:38:20 -07:00
teknium1 643b3f450d fix(tools): stop check_fns swallowing resolver crashes into "returned False"
check_vision_requirements (and five siblings: browser_vision, image/video
generation, x_search, browser_vault) wrapped their whole probe in
`except Exception: return False`. The registry then logged "returned False",
indistinguishable from an unconfigured backend, and the only diagnostic for a
crashed resolver was gone (#87950: named custom provider lookup failing in a
long-lived multi-profile process, reported as vision tools silently vanishing).

The registry owns the verdict: _run_check_fn_uncached and _check_fn_cached both
catch, log with traceback, and return False. Let the exception reach them.
Behaviour for the model is unchanged (tool hidden either way); agent.log now
says why.
2026-09-13 12:35:36 -07:00
teknium1 aa7980d777 fix(tools): log the traceback when a check_fn raises on the cached path
The TTL-cached availability path collapsed a raising check_fn into the bare
verdict "check_fn X raised; dependent tools will be unavailable this turn" with
no exception attached, while the uncached path already logged exc_info. A probe
that crashes is a bug in the probe or its resolver, and without the traceback
the log reads exactly like "nothing configured" — that gap turned #87950 from a
ten-minute diagnosis into a multi-day one. Carry the exception out of the
try block and attach it to the verdict.
2026-09-13 12:35:36 -07:00
teknium1 093c58ca77 test: cover adapter-wrapped probe stubs and repeat-probe stability
Trim the salvaged test to the ≤2-invariant bar: one parametrized test asserting
repeat probes stay resolvable (the user-visible contract from #87654) and that
neither a bare stub nor a Codex-adapter-wrapped stub lands under the runtime
cache key. The guard now keys on aux_probe_mode being active rather than on the
client's type, which is what covers the wrapped case.
2026-09-13 11:22:08 -07:00
Steve Ahlstrom a714ff9823 fix(aux): stop probe stubs poisoning the client cache
`_store_cached_client()` refuses an `_AuxProbeClientStub`, but
`_get_cached_client()` assigns to `_client_cache` directly and so never
reaches that guard. check_fns resolve through this path inside
`aux_probe_mode()` during tool-schema assembly, and the cache key carries no
probe/runtime distinction — so the stored stub is returned to the next real
caller sharing that key, which dies on attribute access with
`_AuxProbeClientStub used as a real client (attribute 'chat')`.

The `async_mode` field in the cache key is what kept this latent: the probe
caches the sync variant, so async consumers (`analyze_image`) miss the entry
and build a real client, while sync consumers (`browser_vision`) hit the
poisoned one and fail on every call.

Observed against a local OpenAI-compatible vision endpoint, where
`browser_vision` failed every call with that RuntimeError while
`analyze_image` against the same provider worked.

Guard the inline store the same way `_store_cached_client()` does, and
return the stub to the probe caller without caching it.

The existing `test_probe_stub_never_cached` pins the invariant only on
`_store_cached_client()`, which is why the unguarded door went unnoticed;
the new test exercises `_get_cached_client()` and fails without this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 11:22:08 -07:00
Teknium adaa6643b9 test(vision): import HEIC helpers from the defining sibling module
The Sep-2026 facade decomposition moved _detect_image_mime_type_from_bytes and
_normalize_to_supported_image into tools/vision_tools_image_prep.py; import from
the defining module per repo convention (compat pointers are off limits in-tree).
2026-09-13 09:18:54 -07:00
Sam Painter c83c413214 fix(vision): fail closed on a malformed ftyp box size
Self-review finding (hostile redteam pass), and it defeats the exact guarantee
the previous commit advertised.

The brand scan is bounded by the declared ftyp box size so a token in a FOLLOWING
box cannot upgrade a HEIC to AVIF. But the bound fell back to the whole 64-byte
sniff window whenever the declared size was outside `16 <= size <= len(header)`.
That failed OPEN in precisely the attacker-controlled case:

    declared size  0 -> image/avif   (genuine HEIC, 'avif' in a later box)
    declared size  4 -> image/avif
    declared size  8 -> image/avif
    declared size 12 -> image/avif
    declared size 15 -> image/avif
    declared size 24 -> image/heic   (truthful size, correct)

A size too small to hold any compatible brand now scans none, rather than
scanning everything. A size >= 16 that overruns the sniffed window is a truncated
read rather than an attack, so that case clamps to the bytes actually available
and a genuine oversized-ftyp AVIF is still detected.

The prior test only exercised the truthful size, so it proved the happy path
rather than the invariant it was named for. Added parameterized malformed-size
cases (0/1/4/8/12/15), a misaligned-size case (17-23, non-multiples of 4), and an
honest-oversized case. All six malformed-size assertions fail if the fail-open
bound is restored; verified by mutation, not assumed.
2026-09-13 09:18:54 -07:00
Sam Painter f5a7dc14c7 fix(vision): parse the ftyp compatible-brand list; don't gate AVIF on pillow-heif
Two defects in the HEIF/AVIF sniffing added by this PR.

1. Major-brand-only detection mislabeled AVIF as HEIC.
   AVIF encoders routinely stamp the generic still-image brand 'mif1' as the
   MAJOR brand and declare the AV1 codec only in the compatible-brand list, so
   every mif1-major file was reported as HEVC-coded HEIC. Parse the
   compatible-brand list (bounded by the declared ftyp box size, so an 'avif'
   token in a following box can't upgrade a real HEIC) and let AV1 brands win
   when both families appear.

2. AVIF was rejected outright when pillow-heif was missing.
   Pillow >= 11.3 bundles a native AvifImagePlugin, while pillow-heif wheels
   are commonly built with no AV1 codec at all (libheif_info() reports
   AVIF: ''). Gating AVIF on a pillow-heif import therefore refused files
   Pillow could already decode, and pointed the user at a library that cannot
   decode them. Registration is now best-effort: the decode attempt is the
   arbiter, and failures emit codec-specific guidance naming the backend that
   actually serves that format.

Tests: mif1+avif and mif1+av01 regressions, a box-size bound guard, a
truncated/bogus-size guard, AVIF-converts-without-pillow-heif, and AVIF error
text. Verified against real pillow-heif-encoded HEIC and Pillow-encoded AVIF
files, not just synthetic headers; both new brand tests fail if the
compatible-brand scan is reverted.
2026-09-13 09:18:54 -07:00
Samuel Painter 0a545f7fa5 feat(vision): decode HEIF/HEIC/AVIF images (iPhone photos)
vision_analyze rejected iPhone photos with 'source is not a recognized
image'. iPhones capture HEIC (ISO-BMFF/HEVC) and upload pipelines often
mislabel it .jpg; the magic-byte sniff had no HEIF branch, so the bytes
were rejected before reaching normalization.

- _detect_image_mime_type_from_bytes: sniff the ISO-BMFF 'ftyp' box and
  recognize HEIF/HEIC (heic/heix/heim/heis/hevc/hevx/mif1/msf1) and AVIF
  (avif/avis) brands, returning image/heic or image/avif. Non-image ftyp
  brands (mp4, ...) stay unrecognized.
- _normalize_to_supported_image: register the pillow-heif Pillow opener,
  then re-encode HEIF/AVIF to PNG before embed (same soft-dependency
  pattern as SVG rasterization). Missing decoder yields an actionable
  'pip install pillow-heif' error, not a generic failure.
- Add pillow-heif==1.4.0 as a core dep alongside Pillow (prebuilt wheels
  bundle libheif; no system libs needed).
- Tests: brand detection (incl. AVIF + mp4-not-misdetected), end-to-end
  HEIC->PNG normalization, and the decoder-missing error path.

The vision resolver's magic-byte sniff stays authoritative (no extension
trust); HEIF now reaches the existing normalize step instead of 400ing.
2026-09-13 09:18:54 -07:00
teknium1 d10bb2ab6f test: make tests/ mirror the source tree; drop issue numbers from filenames
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.

Parallel directories for one source package, folded into the mirror:
  tests/acp        -> tests/acp_adapter   (its __init__/conftest move with it)
  tests/cli        -> tests/hermes_cli    (prompt_toolkit fixture merged into
                                           hermes_cli/conftest.py)
  tests/run_agent  -> tests/agent         (backoff fixture becomes
                                           agent/conftest.py)
  tests/relay      -> tests/gateway/relay
  tests/state      -> tests/hermes_state

246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.

Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.

Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).

Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
2026-09-13 09:18:02 -07:00
teknium1 735831f776 fix(mcp): reconcile chore lives in run_profile_reconcile; prune lazy + mid-connect servers
The MCP config reconciler was appended to the gateway/run.py facade; it moves to
gateway/run_profile_reconcile.py, which already owns post-boot MCP discovery, and
run.py keeps only the chore-table entry.

reconcile_mcp_servers_with_config() also drops a schema-cache (lazy) registration
whose entry is gone (its cached tools would otherwise stay callable and spawn the
server on first use) and reports a dropped server still mid-connect as "pending";
the chore retries on the next tick without waiting for another config edit.

test_cron_delivery_housekeeping neutralizes the chore: it pins the exact
scope/drain sequence of the housekeeping loop and the new chore enters each
profile's scope once per tick.
2026-09-13 06:29:12 -07:00
Teknium 4beb7e29a2 fix(mcp): unattended paths never open browser OAuth; gateway follows mcp_servers edits
A gateway with an OAuth MCP server whose refresh token expired opened a new
authorize tab every 300s, all night (92 tabs). Four defects stacked:

- The parked-server self-probe re-entered the SDK's authorization-code flow
  with interactive OAuth enabled. The timed wake is unattended by definition:
  `_wait_for_reconnect_or_shutdown` now distinguishes "self-probe" from an
  explicit "reconnect", and `_park` flips the task-local
  `_oauth_interactive_enabled` off before a self-probe revival.
- Gateway MCP discovery (startup, `/reload-mcp`, hot-added multiplex
  profiles) ran interactive, unlike the CLI's background discovery. All three
  now run under `suppress_interactive_oauth()`; an expired token parks with
  the `hermes mcp login` hint instead of a browser.
- `_is_interactive()` trusted `sys.stdin.isatty()`, which the Windows CRT
  reports True for a DEVNULL/detached stdin. `_stdin_is_console()` confirms
  with `GetConsoleMode` on Windows.
- Removing an `mcp_servers` entry (or `enabled: false`) never reached a
  running gateway; the parked server probed forever. New
  `reconcile_mcp_servers_with_config()` tears down dropped/disabled servers
  (via `shutdown_mcp_servers(names=...)`) and connects new ones; a
  housekeeping chore runs it when config.yaml's (mtime, size) changes.
  `_select_new_servers` also stops nudging disabled parked servers.

Fixes #81830. Fixes the browser-storm item of #96320.
2026-09-13 06:29:12 -07:00
teknium1 3e066dfedd fix(computer_use): approval goes through the shared gate; no callback now fails closed
computer_use kept its own approval decision: two module dicts
(_session_auto_approve / _always_allow) mirroring tools.approval's
session store and _persist_choice, a private verdict vocabulary
(approve_once/approve_session/always_approve) that hermes_cli mapped
back to once/session/always, and — the real problem — `if
_approval_callback is None: return None`. Only the interactive CLI ever
installed that callback, so every other host (gateway turns, cron,
api_server, tui_gateway, ACP) ran destructive desktop input with no
approval at all, ignoring cron_mode / unattended_mode / the permanent
allowlist, and "always" grants were invisible to `is_approved`,
`clear_session` and the messaging-platform approval buttons.

_request_approval now calls tools.approval._run_approval_gate with
pattern_key `cua:<action>:<background|foreground>` (the old scope shape,
so a background grant still never covers the visible foreground variant)
and fail_closed_when_no_human=True, the same posture as
request_tool_approval / the SSH-config write gate. The private dicts,
their release/atexit clearing, the verdict mapping in
hermes_cli/cli_modal_mixin.py and the extra callback install in cli.py
are deleted: the CLI's terminal_tool callback answers computer_use
prompts like any other tool. set_approval_callback stays as an optional
explicit-callback hook with the shared callback contract
(cb(command, description, **kw) -> once|session|always|deny|timeout);
no in-tree host uses it.

Behavior change:
- No approval callback and no gateway (cron, api_server/webhook,
  headless -q, plain library use): destructive actions are now REFUSED
  with a BLOCKED error and never reach the backend. Previously they
  silently ran. cron honors approvals.cron_mode, unattended platforms
  approvals.unattended_mode, -q approvals.single_query_mode.
- --yolo / gateway /yolo / approvals.mode: off still allow (unchanged).
- Gateway sessions (Telegram/Discord/Slack/...) now get a real pending
  approval with once/session/always buttons instead of default-allow.
- session/always grants live in tools.approval's store; "always" is one
  command_allowlist entry (`cua:click:background`) and is scoped to that
  action+mode — the old blanket "always_approve unlocks everything for
  the session" no longer exists.
- Denial wording is the shared gate's ("BLOCKED: User denied ...",
  "BLOCKED: Action timed out ..."); the error JSON keeps `action`.

Tests: tests/tools/test_computer_use_approval_isolation.py
::test_no_callback_refuses_unless_yolo (blocked + no backend call, then
yolo executes) and ::test_always_grant_lands_in_the_shared_store
(is_approved sees the cua:<action>:<mode> key; second call served from
the store). Sabotage: restoring the `callback is None -> allow`
short-circuit fails the first; swapping the shared gate for a private
grant set fails the second plus the three delivery-ladder scope tests.
tests/tools/conftest.py gains `grant_computer_use_approvals` for
dispatch tests that only care about routing.
2026-09-13 05:21:02 -07:00
teknium1 398234748f refactor(agent): one Retry-After parser and one reset-grammar table feed every retry wait
Seven sites hand-rolled `float(headers.get("Retry-After"))` (anon_auth,
shared_metrics_sender, gemini_native_adapter, extract_api_error_context,
nous_rate_guard, skills_hub_github, skills_hub_clawhub x2) and silently
dropped RFC 7231 HTTP-date values that the conversation loop already honours
via agent/retry_utils.py::parse_retry_after_seconds. They now call it; per-site
caps/floors stay at the call site.

The free-text "resets in / quotaResetDelay / retry after N s" regexes lived in
two tables (agent_runtime_helpers vs credential_pool) whose "resets in"
grammars diverged: the pool accepted only integer `Nhr Nmin` while the error
context accepted h/hr/hours + m/min/minutes + s/seconds with decimals. One table
(agent/retry_utils.py::RETRY_DELAY_PATTERNS / reset_delay_from_message) using
the wider grammar, so a pooled credential's cooldown and the UI's reset time
now agree.
2026-09-13 05:09:43 -07:00
teknium1 469a87f7a5 refactor(config): collapse thin _load_config copies onto the canonical readers
doctor_live and kanban_decompose carried byte-identical
`try: load_config() or {}` wrappers; local_models wrapped load_config in
_quiet; each is now a direct load_config_readonly() call (read-only callers;
the canonical already fails open and returns a mapping). Tests that patched the
local wrappers patch hermes_cli.config.load_config_readonly instead.

tools/code_execution_tool._load_config read the RAW file, so a managed-pinned
`code_execution.mode` and the DEFAULT_CONFIG keys were invisible at tool
discovery — it now reads load_config_readonly() (behavior change: the managed
overlay applies to execute_code's mode/timeout). onboarding.mark_seen and
credential_lifecycle's config mirror scrub parsed config.yaml with a bare
safe_load; both are read→mutate→write round-trips and use read_user_config_raw,
the documented write-back primitive.
2026-09-13 05:09:06 -07:00