The "screen unchanged" result points the model at its previous capture. After
context compression that capture may be summarized away, so the note would refer
to pixels no longer in context. Mirror read_file's reset_file_dedup: the
compaction boundary (both the summary path and the codex app-server path) now
clears the session's screenshot digest, and the first capture afterwards delivers
the image again even when the screen is byte-identical.
Key the screenshot-dedup state by the same profile-scoped session id the
backend cache uses, so two multiplexed profiles sharing a session id (or a
DISPLAY) never dedup against each other's frames, and forget the state in
release_computer_use_session so a re-created session's first capture
always delivers pixels. Reword the unchanged note to cover the aux-vision
path (where the prior result was an analysis, not an image). Tests: the
dispatch path (explicit capture + capture_after) honours the streak cap;
release forgets state.
Port from openclaw/openclaw#129924: a capture whose pixels are
byte-identical to the previous capture of the same target in the same
session returns its full text metadata (element index included) plus an
explicit 'screen unchanged' note instead of the multimodal image block.
Adapted for Hermes: openclaw gates dedup on per-frame context-epoch
tracking; Hermes bounds staleness with a consecutive-omission streak cap
(2) so full pixels are re-delivered before compaction could evict the
referenced image. Dedup state is per-session (no cross-session leaks),
append-only (no history rewrites — prompt cache prefixes untouched),
and skipped entirely when no session_id is present.
tests/contracts -> tests/tui_gateway/contracts (tree-layout rule: tests mirror a source
package). test_rpc_params_cannot_spoof_runtime_artifacts: forged owner_transport /
owner_session_record / owner_token keys are now refused at the wire (4000 + key path)
instead of silently dropped before the handler; the invariant (no steer reaches the
agent) is unchanged and asserted directly.
The rebased synthetic guard built a bare dict that main's _build_payload
no longer accepts (it indexes aspect_ratios/resolutions directly) and never
exercised generate(); it now builds the family via _family() and asserts
generate() returns modality_unsupported without submitting. The surface
matrix's i2v-only parametrization collected zero cases after Gemini Omni
Flash 1.1 gained t2v, so it is removed along with the dead branch in the
t2v matrix.
_attach_child now mirrors a pending parent stop with the same split
interrupt() uses for its own fan-out (hard -> hard_interrupt, soft ->
interrupt), so a redirect is not turned into a cancel on a child that was
attached late. _restore_parent_cancellation collapses to re-attaching the
rejected unit's children: the replay is the attach step's job now.
Test fixture: _Batch gained origin_session_history_delivery on main after
the salvaged PR was written.
Co-authored-by: illidan <noequal666@gmail.com>
Rebased onto main (~9700 commits): the branch predates the god-file
decomposition and the tests/state -> tests/hermes_state move.
Fixes for seams that moved:
- cron memory contract patches cron.scheduler_delivery._resolve_origin and
hermes_state_registry.acquire (where run_job now reads them).
- state.db conformance imports _live_writer_holds_db from
hermes_state_repair and drives the public create_quick_snapshot.
- update receipt: serve runtimes are reconciled in their own unit
vocabulary since #100479, so the "full accounting" row names the serve
unit instead of borrowing a gateway relaunch.
Deleted (change-detectors / source-greps / duplicates / dead code):
- source-text scan of hermes_state*.py for "maintenance-shaped" defs and
the symbol REGISTRY it fed (renames are not regressions).
- per-op exact-outcome table under a live writer -> one invariant: refuse
with a lock error or report zero work, DB stays intact.
- copy_db_and_verify pins: the symbol is a revert-scheduled plugin-compat
pointer, not production code (check_compat_pointers.py).
- cron ON-direction tests already pinned by tests/cron/test_scheduler.py,
plus a tautology that never called production code.
- exact warning-wording asserts in the env deprecation truth table.
- same-file duplicates: fixed-point (implied by idempotence), tripwire
round-trips, absent-registry, sequential "race" re-enactment.
Two weeks of closed issues/merged PRs show the same areas regenerating:
each salvage pinned its instance while the class invariant had no test.
These suites pin the invariants themselves:
- tests/conformance/test_profile_write_tripwire.py — no writes to the
default profile tree while a profile is active (#88532#92662#89190#89625#92156); reusable tripwire fixture, 4 surfaces
- tests/hermes_cli/test_env_deprecation_truthtable.py — 18-row truth
table for the Deprecated-.env warning (#88829#89016#89389#90299)
- tests/cron/test_cron_memory_contract.py — cron<->memory contract that
flipped twice in Aug (#91269 -> #91384 -> #91447)
- tests/agent/test_injected_param_strip_retry_registry.py — every
strippable injected param x real 400 shapes must strip-and-retry;
unknown params must still fail (#90257#89897#91164#89503)
- tests/agent/test_transcript_decoration_idempotence.py — f(f(x))==f(x)
law + 4-breakpoint budget for apply_anthropic_cache_control (#90971)
- tests/state/test_state_db_maintenance_conformance.py — registry-
enumerated maintenance ops refuse/degrade under a live writer; copies
of corrupt DBs are refused or flagged (#91839#90806#90613#88235)
- tests/tools/test_bot_mode_canonical_chat_resolution.py — canonical
Bot Chat resolution is idempotent, never mints, unique per profile,
race-safe (#92040#90705#92692#90005#90732, PR #92129)
- tests/hermes_cli/test_update_receipt_truthfulness.py — receipts:
crash never claims success; success requires full fleet accounting;
refusal != failure (#91283#91439#92902#92780)
117 tests, all sabotage-verified (each suite proven to FAIL when its
bug class is reintroduced).
data:image/... entries were treated as local paths and silently skipped as
"unreadable". They now ride as image_url parts only (never pasted into the
text hint, never appended to a text-mode goal). Skips and forwarding
failures log at warning since the caller explicitly asked for the images;
decide_image_input_mode gets the child's requested_provider like the CLI
and gateway callers.
Tests collapse to five invariants, including one that drives _ChildRun
and asserts the multimodal content list reaches run_conversation as the
first user turn. Docs mention data: URLs and the read guard.
Subagents can now SEE images. Each delegate_task task accepts an optional
images list (max 8; local paths or http(s) URLs). Vision-capable children
receive native image_url content parts on their goal turn (local files as
data URLs, remote URLs verbatim); non-vision children get
[Image attached at: ...] hints plus a vision_analyze pointer. Routing
reuses agent.image_routing (decide_image_input_mode /
build_native_content_parts), so agent.image_input_mode governs delegation
exactly like inbound gateway images.
Best-effort by contract: malformed images arrays fail the call loudly
before any child spawns; unreadable paths are skipped with a log line;
any exception in the forwarding path degrades to the text-only goal.
Adapted from RooCodeInc/Roomote#1796 / #1767 (Fast agent forwards bounded
current-turn attachments to delegated coding tasks).
Drop the no-op and still-running change-detectors (4 invariant tests remain:
pipe closed on finish, PTY closed on finish, poll still serves buffered
output, prune releases handles). The accretion-caps fake session now
carries process/_pty like the real dataclass, since prune reads them.
Widen #75162: _prune_if_needed() drops finished sessions (TTL expiry and
oldest-finished eviction at MAX_PROCESSES) — release their Popen/PTY
handles there too, covering sessions inserted into _finished without
passing through _move_to_finished(). The release helper is idempotent,
so double-close on the normal path is a no-op. Adds two tests: prune
releases handles of dropped sessions, and a still-running session's
pipe stays open.
Finished sessions retained their subprocess.Popen pipe objects (and PTY
masters) until the finished-process TTL (FINISHED_TTL_SECONDS, default 30
minutes) elapsed. Under heavy background churn — deployments, archivers,
watchers — finished-but-unpruned sessions accumulated one open pipe FD
each, exhausting the gateway process's file descriptor budget and
surfacing as a 'file descriptor limit' error on new background spawns.
The registry never rejects spawns (it prunes oldest-finished at
MAX_PROCESSES), so the real defect was the retained-handle leak, not a
registry-cap rejection. The fix closes each finished session's Popen
stdout/stderr/stdin streams and PTY master in _move_to_finished(), right
after the reader loop drains EOF. poll()/wait()/read_log() serve output
from the buffered output_buffer — never from the pipe — so the release is
lossless.
Tests: 4 new cases in TestFinishedHandleRelease — Popen pipes closed,
PTY closed, no-handle sessions safe, and poll() still serves buffered
output after the release. All 4 fail on main (reproduction) and pass
with the fix.
Port from nearai/ironclaw#7965: BM25 admits any document scoring above
zero, i.e. sharing ONE term with the query. A long descriptive search
for a capability that does not exist therefore returned a plausible-
looking ranked list, and the model read 'results exist' as 'it is in
here somewhere' and rephrased instead of stopping (IronClaw production
trace: 652 tool calls, 216 of them tool_search, hunting a 'data' tool
that did not exist).
A document must now match at least half the query's ANSWERABLE terms
(terms present anywhere in the index) before it is offered. Coverage
only engages from four answerable terms up, preserving recall on short
queries; exact tool-name matches remain authoritative; the substring
fallback is unchanged.
Docs: relevance-floor bullet added to tool-search.md implementation
details.
Adds fal-ai/kling-image/v3/text-to-image ($0.028/img, native 2K default,
8 aspect ratios) with its image-to-image edit endpoint. The i2i schema
takes a SINGULAR `image_url` string instead of the usual `image_urls`
list, so the catalog gains an `edit_image_param` knob that
_build_fal_edit_payload honors (first source image only); the
edit-contract test now validates whichever image key the entry declares.
uvx.exe (uv's actual Windows shim) and pipx.exe bypassed the MCP OSV
malware check entirely, and backslash-qualified commands only resolved
when running under ntpath. Basename now splits on both separators;
matching stays exact (npx.cmd / uvx.exe / uvx.cmd / pipx.exe) so
lookalikes like npx.exe or npx.cmd.bak remain fail-open.
Port from openclaw/openclaw#123194: a hostile or misbehaving remote MCP
server could stream an unbounded HTTP catalog/tool-result body that the
MCP SDK buffers and JSON-parses before any of Hermes' post-parse limits
(resource cap, tool-result truncation) run.
New _make_mcp_body_cap_transport wraps the owned httpx AsyncClient's
transport on the Streamable HTTP (mcp >= 1.24) and SSE paths:
- finite HTTP bodies capped at 10 MiB (Content-Length rejected up front,
streamed bodies capped chunk-by-chunk);
- each SSE event capped at 10 MiB, with accounting reset at completed
event boundaries so long-lived streams/keepalives are unlimited;
- violations raise httpx.ReadError naming the byte cap, handled by the
existing transport teardown/reconnect path (#66092).
verify/cert now live on the inner AsyncHTTPTransport (client-level TLS
kwargs are inert once a custom transport is passed); the SSE
httpx_client_factory is always injected so the cap applies with default
TLS too. Legacy mcp < 1.24 path (SDK-internal client, no hook) stays
uncapped — same degradation as strict_redirect_headers.
Portable Agent Plugin packages fold the plugin name into the MCP
registry name three times over (slug, digest, and again as the server
key), so mcp__<server>__<tool> routinely exceeds the 64-char function
name limit OpenAI-compatible providers enforce — while the same server
registered via `hermes mcp add` stays well under it. The oversized name
is never rejected loudly; the tool just becomes unreachable. Clamp
mcp_prefixed_tool_name() to 64 chars with a deterministic, collision-
safe hash suffix, mirroring the existing property-key clamp in
schema_sanitizer.py. Dispatch is unaffected since handlers already
close over the original unprefixed tool name.
Fixes#81331
Port from openai/codex#39615: the authorization server discovered for an
MCP server can change (protected-resource metadata edit, server
migration, DNS takeover). Without binding, Hermes would send the stored
refresh token to whatever issuer the server now advertises — handing a
long-lived credential to a different authorization server.
- HermesTokenStorage records hermes_issuer alongside cached tokens
(stripped before OAuthToken.model_validate; never sent on the wire).
- Both provider classes (tools/mcp_oauth.py legacy path and
tools/mcp_oauth_manager.py managed path) stamp the discovered issuer
on every token save and enforce the binding on _initialize.
- On mismatch: refresh token is stripped from memory and disk; the
unexpired access token keeps working; full re-auth happens at expiry.
- Legacy token files without an issuer adopt the current one once
(no forced re-login for existing installs — deliberate divergence
from Codex, which requires reauth).
Validated: 12 new tests + 163 existing MCP OAuth tests green; sabotage
run confirms the new tests fail without the enforcement; E2E against
the real manager provider class with a temp HERMES_HOME confirms
mismatch strips and match preserves.
AIAgent.interrupt() fans the stop out to a snapshot of _active_children. A child
that is attached after that snapshot — an orchestrator subagent still building
its fan-out siblings, or one that has not yet hit its next iteration check —
started with no signal and ran to completion as an orphan while its parent had
already reported `interrupted`. Live repro (mid orchestrator, stop delivered
between grandchild A and B builds): grandchild B kept its `sleep 20` alive and
stayed in the subagent registry after the mid returned.
_attach_child now mirrors a pending parent stop onto the newcomer, so the whole
spawn tree dies with the node that was stopped. Tests: unit (late attach gets
the stop, normal attach does not) + the real _build_child_agent path with a
stopped orchestrator.
The Desktop "Approvals: off" toggle persists approvals.mode: off. The shell
guards (check_all_command_guards / execute_code) honour it as a bypass, but
_run_approval_gate, the shared gate that computer_use, plugin approval rules,
SSH-config writes and the dangerous-pattern prompt all route through, only
checked _yolo_active() (process --yolo / session /yolo). So with approvals
off, every destructive computer_use action still prompted.
Regressed when 3e066dfedd moved computer_use onto the shared gate: its old
private gate never consulted mode at all, and the shared gate had never been
given the third bypass source. Gate now mirrors the shell guards:
yolo OR approvals.mode == "off" -> approved. Hardline blocks and deny rules
still run before it.
_parallel_safe_servers was keyed by the raw server name while _servers
moved to (scope, name) connection keys (ceaf622c6d). With profile A's `x`
serial and profile B's same-named `x` opted into parallel calls, B's
discovery pass flipped A's tool to parallel-safe and the batch planner put
two A calls in one parallel segment against a server that never opted in
(review of #108352, finding B).
The opt-in is now recorded under the discovering profile's own key and
is_mcp_tool_parallel_safe() looks it up under the calling profile's key,
so one profile's policy never reaches another's same-named server.
Single-profile processes keep the bare-name key, byte for byte.
Under gateway.multiplex_profiles a profile whose mcp_servers entry has the
same route and credentials as another profile's live connection adopts that
connection instead of opening its own. _same_server_route() compares only
the connection identity, so a `trust: untrusted` profile adopted a
`trust: full` profile's connection; _trust_gate_check() then resolved the
OWNER's connection key, read `full`, and let the untrusted profile run
write-capable tools without the approval prompt its config demands
(review of #108352, finding A; regression from ceaf622c6d, where the
name-keyed ledger let the last registrant's trust win instead).
`trust` is the consuming profile's policy, not a property of the
connection: _server_trust_levels is now keyed by the calling profile's own
key (recorded at its own registration and at adoption, dropped when its
overlay is removed), while readOnlyHint stays under the connection key
because it describes the server's tools. Sharing the connection is still
allowed — only the gate is per profile.
Skipping `tests/`, `spec/`, ... in EXCLUDED_DIRS made those trees
invisible to the guard, but `plugins_loader._load_directory_module`
sets `submodule_search_locations=[plugin_dir]`, so a plugin
`__init__.py` doing `from .tests import evil` imports and runs whatever
lives there: a `tests/evil.py` with a destructive root remove scanned
`dangerous` on main and `safe` on this branch. `_walk` also matched the
names at any depth, so `src/spec/handler.py` — plain runtime code — went
unscanned.
Keep scanning everything; instead cap a critical finding located under a
ROOT-level test dir at `high`, so the verdict is `caution` (confirmation
required, `--force` overridable) rather than the un-overridable
`dangerous`. Fixture strings still cannot brick an install, which was
the reported problem, while a critical in any runtime file (`setup.sh`,
`src/spec/...`) still yields `dangerous`. Trade-off stated in the PR
body: hostile code deliberately placed under `tests/` is now
force-installable rather than blocked outright.
Docs no longer claim test code never runs.
plugin_guard walks the whole plugin clone, and EXCLUDED_DIRS skipped
caches and vendored dirs but not tests/. A security-conscious plugin's
test suite SHOULD contain adversarial fixtures — a test asserting the
trust boundary holds round-trips the injection string verbatim — and
any single critical finding makes the verdict dangerous, which --force
explicitly cannot override. Scanning tests therefore made exactly the
plugins that test their security unconditionally uninstallable, and the
only workaround was obfuscating the payload strings, weakening the
tests and inverting the incentive. Fixtures are never loaded into an
agent's context at runtime the way README/plugin.yaml are.
Add the conventional test/spec/fixture directory names to
EXCLUDED_DIRS, alongside the existing cache/vendored skips.
Both process-global caches were keyed by the caller's session/task id alone, so under
gateway.multiplex_profiles two profiles using the same id — a shared `browser_exec session=`
name, or two Hermes sessions whose screens report the same DISPLAY — resolved to the FIRST
profile's cloud browser / cua-driver, and a command issued in one bot's chat could act on
another bot's screen.
The key now carries the routed profile's home key whenever a served-profile scope is active
(`get_hermes_home_override()` set), the same shape `tools/approval.py::_baseline_key` and the
camofox/cloud caches already use; outside a scope every key is byte-identical to before. The
computer_use lookup, install and release paths all go through one `_scoped_sid`, so a release
under profile B never stops profile A's driver; approval-bypass state keeps the bare session id.
Fixes#110032 (report by @wolfyy970, from @vandaimer's manual test on #108914).
browser_exec with browser.cloud_provider: nous (the hermes tools picker row) fell into
the direct-API Browser Use branch and reported chrome-not-running, because
_resolve_backend_cdp gated on _use_gateway(), which only read the pre-picker
use_gateway: true flag. Recognize the picker selection too.
On WSL2 the only input device is the ALSA->PulseAudio bridge; with the
WSLg RDP source SUSPENDED the first InputStream.start() can exceed
PortAudio's 1 s thread-start window and fail with paTimedOut (-9987).
The failed open itself wakes the bridge, which is why the user's second
key press always worked. Retry the open exactly once when the error is a
timeout, on every platform: no WSL detection, no external parecord
warm-up. Any other error, or a second timeout, raises the same
RuntimeError as before.
Generic slim redo of #109313 by @liuhao1024 (WSL-gated parecord warm-up
and retry); diagnosis by @rugscan2021 in #109303.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
The regression only exercised register_connected_into_current_scope() and
_select_new_servers() directly. Closed#109430 covered the user-visible
path: profile B, driven through the public register_mcp_servers() entry
point, must open its own connection (its own OAuth token) instead of
adopting A's session. Fold that drive into the existing test rather than
adding a second one.
The direct _select_new_servers() assertion is dropped because it marks
B's key as connecting as a side effect, which would make the subsequent
entry-point drive skip the server; the end-to-end drive subsumes it.
check_vision_requirements (and five siblings: browser_vision, image/video
generation, x_search, browser_vault) wrapped their whole probe in
`except Exception: return False`. The registry then logged "returned False",
indistinguishable from an unconfigured backend, and the only diagnostic for a
crashed resolver was gone (#87950: named custom provider lookup failing in a
long-lived multi-profile process, reported as vision tools silently vanishing).
The registry owns the verdict: _run_check_fn_uncached and _check_fn_cached both
catch, log with traceback, and return False. Let the exception reach them.
Behaviour for the model is unchanged (tool hidden either way); agent.log now
says why.
The TTL-cached availability path collapsed a raising check_fn into the bare
verdict "check_fn X raised; dependent tools will be unavailable this turn" with
no exception attached, while the uncached path already logged exc_info. A probe
that crashes is a bug in the probe or its resolver, and without the traceback
the log reads exactly like "nothing configured" — that gap turned #87950 from a
ten-minute diagnosis into a multi-day one. Carry the exception out of the
try block and attach it to the verdict.
Trim the salvaged test to the ≤2-invariant bar: one parametrized test asserting
repeat probes stay resolvable (the user-visible contract from #87654) and that
neither a bare stub nor a Codex-adapter-wrapped stub lands under the runtime
cache key. The guard now keys on aux_probe_mode being active rather than on the
client's type, which is what covers the wrapped case.
`_store_cached_client()` refuses an `_AuxProbeClientStub`, but
`_get_cached_client()` assigns to `_client_cache` directly and so never
reaches that guard. check_fns resolve through this path inside
`aux_probe_mode()` during tool-schema assembly, and the cache key carries no
probe/runtime distinction — so the stored stub is returned to the next real
caller sharing that key, which dies on attribute access with
`_AuxProbeClientStub used as a real client (attribute 'chat')`.
The `async_mode` field in the cache key is what kept this latent: the probe
caches the sync variant, so async consumers (`analyze_image`) miss the entry
and build a real client, while sync consumers (`browser_vision`) hit the
poisoned one and fail on every call.
Observed against a local OpenAI-compatible vision endpoint, where
`browser_vision` failed every call with that RuntimeError while
`analyze_image` against the same provider worked.
Guard the inline store the same way `_store_cached_client()` does, and
return the stub to the probe caller without caching it.
The existing `test_probe_stub_never_cached` pins the invariant only on
`_store_cached_client()`, which is why the unguarded door went unnoticed;
the new test exercises `_get_cached_client()` and fails without this change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The Sep-2026 facade decomposition moved _detect_image_mime_type_from_bytes and
_normalize_to_supported_image into tools/vision_tools_image_prep.py; import from
the defining module per repo convention (compat pointers are off limits in-tree).
Self-review finding (hostile redteam pass), and it defeats the exact guarantee
the previous commit advertised.
The brand scan is bounded by the declared ftyp box size so a token in a FOLLOWING
box cannot upgrade a HEIC to AVIF. But the bound fell back to the whole 64-byte
sniff window whenever the declared size was outside `16 <= size <= len(header)`.
That failed OPEN in precisely the attacker-controlled case:
declared size 0 -> image/avif (genuine HEIC, 'avif' in a later box)
declared size 4 -> image/avif
declared size 8 -> image/avif
declared size 12 -> image/avif
declared size 15 -> image/avif
declared size 24 -> image/heic (truthful size, correct)
A size too small to hold any compatible brand now scans none, rather than
scanning everything. A size >= 16 that overruns the sniffed window is a truncated
read rather than an attack, so that case clamps to the bytes actually available
and a genuine oversized-ftyp AVIF is still detected.
The prior test only exercised the truthful size, so it proved the happy path
rather than the invariant it was named for. Added parameterized malformed-size
cases (0/1/4/8/12/15), a misaligned-size case (17-23, non-multiples of 4), and an
honest-oversized case. All six malformed-size assertions fail if the fail-open
bound is restored; verified by mutation, not assumed.
Two defects in the HEIF/AVIF sniffing added by this PR.
1. Major-brand-only detection mislabeled AVIF as HEIC.
AVIF encoders routinely stamp the generic still-image brand 'mif1' as the
MAJOR brand and declare the AV1 codec only in the compatible-brand list, so
every mif1-major file was reported as HEVC-coded HEIC. Parse the
compatible-brand list (bounded by the declared ftyp box size, so an 'avif'
token in a following box can't upgrade a real HEIC) and let AV1 brands win
when both families appear.
2. AVIF was rejected outright when pillow-heif was missing.
Pillow >= 11.3 bundles a native AvifImagePlugin, while pillow-heif wheels
are commonly built with no AV1 codec at all (libheif_info() reports
AVIF: ''). Gating AVIF on a pillow-heif import therefore refused files
Pillow could already decode, and pointed the user at a library that cannot
decode them. Registration is now best-effort: the decode attempt is the
arbiter, and failures emit codec-specific guidance naming the backend that
actually serves that format.
Tests: mif1+avif and mif1+av01 regressions, a box-size bound guard, a
truncated/bogus-size guard, AVIF-converts-without-pillow-heif, and AVIF error
text. Verified against real pillow-heif-encoded HEIC and Pillow-encoded AVIF
files, not just synthetic headers; both new brand tests fail if the
compatible-brand scan is reverted.
vision_analyze rejected iPhone photos with 'source is not a recognized
image'. iPhones capture HEIC (ISO-BMFF/HEVC) and upload pipelines often
mislabel it .jpg; the magic-byte sniff had no HEIF branch, so the bytes
were rejected before reaching normalization.
- _detect_image_mime_type_from_bytes: sniff the ISO-BMFF 'ftyp' box and
recognize HEIF/HEIC (heic/heix/heim/heis/hevc/hevx/mif1/msf1) and AVIF
(avif/avis) brands, returning image/heic or image/avif. Non-image ftyp
brands (mp4, ...) stay unrecognized.
- _normalize_to_supported_image: register the pillow-heif Pillow opener,
then re-encode HEIF/AVIF to PNG before embed (same soft-dependency
pattern as SVG rasterization). Missing decoder yields an actionable
'pip install pillow-heif' error, not a generic failure.
- Add pillow-heif==1.4.0 as a core dep alongside Pillow (prebuilt wheels
bundle libheif; no system libs needed).
- Tests: brand detection (incl. AVIF + mp4-not-misdetected), end-to-end
HEIC->PNG normalization, and the decoder-missing error path.
The vision resolver's magic-byte sniff stays authoritative (no extension
trust); HEIF now reaches the existing normalize step instead of 400ing.
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
The MCP config reconciler was appended to the gateway/run.py facade; it moves to
gateway/run_profile_reconcile.py, which already owns post-boot MCP discovery, and
run.py keeps only the chore-table entry.
reconcile_mcp_servers_with_config() also drops a schema-cache (lazy) registration
whose entry is gone (its cached tools would otherwise stay callable and spawn the
server on first use) and reports a dropped server still mid-connect as "pending";
the chore retries on the next tick without waiting for another config edit.
test_cron_delivery_housekeeping neutralizes the chore: it pins the exact
scope/drain sequence of the housekeeping loop and the new chore enters each
profile's scope once per tick.
A gateway with an OAuth MCP server whose refresh token expired opened a new
authorize tab every 300s, all night (92 tabs). Four defects stacked:
- The parked-server self-probe re-entered the SDK's authorization-code flow
with interactive OAuth enabled. The timed wake is unattended by definition:
`_wait_for_reconnect_or_shutdown` now distinguishes "self-probe" from an
explicit "reconnect", and `_park` flips the task-local
`_oauth_interactive_enabled` off before a self-probe revival.
- Gateway MCP discovery (startup, `/reload-mcp`, hot-added multiplex
profiles) ran interactive, unlike the CLI's background discovery. All three
now run under `suppress_interactive_oauth()`; an expired token parks with
the `hermes mcp login` hint instead of a browser.
- `_is_interactive()` trusted `sys.stdin.isatty()`, which the Windows CRT
reports True for a DEVNULL/detached stdin. `_stdin_is_console()` confirms
with `GetConsoleMode` on Windows.
- Removing an `mcp_servers` entry (or `enabled: false`) never reached a
running gateway; the parked server probed forever. New
`reconcile_mcp_servers_with_config()` tears down dropped/disabled servers
(via `shutdown_mcp_servers(names=...)`) and connects new ones; a
housekeeping chore runs it when config.yaml's (mtime, size) changes.
`_select_new_servers` also stops nudging disabled parked servers.
Fixes#81830. Fixes the browser-storm item of #96320.
computer_use kept its own approval decision: two module dicts
(_session_auto_approve / _always_allow) mirroring tools.approval's
session store and _persist_choice, a private verdict vocabulary
(approve_once/approve_session/always_approve) that hermes_cli mapped
back to once/session/always, and — the real problem — `if
_approval_callback is None: return None`. Only the interactive CLI ever
installed that callback, so every other host (gateway turns, cron,
api_server, tui_gateway, ACP) ran destructive desktop input with no
approval at all, ignoring cron_mode / unattended_mode / the permanent
allowlist, and "always" grants were invisible to `is_approved`,
`clear_session` and the messaging-platform approval buttons.
_request_approval now calls tools.approval._run_approval_gate with
pattern_key `cua:<action>:<background|foreground>` (the old scope shape,
so a background grant still never covers the visible foreground variant)
and fail_closed_when_no_human=True, the same posture as
request_tool_approval / the SSH-config write gate. The private dicts,
their release/atexit clearing, the verdict mapping in
hermes_cli/cli_modal_mixin.py and the extra callback install in cli.py
are deleted: the CLI's terminal_tool callback answers computer_use
prompts like any other tool. set_approval_callback stays as an optional
explicit-callback hook with the shared callback contract
(cb(command, description, **kw) -> once|session|always|deny|timeout);
no in-tree host uses it.
Behavior change:
- No approval callback and no gateway (cron, api_server/webhook,
headless -q, plain library use): destructive actions are now REFUSED
with a BLOCKED error and never reach the backend. Previously they
silently ran. cron honors approvals.cron_mode, unattended platforms
approvals.unattended_mode, -q approvals.single_query_mode.
- --yolo / gateway /yolo / approvals.mode: off still allow (unchanged).
- Gateway sessions (Telegram/Discord/Slack/...) now get a real pending
approval with once/session/always buttons instead of default-allow.
- session/always grants live in tools.approval's store; "always" is one
command_allowlist entry (`cua:click:background`) and is scoped to that
action+mode — the old blanket "always_approve unlocks everything for
the session" no longer exists.
- Denial wording is the shared gate's ("BLOCKED: User denied ...",
"BLOCKED: Action timed out ..."); the error JSON keeps `action`.
Tests: tests/tools/test_computer_use_approval_isolation.py
::test_no_callback_refuses_unless_yolo (blocked + no backend call, then
yolo executes) and ::test_always_grant_lands_in_the_shared_store
(is_approved sees the cua:<action>:<mode> key; second call served from
the store). Sabotage: restoring the `callback is None -> allow`
short-circuit fails the first; swapping the shared gate for a private
grant set fails the second plus the three delivery-ladder scope tests.
tests/tools/conftest.py gains `grant_computer_use_approvals` for
dispatch tests that only care about routing.
Seven sites hand-rolled `float(headers.get("Retry-After"))` (anon_auth,
shared_metrics_sender, gemini_native_adapter, extract_api_error_context,
nous_rate_guard, skills_hub_github, skills_hub_clawhub x2) and silently
dropped RFC 7231 HTTP-date values that the conversation loop already honours
via agent/retry_utils.py::parse_retry_after_seconds. They now call it; per-site
caps/floors stay at the call site.
The free-text "resets in / quotaResetDelay / retry after N s" regexes lived in
two tables (agent_runtime_helpers vs credential_pool) whose "resets in"
grammars diverged: the pool accepted only integer `Nhr Nmin` while the error
context accepted h/hr/hours + m/min/minutes + s/seconds with decimals. One table
(agent/retry_utils.py::RETRY_DELAY_PATTERNS / reset_delay_from_message) using
the wider grammar, so a pooled credential's cooldown and the UI's reset time
now agree.
doctor_live and kanban_decompose carried byte-identical
`try: load_config() or {}` wrappers; local_models wrapped load_config in
_quiet; each is now a direct load_config_readonly() call (read-only callers;
the canonical already fails open and returns a mapping). Tests that patched the
local wrappers patch hermes_cli.config.load_config_readonly instead.
tools/code_execution_tool._load_config read the RAW file, so a managed-pinned
`code_execution.mode` and the DEFAULT_CONFIG keys were invisible at tool
discovery — it now reads load_config_readonly() (behavior change: the managed
overlay applies to execute_code's mode/timeout). onboarding.mark_seen and
credential_lifecycle's config mirror scrub parsed config.yaml with a bare
safe_load; both are read→mutate→write round-trips and use read_user_config_raw,
the documented write-back primitive.