606903badc made launch-profile turns bind a file-backed terminal scope as
soon as any secondary home is served, so a poisoned ambient bridge can never
be the launch turn's authority. That scope is rebuilt from defaults +
<home>/.env + config.yaml only, which drops the launch process's legitimate
env-only policy: TERMINAL_ENV=ssh TERMINAL_SSH_HOST=example.test with a `{}`
config.yaml became backend=local, ssh_host='' the moment a second profile
was served.
tui_gateway/launch_terminal_policy.py freezes the process TERMINAL_* once,
in _profile_home right before the first secondary home is registered as
served — the last moment ambient env is provably the launch profile's own.
build_profile_terminal_scope takes that snapshot as a trusted env_overlay
sitting where the process env sits in the standalone bridge (explicit YAML
keys still win). Launch turns overlay the snapshot; ambient os.environ is
never re-read after activation, so a later secondary write is still
rejected, and the scope is reset after the turn as before.
Refs #108440 review (andrexibiza), #107442 (ehz0ah).
Under multiplexing tools/terminal_tool.py no longer bridges a profile's
terminal.* into os.environ while a home override is bound (#108440), so any
secondary-profile entrypoint that binds only the home (+ secrets) leaves
terminal_tool on the launch process's ambient TERMINAL_*: a secondary with
`terminal.backend: docker` ran its prompt.background / prompt.btw /
preview.restart side agents, its eager resume and its session.branch build
on the launch `local` backend.
_spawn_side_agent and _profile_build_scope now enter
_session_profile_runtime_scope, the same home -> secrets -> terminal
composition a prompt turn binds (and gateway/run.py::_profile_runtime_scope
mirrors). The terminal scope is installed for the whole worker/build
lifetime and reset on success and on exception; a malformed secondary
config still yields the fail-closed refusal scope. No ambient os.environ
write is restored.
Refs #108440 review (andrexibiza), #107442 (ehz0ah).
Same class as minimax-portal: aliyun, deep-seek, nim/build-nvidia/nemotron
and vertexai resolve in hermes_cli's alias tables but not in
providers.get_provider_profile(), so a profile lookup keyed on the alias
returned None and lost the profile's wire mode/extra_body/headers. qwen is
left out: the catalog maps it to alibaba while the qwen-oauth plugin already
claims it, a pre-existing disagreement outside this fix.
The two salvaged tests overlapped (resolution subsumes membership); one
contract test that every documented alias resolves to minimax-oauth is
the salvage bar.
check_vision_requirements (and five siblings: browser_vision, image/video
generation, x_search, browser_vault) wrapped their whole probe in
`except Exception: return False`. The registry then logged "returned False",
indistinguishable from an unconfigured backend, and the only diagnostic for a
crashed resolver was gone (#87950: named custom provider lookup failing in a
long-lived multi-profile process, reported as vision tools silently vanishing).
The registry owns the verdict: _run_check_fn_uncached and _check_fn_cached both
catch, log with traceback, and return False. Let the exception reach them.
Behaviour for the model is unchanged (tool hidden either way); agent.log now
says why.
The TTL-cached availability path collapsed a raising check_fn into the bare
verdict "check_fn X raised; dependent tools will be unavailable this turn" with
no exception attached, while the uncached path already logged exc_info. A probe
that crashes is a bug in the probe or its resolver, and without the traceback
the log reads exactly like "nothing configured" — that gap turned #87950 from a
ten-minute diagnosis into a multi-day one. Carry the exception out of the
try block and attach it to the verdict.
Trim the salvaged test to the ≤2-invariant bar: one parametrized test asserting
repeat probes stay resolvable (the user-visible contract from #87654) and that
neither a bare stub nor a Codex-adapter-wrapped stub lands under the runtime
cache key. The guard now keys on aux_probe_mode being active rather than on the
client's type, which is what covers the wrapped case.
`_store_cached_client()` refuses an `_AuxProbeClientStub`, but
`_get_cached_client()` assigns to `_client_cache` directly and so never
reaches that guard. check_fns resolve through this path inside
`aux_probe_mode()` during tool-schema assembly, and the cache key carries no
probe/runtime distinction — so the stored stub is returned to the next real
caller sharing that key, which dies on attribute access with
`_AuxProbeClientStub used as a real client (attribute 'chat')`.
The `async_mode` field in the cache key is what kept this latent: the probe
caches the sync variant, so async consumers (`analyze_image`) miss the entry
and build a real client, while sync consumers (`browser_vision`) hit the
poisoned one and fail on every call.
Observed against a local OpenAI-compatible vision endpoint, where
`browser_vision` failed every call with that RuntimeError while
`analyze_image` against the same provider worked.
Guard the inline store the same way `_store_cached_client()` does, and
return the stub to the probe caller without caching it.
The existing `test_probe_stub_never_cached` pins the invariant only on
`_store_cached_client()`, which is why the unguarded door went unnoticed;
the new test exercises `_get_cached_client()` and fails without this change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An MCP server whose OAuth refresh token died used to get one warning toast
per app session (on the transition into needs-auth) and then nothing: the
server sat parked, silently, until the user happened to open Capabilities →
MCP. A dead token is a standing problem that needs an action, so the health
checker now re-nudges once a day while the server stays broken (persisted
24h snooze per profile+server, same pattern as the update/skew toasts) and
the toast offers the second way out: Disable, which writes enabled: false
through PUT /api/mcp/servers/{name}/enabled. The gateway's config reconcile
(#109906) and the serve backend's reload both follow that edit.
Toasts gain an optional secondaryAction (outline button beside the primary).
vitest and tsc -p tsconfig.app.json were green because the app config excludes tests; the image build runs tsc -b, which type-checks them and widened the indexed lookup to a union missing the optional slot.
The desktop and the web dashboard each carried a private copy of the
cyberpunk / ember / midnight / mono palettes and they had drifted: the
dashboard's cyberpunk canvas was #040608 with a mint #9bffcf accent
while the desktop's was #000a00 with #00ff41, ember and midnight
disagreed on both canvas and accent, mono agreed only by luck.
Move the raw palette table for every built-in preset into
@hermes/shared (`THEME_PRESET_PALETTES`, apps/shared/src/theme-presets.ts)
and make it the single source of truth:
- apps/desktop/src/themes/presets.ts spreads its `colors` / `darkColors`
from the shared table; the OKLCH synthesis, terminal palettes and
typography stay in the desktop. Serialised BUILTIN_THEMES are
byte-identical to before, so the existing `--dt-primary-solid`
parity pins stay green untouched.
- web/src/themes/presets.ts projects each shared preset onto its
3-slot model through one pure function, `webPresetFromShared`
(background <- background, midground <- primary, warmGlow <- the
midground/ring accent), so cyberpunk / ember / midnight / mono now
render the desktop's palette. Web-only presets (default,
default-large, nous-blue, rose) are untouched.
- Invariant test (web): for every preset shared by both surfaces the
dashboard canvas equals the shared background and the projected text
colour keeps >= 3:1 contrast against it. Red on the previous hexes,
green now.
Why: one edit in one place should recolour a preset on every surface;
two hand-maintained tables guarantee the drift the audit found.
The autosave test did `await import('./config-settings')` inside the test
body. That module graph is large: 1.5s cold on an idle machine, and on a
saturated CI runner (import 1682s aggregate across the run) it alone
crossed the 15s testTimeout twice today on PRs that never touched the file.
The assertions themselves take ~100ms. Hoisting the import into a
beforeAll with its own 60s hookTimeout keeps the test's budget for the
behaviour under test; the fault-injection mocks are hoisted vi.mock calls
and still apply.
The cookie gate (middleware._attempt_refresh) never coalesced concurrent requests
carrying the same stale refresh token, so a browser/desktop burst after access-token
expiry replayed a just-rotated RT into the provider's reuse detection and the whole
session was revoked (#55712). Both refresh paths also ran the synchronous provider
HTTP call on the ASGI event loop, which wedged /api/status behind a slow IdP.
Generalise Doud-FR's native-route single-flight (#71548) into
refresh_singleflight.refresh_session_coalesced and use it from both paths, each in
run_in_threadpool. The replay key is the RT alone: a burst that straddles a network
change must still coalesce, and whoever presents the RT already owns the session.
Middleware keeps its refresh_expired / provider_unreachable audit events via callbacks.
Live E2E (evals/dashboard_auth/refresh_singleflight_live_e2e.py, real uvicorn, stub
rotating IdP with reuse detection): base 1/4 requests survive each burst, 4 provider
calls, /api/status 2.7 s behind one refresh; fixed 4/4, 1 call, 10 ms.
Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
The Sep-2026 facade decomposition moved _detect_image_mime_type_from_bytes and
_normalize_to_supported_image into tools/vision_tools_image_prep.py; import from
the defining module per repo convention (compat pointers are off limits in-tree).
Self-review finding (hostile redteam pass), and it defeats the exact guarantee
the previous commit advertised.
The brand scan is bounded by the declared ftyp box size so a token in a FOLLOWING
box cannot upgrade a HEIC to AVIF. But the bound fell back to the whole 64-byte
sniff window whenever the declared size was outside `16 <= size <= len(header)`.
That failed OPEN in precisely the attacker-controlled case:
declared size 0 -> image/avif (genuine HEIC, 'avif' in a later box)
declared size 4 -> image/avif
declared size 8 -> image/avif
declared size 12 -> image/avif
declared size 15 -> image/avif
declared size 24 -> image/heic (truthful size, correct)
A size too small to hold any compatible brand now scans none, rather than
scanning everything. A size >= 16 that overruns the sniffed window is a truncated
read rather than an attack, so that case clamps to the bytes actually available
and a genuine oversized-ftyp AVIF is still detected.
The prior test only exercised the truthful size, so it proved the happy path
rather than the invariant it was named for. Added parameterized malformed-size
cases (0/1/4/8/12/15), a misaligned-size case (17-23, non-multiples of 4), and an
honest-oversized case. All six malformed-size assertions fail if the fail-open
bound is restored; verified by mutation, not assumed.
Two defects in the HEIF/AVIF sniffing added by this PR.
1. Major-brand-only detection mislabeled AVIF as HEIC.
AVIF encoders routinely stamp the generic still-image brand 'mif1' as the
MAJOR brand and declare the AV1 codec only in the compatible-brand list, so
every mif1-major file was reported as HEVC-coded HEIC. Parse the
compatible-brand list (bounded by the declared ftyp box size, so an 'avif'
token in a following box can't upgrade a real HEIC) and let AV1 brands win
when both families appear.
2. AVIF was rejected outright when pillow-heif was missing.
Pillow >= 11.3 bundles a native AvifImagePlugin, while pillow-heif wheels
are commonly built with no AV1 codec at all (libheif_info() reports
AVIF: ''). Gating AVIF on a pillow-heif import therefore refused files
Pillow could already decode, and pointed the user at a library that cannot
decode them. Registration is now best-effort: the decode attempt is the
arbiter, and failures emit codec-specific guidance naming the backend that
actually serves that format.
Tests: mif1+avif and mif1+av01 regressions, a box-size bound guard, a
truncated/bogus-size guard, AVIF-converts-without-pillow-heif, and AVIF error
text. Verified against real pillow-heif-encoded HEIC and Pillow-encoded AVIF
files, not just synthetic headers; both new brand tests fail if the
compatible-brand scan is reverted.
AGENTS.md:559-576 requires every PyPI dependency to carry an upper bound
rather than an exact runtime pin. Replace "pillow-heif==1.4.0" with
"pillow-heif>=1.4.0,<2" and regenerate uv.lock (resolves 1.5.0, with
hashes), confirming the range admits later 1.x releases.
vision_analyze rejected iPhone photos with 'source is not a recognized
image'. iPhones capture HEIC (ISO-BMFF/HEVC) and upload pipelines often
mislabel it .jpg; the magic-byte sniff had no HEIF branch, so the bytes
were rejected before reaching normalization.
- _detect_image_mime_type_from_bytes: sniff the ISO-BMFF 'ftyp' box and
recognize HEIF/HEIC (heic/heix/heim/heis/hevc/hevx/mif1/msf1) and AVIF
(avif/avis) brands, returning image/heic or image/avif. Non-image ftyp
brands (mp4, ...) stay unrecognized.
- _normalize_to_supported_image: register the pillow-heif Pillow opener,
then re-encode HEIF/AVIF to PNG before embed (same soft-dependency
pattern as SVG rasterization). Missing decoder yields an actionable
'pip install pillow-heif' error, not a generic failure.
- Add pillow-heif==1.4.0 as a core dep alongside Pillow (prebuilt wheels
bundle libheif; no system libs needed).
- Tests: brand detection (incl. AVIF + mp4-not-misdetected), end-to-end
HEIC->PNG normalization, and the decoder-missing error path.
The vision resolver's magic-byte sniff stays authoritative (no extension
trust); HEIF now reaches the existing normalize step instead of 400ing.
Twenty-three tests imported helpers from sibling test modules by dotted
path (tests.run_agent.test_run_agent, tests.cli.test_cli_init,
tests.gateway.test_42039_duplicate_user_message); those follow the moves
and renames. The run_agent autouse fixture that zeroes jittered_backoff
now covers all of tests/agent, where one pre-existing test asserts the
real backoff is positive — it opts out via a real_retry_backoff marker,
the same pattern real_concurrent_gate and real_agent_prewarm use.
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
Follow-ups from the #109962 review that landed after the merge:
- A hand-edited manifest (`"secondaries": null`, a record without profile/home,
`"default": []`) crashed with a traceback AFTER the flag had already been
flipped. It is now refused up front, with the same message from the dry-run
plan and the real run.
- The dry-run plan said "restore its recorded standalone gateway" for a record
with neither service nor pid, where the run did nothing; both now say so.
- `--standalone --dry-run` with no manifest exits 1 like the real path and
prints the same hint.
- One "Rollback incomplete" message; a failed default restart now tells the
user the multiplexer is still running and how to stop it.
- Drop the dead `prefixes and` guard; an empty profile name can no longer
produce a bare ":" platform prefix.
`_finite_positive_config_float` and `_config_int` were the same shape with the warn call
pasted six times and asymmetric sign checks. Collapse both onto `_liveness_knob(key, default,
cast)`: usable iff finite, >= 0 and exact for the cast; else warn and return 0; explicit 0
stays silent.
Two int-path holes closed on the way: `websocket_liveness_failure_threshold: .inf` raised
OverflowError inside `DiscordAdapter.__init__`, and `0.5` truncated to 0 and disabled the
probe silently — the bug class this PR exists to remove. `2.0` / `"2"` still resolve to 2.
The cherry-picked commit added an `event_silence` probe dimension stamped from
`on_socket_raw_receive`. Two verified problems make it a regression rather than a fix:
- discord.py 2.7.1 dispatches `socket_raw_receive` only when the client is built with
`enable_debug_events=True` (client.py:330, gateway.py:410-412; the default
`log_receive` is a no-op). The adapter never sets it, so the stamp only ever moves at
`on_ready` and every healthy connection reads `event_silence` 300s later — a forced
reconnect every ~5 min. Live-verified against a real `commands.Bot` +
`DiscordWebSocket.received_message`: 6 frames delivered, stamp unchanged, probe unhealthy.
- discord.py already keeps a per-frame clock (`KeepAliveHandler._last_recv`) and closes the
socket itself after `heartbeat_timeout` without frames; and because ACKs are frames,
`ack_stale` (60s) always trips before `event_silence` (300s). A raw-frame stamp cannot
detect the "ESTAB + ACKing + zero events" incident by construction.
Kept and tightened the warning half: bool values (`float(True) == 1.0` silently enabled a
knob at 1s), negative ints, and unparsable strings now warn; an explicit `0` is the documented
opt-out and stays silent. Tests trimmed to the two invariant contracts (warn / don't warn),
proven red on origin/main. Docs updated to match.
The Gateway WS health probe sampled only transport state — ready, open,
heartbeat-ACK age, latency. A socket that stays ESTAB and keeps ACKing while
zero gateway frames arrive (the #109521 "connected-but-deaf" incident) read
healthy indefinitely, and the adapter went silent for hours with no log line
and no watchdog firing.
Two defects fixed:
1. Dispatch-side dimension. `on_socket_raw_receive` now stamps
`_last_gateway_frame_at` for every inbound raw gateway frame — heartbeats
and ACKs included, so a legitimately quiet server is not flagged. The
health check gains an `event_silence` reason with its own bound,
`websocket_event_max_silence_seconds` (default 300s, 0 disables the
dimension alone). The stamp resets on `on_ready` so a reconnect never
inherits pre-restart silence. Trip path is unchanged: consecutive
failures -> retryable `discord_websocket_health_stale` -> the existing
reconnect watcher builds a fresh adapter.
2. Silent probe disable. `_finite_positive_config_float` / `_config_int`
mapped anything `float()` rejects ("15s", "nan", "true") to 0.0 with no
log line, permanently disabling the watchdog invisibly. Unparsable and
non-positive values now log one WARNING naming the knob and raw value.
The new key rides the existing `_YAML_WEBSOCKET_LIVENESS_KEYS` seeding and is
documented in the Discord guide's liveness section.
The manifest is kept after a partial rollback so it can be re-run, but the
detached-gateway branch spawned unconditionally, so every re-run doubled the
secondary (service start/restart is idempotent, Popen is not).
Tests: dedupe the fixture-local profile-name helper to module scope.
Rollback started each per-profile gateway while the still-live multiplexer's
gateway_state.json listed it as served, so `hermes -p X gateway run` refused
with exit 78 and RestartPreventExitStatus parked the unit for good — the exact
symptom #109473 reports, one step later. Clear served_profiles (and the
profile's platform entries) first; recorded [] is authoritative for the guard.
A failed secondary previously skipped the default restart with the flag already
off, leaving config saying standalone while the process kept multiplexing. The
default now restarts regardless (it is the last operation either way, so a
cgroup-kill of this process can no longer strand later steps); the manifest is
kept for a re-run only when something failed.
Tests: the fake service layer now runs the real served-by-multiplexer guard at
each secondary start, and a failed-secondary contract test; both red on the
prior stack.
An interruption during the first secondary stop/uninstall left no recovery
metadata on disk. Write the complete manifest first; the dict never changes
afterwards, so the per-secondary and post-flag rewrites were byte-identical
and are dropped.
Based on #109490 by @JoaoMarcos44.
resolve_provider_client("commandcode-anthropic") with no explicit api_mode (a bare
``auxiliary.<task>.provider`` entry) built a plain OpenAI client: _wrap_transport
only consulted req.api_mode and URL heuristics, and api.commandcode.ai/provider/v1
matches none. With _reasoning_config now emitted for that profile, an unwrapped
client would TypeError on the unknown kwarg; before, the OpenAI-wire call simply
misfired against a Messages endpoint. Fall back to the registered profile's
api_mode so the wrap and the reasoning gate agree.
The delegation imported ``plugins.model_providers.deepseek`` — the only cross-plugin
module import in the tree, and one that resolves solely through the loader's
sys.modules shim (popped again if the deepseek plugin fails to load). Look the
profile up with get_provider_profile("deepseek") instead: already imported from
``providers``, honours a user override of the profile, degrades to the base no-op
without a try/except.
Drop the ``len(m) <= len("deepseek/")`` guard — the native profile returns
({}, {}) for an empty id anyway. Bind the expected value in the parity test and
assert it is non-empty so the equality cannot pass as ({}, {}) == ({}, {}).
commandcode-anthropic (api_mode=anthropic_messages, OpenAI-shaped base URL) lost
thinking control on compression/title/vision calls after #109530: once its class
overrides build_api_kwargs_extras, _project_provider_profile marks the profile as
handling reasoning and drops the generic extra_body.reasoning fallback that the
Anthropic adapter used to read — and the _reasoning_config gate only fired for
provider=anthropic, Portal /v1/messages ids, /anthropic URLs and MiniMax.
Carry the profile's api_mode on the projection and include it in that gate, so
any anthropic_messages profile reaches the adapter regardless of URL shape.
Chat-completions profiles are unchanged.
Before: _build_call_kwargs("commandcode-anthropic", claude-haiku, {"enabled": False})
-> no _reasoning_config, no extra_body.reasoning; adapter defaults thinking.
After: -> _reasoning_config={"enabled": False}.
Gate findings on the #108614 salvage (2c + simplify-code):
- CopilotACPProfile.fetch_models promised None on failure but raised (AuthError on a
missing CLI, RuntimeError/TimeoutError from the probe) — only the caller in
hermes_cli/models.py caught it; any other caller got exceptions. Wrap the body and
return None so the docstring and the base ProviderProfile.fetch_models contract hold.
- timeout_seconds was a PER-REQUEST budget (initialize + session/new each got the full
timeout => ~30s worst-case foreground stall on a hung CLI). One shared session deadline.
- /model --refresh (clear_provider_models_cache) wiped the disk cache but not the new
session memo, so a fresh CLI login stayed stale for 5 min past an explicit refresh.
- A failed probe is now memoized for 30s instead of the full 5 min, so signing in to
the CLI is picked up on the next switch.
- _fresh_acp_memo fixture restores the memo on teardown — no cross-test state leak.
- build_api_messages reads agent._current_turn_timestamp directly: a caller that skipped the
turn prologue now raises instead of silently falling back to per-request wall time, which
would re-create the mid-turn drift the fix removes. Only production caller
(assemble_api_request) runs after _reset_per_turn_agent_state; cross-reference to the
tripwire _inflight_turn_started so the two clocks are not "unified" by mistake.
- is_dangerous_confirmation lowercases once instead of once per pattern (now on the per-request path).
- Tests: one _send(idx=) helper instead of three spellings of the builder call; the
untrustworthy-stamp contract is its own test.
Two replay-boundary findings from the #109320 review (gaoanze888, andrexibiza):
- is_interrupted_tool_result matched the bracketed marker anywhere in the payload, so a
successful `read_file`/`search_files` result QUOTING "[Command interrupted]" (a doc, a
grep hit) was classified as a killed run — the read-only block vanished from the
model-facing history, a `terminal` result became the UNKNOWN-effect notice — while
send/replay parity still held because both consumers made the same mistake. Require
the shape every executor actually produces: the marker is the LAST line of `output`,
inside a JSON envelope with a non-zero exit code (or a bare text result). Real
interruptions from tools/environments (130), managed_modal (130) and
code_execution_tool (-1) all keep matching.
- strip_stale_dangerous_confirmations only bounded age from above, so a finite FUTURE
stamp (clock skew, corruption) had negative age and kept the confirmation plus its
api_content sidecar live. Freshness is now 0 <= admission - ts <= expiry; a future
stamp expires like a corrupt one. Missing stamps (legacy rows) stay untouched.
Test: the parity fixture carries both quoted-marker results (JSON grep hit, bare doc text)
and a genuine execute_code interruption envelope; the expiry test covers "nan" and two
future epochs. Red under: marker-anywhere, any-line, old loose heuristic, upper-bound-only.
Follow-up to the two cherry-picked commits from #105308 (@JoaoMarcos44), closing the
three blockers raised on that thread plus one regression the salvage found:
- Prefix-only on the send path. build_api_messages now canonicalizes only
messages[:current_turn_user_idx]; rows the current turn appended (its own tool
calls/results) pass through verbatim. Canonicalizing the live tail rewrote a block
the previous iteration had already sent whenever a tool result matched the
interrupt heuristic, which is exactly the mid-turn prefix rewrite this fix exists
to remove, and it also made the dangling-tail transform order-dependent on when
the user row was appended.
- Exact interrupt marker. is_interrupted_tool_result matched
"exit_code" + ("130" | "-1") + "interrupt" as substrings, so an ordinary
`grep KeyboardInterrupt` result next to a diff hunk header rewrote a terminal
result to an orphan notice (or dropped a read-only block). That heuristic was
tolerable at resume time only; it now runs per request. Match the executors'
bracketed markers ("[Command interrupted", "[execution interrupted") and nothing else.
- Admission-time clock. The frozen expiry clock was the input's platform-event
stamp, so a message queued 70 s before the turn ran kept a 129 s-old confirmation
live on the send path while replay expired it. _reset_per_turn_agent_state stamps
time.time() once at admission; the three other writes (bind identity, stage
message, build_api_messages write-back under suppress(Exception)) are gone.
- Fail closed on corrupt stamps. A present-but-unparseable timestamp (`"nan"`,
`"not_a_number"`) made strip_stale_dangerous_confirmations keep the confirmation
and its api_content sidecar. Coerce through hermes_cli.timefmt.coerce_epoch and
treat an unknowable age as expired; missing stamps (legacy rows) are still left
alone.
- Shape: drop the canonicalize_history_for_send alias (no consumer, never existed on
main), the `now=` kwarg (no production caller), and the getattr/hasattr rewrite of
_reset_per_turn_agent_state (only the test double needed it).
- Tests: 17 → 2 invariant tests. Real SessionDB round trip → canonicalize →
ChatCompletionsTransport bytes, equal to the send path with sidecars applied and
the durable list untouched, live tail preserved; admission-clock freeze across
iterations + corrupt-stamp fail-closed. Each is red under the matching mutation
(send path unpatched, whole-list canonicalization, per-request clock, fail-open,
loose heuristic).
The 1e9 → '1000M' test row read as a snapshot of a missing rung; the
formatter comment now states the cap (token/cost figures stay well under a
billion; a B suffix would collide with the bytes reading) and the row cites it.
scripts/dump_desktop_slash_registry.py --check exits 1 (no write) when the
committed apps/desktop/src/lib/desktop-slash-registry.json differs from
desktop_surface_registry(); the docstring now cites the test that actually
fails (test_desktop_slash_registry.py, not test_commands.py). The pytest
runs --check as a subprocess (green on the committed file) and calls
check() on a hand-edited copy (red).
scripts/ci/classify_changes.py: _PY_SKIP includes apps/, so an apps-only
edit of that JSON (or apps/shared/src/gateway-events.json) skipped the
Python lane and the only equality check never ran in CI. A table of
cross-language contract files now forces python:true; two classifier rows
pin it.
The shared ensureContrast shipped the TUI's fine 0.05×20 ladder, which
changed --dt-primary-solid for 7 of 15 desktop presets (nous #3b6acb →
#3f70d8, cyberpunk #00661a → #008021, slate #505457 → #6f7377) while the PR
body said no preset VALUE changed. The ladder is now the desktop's original
algorithm exactly — pole by luminance < 0.5, accumulating 0.2 steps up to
1.0001, re-mixed from the source colour — with `step` as a parameter. The
only pre-refactor TUI caller (ColorChain.ensureContrast) passes 0.05, so
the terminal palette is byte-identical too.
Test: apps/desktop context.test.tsx iterates every builtin preset × mode,
paints it through ThemeProvider and asserts --dt-primary-solid equals the
value a reference copy of the old desktop algorithm computes. Sabotage
(default step 0.05): 11/30 rows fail. Docs: the SDK table now lists
contrastRatio as `number | null` under sRGB measures, not OKLCH.
The desktop model picker moved from foldIncludes (searchFold: lower-case +
`[-_.]` → space on text AND query) to the shared fuzzyRank, which only
lower-cased. `gpt.4o`, `claude_3` and `qwen3-8` returned zero rows where
main listed gpt-4o / claude-3-opus / qwen3.8-flash, while HighlightMatches
still folded — filter and highlight disagreed. fuzzyScore now folds both
sides with a length-preserving fold, so positions still index the original
target and all three surfaces rank a separator variant identically.
Tests: three separator rows in fuzzy.test.ts and in the desktop picker
test. Sabotage (lower-case only): all six fail.