Commit Graph

34331 Commits

Author SHA1 Message Date
teknium1 06bfcae47f fix(tui-gateway): launch-profile turns keep their env-only terminal policy once multiplexing is active
606903badc made launch-profile turns bind a file-backed terminal scope as
soon as any secondary home is served, so a poisoned ambient bridge can never
be the launch turn's authority. That scope is rebuilt from defaults +
<home>/.env + config.yaml only, which drops the launch process's legitimate
env-only policy: TERMINAL_ENV=ssh TERMINAL_SSH_HOST=example.test with a `{}`
config.yaml became backend=local, ssh_host='' the moment a second profile
was served.

tui_gateway/launch_terminal_policy.py freezes the process TERMINAL_* once,
in _profile_home right before the first secondary home is registered as
served — the last moment ambient env is provably the launch profile's own.
build_profile_terminal_scope takes that snapshot as a trusted env_overlay
sitting where the process env sits in the standalone bridge (explicit YAML
keys still win). Launch turns overlay the snapshot; ambient os.environ is
never re-read after activation, so a later secondary write is still
rejected, and the scope is reset after the turn as before.

Refs #108440 review (andrexibiza), #107442 (ehz0ah).
2026-09-13 14:31:21 -07:00
teknium1 484b8e96e8 fix(tui-gateway): secondary side workers and agent builds bind their own terminal scope
Under multiplexing tools/terminal_tool.py no longer bridges a profile's
terminal.* into os.environ while a home override is bound (#108440), so any
secondary-profile entrypoint that binds only the home (+ secrets) leaves
terminal_tool on the launch process's ambient TERMINAL_*: a secondary with
`terminal.backend: docker` ran its prompt.background / prompt.btw /
preview.restart side agents, its eager resume and its session.branch build
on the launch `local` backend.

_spawn_side_agent and _profile_build_scope now enter
_session_profile_runtime_scope, the same home -> secrets -> terminal
composition a prompt turn binds (and gateway/run.py::_profile_runtime_scope
mirrors). The terminal scope is installed for the whole worker/build
lifetime and reset on success and on exception; a malformed secondary
config still yields the fail-closed refusal scope. No ambient os.environ
write is restored.

Refs #108440 review (andrexibiza), #107442 (ehz0ah).
2026-09-13 14:31:21 -07:00
teknium1 3f86ed75da chore: map juliancruzet@users.noreply.github.com to @JulianCruzet for contributor attribution 2026-09-13 12:52:43 -07:00
teknium1 130284fdcd fix(providers): declare catalog-only aliases on their plugin profiles
Same class as minimax-portal: aliyun, deep-seek, nim/build-nvidia/nemotron
and vertexai resolve in hermes_cli's alias tables but not in
providers.get_provider_profile(), so a profile lookup keyed on the alias
returned None and lost the profile's wire mode/extra_body/headers. qwen is
left out: the catalog maps it to alibaba while the qwen-oauth plugin already
claims it, a pre-existing disagreement outside this fix.
2026-09-13 12:52:43 -07:00
teknium1 117b01afa7 test: collapse minimax alias tests to one registry invariant
The two salvaged tests overlapped (resolution subsumes membership); one
contract test that every documented alias resolves to minimax-oauth is
the salvage bar.
2026-09-13 12:52:43 -07:00
Julian Cruzet d55c3ad70b fix(minimax): register minimax-portal and minimax-global aliases for minimax-oauth (#107928) 2026-09-13 12:52:43 -07:00
teknium1 643b3f450d fix(tools): stop check_fns swallowing resolver crashes into "returned False"
check_vision_requirements (and five siblings: browser_vision, image/video
generation, x_search, browser_vault) wrapped their whole probe in
`except Exception: return False`. The registry then logged "returned False",
indistinguishable from an unconfigured backend, and the only diagnostic for a
crashed resolver was gone (#87950: named custom provider lookup failing in a
long-lived multi-profile process, reported as vision tools silently vanishing).

The registry owns the verdict: _run_check_fn_uncached and _check_fn_cached both
catch, log with traceback, and return False. Let the exception reach them.
Behaviour for the model is unchanged (tool hidden either way); agent.log now
says why.
2026-09-13 12:35:36 -07:00
teknium1 aa7980d777 fix(tools): log the traceback when a check_fn raises on the cached path
The TTL-cached availability path collapsed a raising check_fn into the bare
verdict "check_fn X raised; dependent tools will be unavailable this turn" with
no exception attached, while the uncached path already logged exc_info. A probe
that crashes is a bug in the probe or its resolver, and without the traceback
the log reads exactly like "nothing configured" — that gap turned #87950 from a
ten-minute diagnosis into a multi-day one. Carry the exception out of the
try block and attach it to the verdict.
2026-09-13 12:35:36 -07:00
teknium1 b9271bcb34 chore: map steveahlstrom@gmail.com to @sjahlstrom for contributor attribution 2026-09-13 11:22:08 -07:00
teknium1 093c58ca77 test: cover adapter-wrapped probe stubs and repeat-probe stability
Trim the salvaged test to the ≤2-invariant bar: one parametrized test asserting
repeat probes stay resolvable (the user-visible contract from #87654) and that
neither a bare stub nor a Codex-adapter-wrapped stub lands under the runtime
cache key. The guard now keys on aux_probe_mode being active rather than on the
client's type, which is what covers the wrapped case.
2026-09-13 11:22:08 -07:00
Steve Ahlstrom a714ff9823 fix(aux): stop probe stubs poisoning the client cache
`_store_cached_client()` refuses an `_AuxProbeClientStub`, but
`_get_cached_client()` assigns to `_client_cache` directly and so never
reaches that guard. check_fns resolve through this path inside
`aux_probe_mode()` during tool-schema assembly, and the cache key carries no
probe/runtime distinction — so the stored stub is returned to the next real
caller sharing that key, which dies on attribute access with
`_AuxProbeClientStub used as a real client (attribute 'chat')`.

The `async_mode` field in the cache key is what kept this latent: the probe
caches the sync variant, so async consumers (`analyze_image`) miss the entry
and build a real client, while sync consumers (`browser_vision`) hit the
poisoned one and fail on every call.

Observed against a local OpenAI-compatible vision endpoint, where
`browser_vision` failed every call with that RuntimeError while
`analyze_image` against the same provider worked.

Guard the inline store the same way `_store_cached_client()` does, and
return the stub to the probe caller without caching it.

The existing `test_probe_stub_never_cached` pins the invariant only on
`_store_cached_client()`, which is why the unguarded door went unnoticed;
the new test exercises `_get_cached_client()` and fails without this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 11:22:08 -07:00
hermes-seaeye[bot] 61304d5923 fmt(js): npm run fix on merge (#110132)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-13 17:58:46 +00:00
teknium1 68eb059863 feat(desktop): daily re-auth nudge for MCP servers with a Disable action
An MCP server whose OAuth refresh token died used to get one warning toast
per app session (on the transition into needs-auth) and then nothing: the
server sat parked, silently, until the user happened to open Capabilities →
MCP. A dead token is a standing problem that needs an action, so the health
checker now re-nudges once a day while the server stays broken (persisted
24h snooze per profile+server, same pattern as the update/skew toasts) and
the toast offers the second way out: Disable, which writes enabled: false
through PUT /api/mcp/servers/{name}/enabled. The gateway's config reconcile
(#109906) and the serve backend's reload both follow that edit.

Toasts gain an optional secondaryAction (outline button beside the primary).
2026-09-13 10:53:29 -07:00
teknium1 d8668e25ef test(web): type the shared preset lookup as ThemePresetPalette so tsc -b (the Docker build) accepts darkColors
vitest and tsc -p tsconfig.app.json were green because the app config excludes tests; the image build runs tsc -b, which type-checks them and widened the indexed lookup to a union missing the optional slot.
2026-09-13 10:52:11 -07:00
teknium1 e7657792df refactor(themes): web dashboard presets derive from the desktop palette table
The desktop and the web dashboard each carried a private copy of the
cyberpunk / ember / midnight / mono palettes and they had drifted: the
dashboard's cyberpunk canvas was #040608 with a mint #9bffcf accent
while the desktop's was #000a00 with #00ff41, ember and midnight
disagreed on both canvas and accent, mono agreed only by luck.

Move the raw palette table for every built-in preset into
@hermes/shared (`THEME_PRESET_PALETTES`, apps/shared/src/theme-presets.ts)
and make it the single source of truth:

- apps/desktop/src/themes/presets.ts spreads its `colors` / `darkColors`
  from the shared table; the OKLCH synthesis, terminal palettes and
  typography stay in the desktop. Serialised BUILTIN_THEMES are
  byte-identical to before, so the existing `--dt-primary-solid`
  parity pins stay green untouched.
- web/src/themes/presets.ts projects each shared preset onto its
  3-slot model through one pure function, `webPresetFromShared`
  (background <- background, midground <- primary, warmGlow <- the
  midground/ring accent), so cyberpunk / ember / midnight / mono now
  render the desktop's palette. Web-only presets (default,
  default-large, nous-blue, rose) are untouched.
- Invariant test (web): for every preset shared by both surfaces the
  dashboard canvas equals the shared background and the projected text
  colour keeps >= 3:1 contrast against it. Red on the previous hexes,
  green now.

Why: one edit in one place should recolour a preset on every surface;
two hand-maintained tables guarantee the drift the audit found.
2026-09-13 10:52:11 -07:00
teknium1 7d2bb463c4 test(desktop): load ConfigSettings once in beforeAll so the autosave test's 15s budget is not spent on the import
The autosave test did `await import('./config-settings')` inside the test
body. That module graph is large: 1.5s cold on an idle machine, and on a
saturated CI runner (import 1682s aggregate across the run) it alone
crossed the 15s testTimeout twice today on PRs that never touched the file.
The assertions themselves take ~100ms. Hoisting the import into a
beforeAll with its own 60s hookTimeout keeps the test's budget for the
behaviour under test; the fault-injection mocks are hoisted vi.mock calls
and still apply.
2026-09-13 10:51:50 -07:00
teknium1 5dea46d13d fix(dashboard-auth): one refresh single-flight for the cookie gate and the native route, off the event loop
The cookie gate (middleware._attempt_refresh) never coalesced concurrent requests
carrying the same stale refresh token, so a browser/desktop burst after access-token
expiry replayed a just-rotated RT into the provider's reuse detection and the whole
session was revoked (#55712). Both refresh paths also ran the synchronous provider
HTTP call on the ASGI event loop, which wedged /api/status behind a slow IdP.

Generalise Doud-FR's native-route single-flight (#71548) into
refresh_singleflight.refresh_session_coalesced and use it from both paths, each in
run_in_threadpool. The replay key is the RT alone: a burst that straddles a network
change must still coalesce, and whoever presents the RT already owns the session.
Middleware keeps its refresh_expired / provider_unreachable audit events via callbacks.

Live E2E (evals/dashboard_auth/refresh_singleflight_live_e2e.py, real uvicorn, stub
rotating IdP with reuse detection): base 1/4 requests survive each burst, 4 provider
calls, /api/status 2.7 s behind one refresh; fixed 4/4, 1 call, 10 ms.

Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
2026-09-13 09:40:44 -07:00
Doud-FR f561155a70 fix(gateway): coalesce native refresh per concrete provider across hints
(cherry picked from commit 18ce843daba52c427ab5d4b5f21990efc1202486)
2026-09-13 09:40:44 -07:00
teknium1 6c27a289c3 chore: map contributor email for samfoy 2026-09-13 09:18:54 -07:00
Teknium e016e87d20 chore(deps): regenerate uv.lock for pillow-heif on current main 2026-09-13 09:18:54 -07:00
Teknium adaa6643b9 test(vision): import HEIC helpers from the defining sibling module
The Sep-2026 facade decomposition moved _detect_image_mime_type_from_bytes and
_normalize_to_supported_image into tools/vision_tools_image_prep.py; import from
the defining module per repo convention (compat pointers are off limits in-tree).
2026-09-13 09:18:54 -07:00
Sam Painter c83c413214 fix(vision): fail closed on a malformed ftyp box size
Self-review finding (hostile redteam pass), and it defeats the exact guarantee
the previous commit advertised.

The brand scan is bounded by the declared ftyp box size so a token in a FOLLOWING
box cannot upgrade a HEIC to AVIF. But the bound fell back to the whole 64-byte
sniff window whenever the declared size was outside `16 <= size <= len(header)`.
That failed OPEN in precisely the attacker-controlled case:

    declared size  0 -> image/avif   (genuine HEIC, 'avif' in a later box)
    declared size  4 -> image/avif
    declared size  8 -> image/avif
    declared size 12 -> image/avif
    declared size 15 -> image/avif
    declared size 24 -> image/heic   (truthful size, correct)

A size too small to hold any compatible brand now scans none, rather than
scanning everything. A size >= 16 that overruns the sniffed window is a truncated
read rather than an attack, so that case clamps to the bytes actually available
and a genuine oversized-ftyp AVIF is still detected.

The prior test only exercised the truthful size, so it proved the happy path
rather than the invariant it was named for. Added parameterized malformed-size
cases (0/1/4/8/12/15), a misaligned-size case (17-23, non-multiples of 4), and an
honest-oversized case. All six malformed-size assertions fail if the fail-open
bound is restored; verified by mutation, not assumed.
2026-09-13 09:18:54 -07:00
Sam Painter f5a7dc14c7 fix(vision): parse the ftyp compatible-brand list; don't gate AVIF on pillow-heif
Two defects in the HEIF/AVIF sniffing added by this PR.

1. Major-brand-only detection mislabeled AVIF as HEIC.
   AVIF encoders routinely stamp the generic still-image brand 'mif1' as the
   MAJOR brand and declare the AV1 codec only in the compatible-brand list, so
   every mif1-major file was reported as HEVC-coded HEIC. Parse the
   compatible-brand list (bounded by the declared ftyp box size, so an 'avif'
   token in a following box can't upgrade a real HEIC) and let AV1 brands win
   when both families appear.

2. AVIF was rejected outright when pillow-heif was missing.
   Pillow >= 11.3 bundles a native AvifImagePlugin, while pillow-heif wheels
   are commonly built with no AV1 codec at all (libheif_info() reports
   AVIF: ''). Gating AVIF on a pillow-heif import therefore refused files
   Pillow could already decode, and pointed the user at a library that cannot
   decode them. Registration is now best-effort: the decode attempt is the
   arbiter, and failures emit codec-specific guidance naming the backend that
   actually serves that format.

Tests: mif1+avif and mif1+av01 regressions, a box-size bound guard, a
truncated/bogus-size guard, AVIF-converts-without-pillow-heif, and AVIF error
text. Verified against real pillow-heif-encoded HEIC and Pillow-encoded AVIF
files, not just synthetic headers; both new brand tests fail if the
compatible-brand scan is reverted.
2026-09-13 09:18:54 -07:00
Sam Painter 352d8b0929 chore(deps): use bounded pillow-heif constraint and regenerate uv.lock
AGENTS.md:559-576 requires every PyPI dependency to carry an upper bound
rather than an exact runtime pin. Replace "pillow-heif==1.4.0" with
"pillow-heif>=1.4.0,<2" and regenerate uv.lock (resolves 1.5.0, with
hashes), confirming the range admits later 1.x releases.
2026-09-13 09:18:54 -07:00
Samuel Painter 0a545f7fa5 feat(vision): decode HEIF/HEIC/AVIF images (iPhone photos)
vision_analyze rejected iPhone photos with 'source is not a recognized
image'. iPhones capture HEIC (ISO-BMFF/HEVC) and upload pipelines often
mislabel it .jpg; the magic-byte sniff had no HEIF branch, so the bytes
were rejected before reaching normalization.

- _detect_image_mime_type_from_bytes: sniff the ISO-BMFF 'ftyp' box and
  recognize HEIF/HEIC (heic/heix/heim/heis/hevc/hevx/mif1/msf1) and AVIF
  (avif/avis) brands, returning image/heic or image/avif. Non-image ftyp
  brands (mp4, ...) stay unrecognized.
- _normalize_to_supported_image: register the pillow-heif Pillow opener,
  then re-encode HEIF/AVIF to PNG before embed (same soft-dependency
  pattern as SVG rasterization). Missing decoder yields an actionable
  'pip install pillow-heif' error, not a generic failure.
- Add pillow-heif==1.4.0 as a core dep alongside Pillow (prebuilt wheels
  bundle libheif; no system libs needed).
- Tests: brand detection (incl. AVIF + mp4-not-misdetected), end-to-end
  HEIC->PNG normalization, and the decoder-missing error path.

The vision resolver's magic-byte sniff stays authoritative (no extension
trust); HEIF now reaches the existing normalize step instead of 400ing.
2026-09-13 09:18:54 -07:00
teknium1 42e7471c64 test: route two honcho tests that landed at tests/ root during the rebase 2026-09-13 09:18:02 -07:00
teknium1 76e88cae2a test: repoint cross-test imports at the moved modules; opt one backoff test out of the merged agent fixture
Twenty-three tests imported helpers from sibling test modules by dotted
path (tests.run_agent.test_run_agent, tests.cli.test_cli_init,
tests.gateway.test_42039_duplicate_user_message); those follow the moves
and renames. The run_agent autouse fixture that zeroes jittered_backoff
now covers all of tests/agent, where one pre-existing test asserts the
real backoff is positive — it opts out via a real_retry_backoff marker,
the same pattern real_concurrent_gate and real_agent_prewarm use.
2026-09-13 09:18:02 -07:00
teknium1 d10bb2ab6f test: make tests/ mirror the source tree; drop issue numbers from filenames
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.

Parallel directories for one source package, folded into the mirror:
  tests/acp        -> tests/acp_adapter   (its __init__/conftest move with it)
  tests/cli        -> tests/hermes_cli    (prompt_toolkit fixture merged into
                                           hermes_cli/conftest.py)
  tests/run_agent  -> tests/agent         (backoff fixture becomes
                                           agent/conftest.py)
  tests/relay      -> tests/gateway/relay
  tests/state      -> tests/hermes_state

246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.

Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.

Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).

Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
2026-09-13 09:18:02 -07:00
kshitijk4poor 422bc9bde9 fix(migrate): rollback refuses a malformed manifest before mutating; plan and run agree
Follow-ups from the #109962 review that landed after the merge:

- A hand-edited manifest (`"secondaries": null`, a record without profile/home,
  `"default": []`) crashed with a traceback AFTER the flag had already been
  flipped. It is now refused up front, with the same message from the dry-run
  plan and the real run.
- The dry-run plan said "restore its recorded standalone gateway" for a record
  with neither service nor pid, where the run did nothing; both now say so.
- `--standalone --dry-run` with no manifest exits 1 like the real path and
  prints the same hint.
- One "Rollback incomplete" message; a failed default restart now tells the
  user the multiplexer is still running and how to stop it.
- Drop the dead `prefixes and` guard; an empty profile name can no longer
  produce a bare ":" platform prefix.
2026-09-13 20:06:00 +05:30
kshitijk4poor 5dcd4844bd refactor(discord): one liveness-knob resolver; int knobs reject inf and fractions
`_finite_positive_config_float` and `_config_int` were the same shape with the warn call
pasted six times and asymmetric sign checks. Collapse both onto `_liveness_knob(key, default,
cast)`: usable iff finite, >= 0 and exact for the cast; else warn and return 0; explicit 0
stays silent.

Two int-path holes closed on the way: `websocket_liveness_failure_threshold: .inf` raised
OverflowError inside `DiscordAdapter.__init__`, and `0.5` truncated to 0 and disabled the
probe silently — the bug class this PR exists to remove. `2.0` / `"2"` still resolve to 2.
2026-09-13 20:03:40 +05:30
kshitijk4poor 8dd0e80f29 fix(discord): drop the unreachable frame-silence dimension, keep the config warning
The cherry-picked commit added an `event_silence` probe dimension stamped from
`on_socket_raw_receive`. Two verified problems make it a regression rather than a fix:

- discord.py 2.7.1 dispatches `socket_raw_receive` only when the client is built with
  `enable_debug_events=True` (client.py:330, gateway.py:410-412; the default
  `log_receive` is a no-op). The adapter never sets it, so the stamp only ever moves at
  `on_ready` and every healthy connection reads `event_silence` 300s later — a forced
  reconnect every ~5 min. Live-verified against a real `commands.Bot` +
  `DiscordWebSocket.received_message`: 6 frames delivered, stamp unchanged, probe unhealthy.
- discord.py already keeps a per-frame clock (`KeepAliveHandler._last_recv`) and closes the
  socket itself after `heartbeat_timeout` without frames; and because ACKs are frames,
  `ack_stale` (60s) always trips before `event_silence` (300s). A raw-frame stamp cannot
  detect the "ESTAB + ACKing + zero events" incident by construction.

Kept and tightened the warning half: bool values (`float(True) == 1.0` silently enabled a
knob at 1s), negative ints, and unparsable strings now warn; an explicit `0` is the documented
opt-out and stays silent. Tests trimmed to the two invariant contracts (warn / don't warn),
proven red on origin/main. Docs updated to match.
2026-09-13 20:03:40 +05:30
salch-cred cef499fbb5 fix(discord): liveness probe gains a dispatch-side dimension (#109521)
The Gateway WS health probe sampled only transport state — ready, open,
heartbeat-ACK age, latency. A socket that stays ESTAB and keeps ACKing while
zero gateway frames arrive (the #109521 "connected-but-deaf" incident) read
healthy indefinitely, and the adapter went silent for hours with no log line
and no watchdog firing.

Two defects fixed:

1. Dispatch-side dimension. `on_socket_raw_receive` now stamps
   `_last_gateway_frame_at` for every inbound raw gateway frame — heartbeats
   and ACKs included, so a legitimately quiet server is not flagged. The
   health check gains an `event_silence` reason with its own bound,
   `websocket_event_max_silence_seconds` (default 300s, 0 disables the
   dimension alone). The stamp resets on `on_ready` so a reconnect never
   inherits pre-restart silence. Trip path is unchanged: consecutive
   failures -> retryable `discord_websocket_health_stale` -> the existing
   reconnect watcher builds a fresh adapter.

2. Silent probe disable. `_finite_positive_config_float` / `_config_int`
   mapped anything `float()` rejects ("15s", "nan", "true") to 0.0 with no
   log line, permanently disabling the watchdog invisibly. Unparsable and
   non-positive values now log one WARNING naming the knob and raw value.

The new key rides the existing `_YAML_WEBSOCKET_LIVENESS_KEYS` seeding and is
documented in the Discord guide's liveness section.
2026-09-13 20:03:40 +05:30
kshitijk4poor fadcff9227 fix(migrate): rollback re-run skips a detached secondary that is already running
The manifest is kept after a partial rollback so it can be re-run, but the
detached-gateway branch spawned unconditionally, so every re-run doubled the
secondary (service start/restart is idempotent, Popen is not).

Tests: dedupe the fixture-local profile-name helper to module scope.
2026-09-13 19:49:17 +05:30
kshitijk4poor d8cfb149b1 fix(migrate): clear the multiplexer's served record before starting secondaries; always restart the default last
Rollback started each per-profile gateway while the still-live multiplexer's
gateway_state.json listed it as served, so `hermes -p X gateway run` refused
with exit 78 and RestartPreventExitStatus parked the unit for good — the exact
symptom #109473 reports, one step later. Clear served_profiles (and the
profile's platform entries) first; recorded [] is authoritative for the guard.

A failed secondary previously skipped the default restart with the flag already
off, leaving config saying standalone while the process kept multiplexing. The
default now restarts regardless (it is the last operation either way, so a
cgroup-kill of this process can no longer strand later steps); the manifest is
kept for a re-run only when something failed.

Tests: the fake service layer now runs the real served-by-multiplexer guard at
each secondary start, and a failed-secondary contract test; both red on the
prior stack.
2026-09-13 19:49:17 +05:30
joaomarcos f9e47aa6fe fix(migrate): write the rollback manifest before the first destructive operation
An interruption during the first secondary stop/uninstall left no recovery
metadata on disk. Write the complete manifest first; the dict never changes
afterwards, so the per-secondary and post-flag rewrites were byte-identical
and are dropped.

Based on #109490 by @JoaoMarcos44.
2026-09-13 19:49:17 +05:30
KoNit-K 043286ae68 fix(gateway): make standalone rollback recoverable 2026-09-13 19:49:17 +05:30
kshitijk4poor aa40c1d21b fix(auxiliary): wrap the Messages adapter on the profile's declared api_mode
resolve_provider_client("commandcode-anthropic") with no explicit api_mode (a bare
``auxiliary.<task>.provider`` entry) built a plain OpenAI client: _wrap_transport
only consulted req.api_mode and URL heuristics, and api.commandcode.ai/provider/v1
matches none. With _reasoning_config now emitted for that profile, an unwrapped
client would TypeError on the unknown kwarg; before, the OpenAI-wire call simply
misfired against a Messages endpoint. Fall back to the registered profile's
api_mode so the wrap and the reasoning gate agree.
2026-09-13 19:46:52 +05:30
kshitijk4poor ef754d75ba refactor(commandcode): resolve the native DeepSeek profile through the registry
The delegation imported ``plugins.model_providers.deepseek`` — the only cross-plugin
module import in the tree, and one that resolves solely through the loader's
sys.modules shim (popped again if the deepseek plugin fails to load). Look the
profile up with get_provider_profile("deepseek") instead: already imported from
``providers``, honours a user override of the profile, degrades to the base no-op
without a try/except.

Drop the ``len(m) <= len("deepseek/")`` guard — the native profile returns
({}, {}) for an empty id anyway. Bind the expected value in the parity test and
assert it is non-empty so the equality cannot pass as ({}, {}) == ({}, {}).
2026-09-13 19:46:52 +05:30
kshitijk4poor 5dd8fa1daa fix(auxiliary): anthropic_messages profiles keep /reasoning reachable on aux calls
commandcode-anthropic (api_mode=anthropic_messages, OpenAI-shaped base URL) lost
thinking control on compression/title/vision calls after #109530: once its class
overrides build_api_kwargs_extras, _project_provider_profile marks the profile as
handling reasoning and drops the generic extra_body.reasoning fallback that the
Anthropic adapter used to read — and the _reasoning_config gate only fired for
provider=anthropic, Portal /v1/messages ids, /anthropic URLs and MiniMax.

Carry the profile's api_mode on the projection and include it in that gate, so
any anthropic_messages profile reaches the adapter regardless of URL shape.
Chat-completions profiles are unchanged.

Before: _build_call_kwargs("commandcode-anthropic", claude-haiku, {"enabled": False})
        -> no _reasoning_config, no extra_body.reasoning; adapter defaults thinking.
After:  -> _reasoning_config={"enabled": False}.
2026-09-13 19:46:52 +05:30
kshitijk4poor db66b21ae4 fix(copilot-acp): honor the fetch_models None-on-failure contract; bound the probe budget; refresh-proof the memo
Gate findings on the #108614 salvage (2c + simplify-code):
- CopilotACPProfile.fetch_models promised None on failure but raised (AuthError on a
  missing CLI, RuntimeError/TimeoutError from the probe) — only the caller in
  hermes_cli/models.py caught it; any other caller got exceptions. Wrap the body and
  return None so the docstring and the base ProviderProfile.fetch_models contract hold.
- timeout_seconds was a PER-REQUEST budget (initialize + session/new each got the full
  timeout => ~30s worst-case foreground stall on a hung CLI). One shared session deadline.
- /model --refresh (clear_provider_models_cache) wiped the disk cache but not the new
  session memo, so a fresh CLI login stayed stale for 5 min past an explicit refresh.
- A failed probe is now memoized for 30s instead of the full 5 min, so signing in to
  the CLI is picked up on the next switch.
- _fresh_acp_memo fixture restores the memo on teardown — no cross-test state leak.
2026-09-13 19:30:48 +05:30
kshitijk4poor 5c4c31cf4d refactor(agent): fail loudly on a missing turn clock; single lowercase in is_dangerous_confirmation
- build_api_messages reads agent._current_turn_timestamp directly: a caller that skipped the
  turn prologue now raises instead of silently falling back to per-request wall time, which
  would re-create the mid-turn drift the fix removes. Only production caller
  (assemble_api_request) runs after _reset_per_turn_agent_state; cross-reference to the
  tripwire _inflight_turn_started so the two clocks are not "unified" by mistake.
- is_dangerous_confirmation lowercases once instead of once per pattern (now on the per-request path).
- Tests: one _send(idx=) helper instead of three spellings of the builder call; the
  untrustworthy-stamp contract is its own test.
2026-09-13 19:23:09 +05:30
kshitijk4poor f296652a66 fix(agent): interrupt detection needs the executor's result shape; future-stamped confirmations expire
Two replay-boundary findings from the #109320 review (gaoanze888, andrexibiza):

- is_interrupted_tool_result matched the bracketed marker anywhere in the payload, so a
  successful `read_file`/`search_files` result QUOTING "[Command interrupted]" (a doc, a
  grep hit) was classified as a killed run — the read-only block vanished from the
  model-facing history, a `terminal` result became the UNKNOWN-effect notice — while
  send/replay parity still held because both consumers made the same mistake. Require
  the shape every executor actually produces: the marker is the LAST line of `output`,
  inside a JSON envelope with a non-zero exit code (or a bare text result). Real
  interruptions from tools/environments (130), managed_modal (130) and
  code_execution_tool (-1) all keep matching.
- strip_stale_dangerous_confirmations only bounded age from above, so a finite FUTURE
  stamp (clock skew, corruption) had negative age and kept the confirmation plus its
  api_content sidecar live. Freshness is now 0 <= admission - ts <= expiry; a future
  stamp expires like a corrupt one. Missing stamps (legacy rows) stay untouched.

Test: the parity fixture carries both quoted-marker results (JSON grep hit, bare doc text)
and a genuine execute_code interruption envelope; the expiry test covers "nan" and two
future epochs. Red under: marker-anywhere, any-line, old loose heuristic, upper-bound-only.
2026-09-13 19:23:09 +05:30
kshitijk4poor 820d3ca65d fix(agent): send-path canonicalization is prefix-only, clock is admission time, corrupt stamps fail closed
Follow-up to the two cherry-picked commits from #105308 (@JoaoMarcos44), closing the
three blockers raised on that thread plus one regression the salvage found:

- Prefix-only on the send path. build_api_messages now canonicalizes only
  messages[:current_turn_user_idx]; rows the current turn appended (its own tool
  calls/results) pass through verbatim. Canonicalizing the live tail rewrote a block
  the previous iteration had already sent whenever a tool result matched the
  interrupt heuristic, which is exactly the mid-turn prefix rewrite this fix exists
  to remove, and it also made the dangling-tail transform order-dependent on when
  the user row was appended.
- Exact interrupt marker. is_interrupted_tool_result matched
  "exit_code" + ("130" | "-1") + "interrupt" as substrings, so an ordinary
  `grep KeyboardInterrupt` result next to a diff hunk header rewrote a terminal
  result to an orphan notice (or dropped a read-only block). That heuristic was
  tolerable at resume time only; it now runs per request. Match the executors'
  bracketed markers ("[Command interrupted", "[execution interrupted") and nothing else.
- Admission-time clock. The frozen expiry clock was the input's platform-event
  stamp, so a message queued 70 s before the turn ran kept a 129 s-old confirmation
  live on the send path while replay expired it. _reset_per_turn_agent_state stamps
  time.time() once at admission; the three other writes (bind identity, stage
  message, build_api_messages write-back under suppress(Exception)) are gone.
- Fail closed on corrupt stamps. A present-but-unparseable timestamp (`"nan"`,
  `"not_a_number"`) made strip_stale_dangerous_confirmations keep the confirmation
  and its api_content sidecar. Coerce through hermes_cli.timefmt.coerce_epoch and
  treat an unknowable age as expired; missing stamps (legacy rows) are still left
  alone.
- Shape: drop the canonicalize_history_for_send alias (no consumer, never existed on
  main), the `now=` kwarg (no production caller), and the getattr/hasattr rewrite of
  _reset_per_turn_agent_state (only the test double needed it).
- Tests: 17 → 2 invariant tests. Real SessionDB round trip → canonicalize →
  ChatCompletionsTransport bytes, equal to the send path with sidecars applied and
  the durable list untouched, live tail preserved; admission-clock freeze across
  iterations + corrupt-stamp fail-closed. Each is red under the matching mutation
  (send path unpatched, whole-list canonicalization, per-request clock, fail-open,
  loose heuristic).
2026-09-13 19:23:09 +05:30
joaomarcos 401fef6e6f fix(agent): freeze turn confirmation expiry and verify wire parity 2026-09-13 19:23:09 +05:30
joaomarcos e5ca5207de fix(agent): unify replay history canonicalization 2026-09-13 19:23:09 +05:30
teknium1 988d471479 style(ts): sort imports/exports the way perfectionist wants after the rebase 2026-09-13 06:50:57 -07:00
teknium1 c3edad29ba docs(shared): declare M as compactNumber's top rung
The 1e9 → '1000M' test row read as a snapshot of a missing rung; the
formatter comment now states the cap (token/cost figures stay well under a
billion; a B suffix would collide with the bytes reading) and the row cites it.
2026-09-13 06:50:57 -07:00
teknium1 e06efbcff0 fix(ci): desktop slash-registry dump gets --check; apps/ contract JSON edits run the Python lane
scripts/dump_desktop_slash_registry.py --check exits 1 (no write) when the
committed apps/desktop/src/lib/desktop-slash-registry.json differs from
desktop_surface_registry(); the docstring now cites the test that actually
fails (test_desktop_slash_registry.py, not test_commands.py). The pytest
runs --check as a subprocess (green on the committed file) and calls
check() on a hand-edited copy (red).

scripts/ci/classify_changes.py: _PY_SKIP includes apps/, so an apps-only
edit of that JSON (or apps/shared/src/gateway-events.json) skipped the
Python lane and the only equality check never ran in CI. A table of
cross-language contract files now forces python:true; two classifier rows
pin it.
2026-09-13 06:50:57 -07:00
teknium1 c764d6d354 fix(shared): ensureContrast keeps the desktop's 0.2-step ladder; TUI chain opts into 0.05
The shared ensureContrast shipped the TUI's fine 0.05×20 ladder, which
changed --dt-primary-solid for 7 of 15 desktop presets (nous #3b6acb →
#3f70d8, cyberpunk #00661a → #008021, slate #505457 → #6f7377) while the PR
body said no preset VALUE changed. The ladder is now the desktop's original
algorithm exactly — pole by luminance < 0.5, accumulating 0.2 steps up to
1.0001, re-mixed from the source colour — with `step` as a parameter. The
only pre-refactor TUI caller (ColorChain.ensureContrast) passes 0.05, so
the terminal palette is byte-identical too.

Test: apps/desktop context.test.tsx iterates every builtin preset × mode,
paints it through ThemeProvider and asserts --dt-primary-solid equals the
value a reference copy of the old desktop algorithm computes. Sabotage
(default step 0.05): 11/30 rows fail. Docs: the SDK table now lists
contrastRatio as `number | null` under sRGB measures, not OKLCH.
2026-09-13 06:50:57 -07:00
teknium1 3f02259518 fix(shared): fuzzyRank folds [-_.] to space on both sides like the desktop picker did
The desktop model picker moved from foldIncludes (searchFold: lower-case +
`[-_.]` → space on text AND query) to the shared fuzzyRank, which only
lower-cased. `gpt.4o`, `claude_3` and `qwen3-8` returned zero rows where
main listed gpt-4o / claude-3-opus / qwen3.8-flash, while HighlightMatches
still folded — filter and highlight disagreed. fuzzyScore now folds both
sides with a length-preserving fold, so positions still index the original
target and all three surfaces rank a separator variant identically.

Tests: three separator rows in fuzzy.test.ts and in the desktop picker
test. Sabotage (lower-case only): all six fail.
2026-09-13 06:50:57 -07:00