Commit Graph

11861 Commits

Author SHA1 Message Date
f-trycua 5b010f448f fix(computer-use): verify Windows driver repair 2026-08-16 11:34:40 -07:00
f-trycua c83ea25dd4 fix(computer-use): preserve missing driver overrides 2026-08-16 11:34:40 -07:00
f-trycua 3162745276 fix(computer-use): report unusable driver exit status
(cherry picked from commit 8bda6191ca548d65648263dfe30d2d18950a6e60)
2026-08-16 11:34:40 -07:00
Francesco Bonacci 3e0087abe2 fix(computer-use): warn when an approval bypass widens the driver mode
`--yolo` / `-z` read as "don't prompt me", but they also swap computer_use
onto a private `unrestricted` daemon, dropping the ceilings the configured
mode would have applied. Nothing said so. A script picks up `-z` for quiet
output and loses its limits as a side effect, and the only trace is a driver
process nobody inspects.

The mapping itself stays. It is deliberate, and `unrestricted` is reachable
no other way: it is intentionally not a config value so a stale config line
can never silently bypass approvals (see `_cua_configured_permission_mode`).
Removing the mapping would delete the capability rather than fix it, and
splitting it onto a second CLI flag was declined to avoid growing the
surface.

So the widening is now stated instead: one warning per session naming the
configured mode it left, what stopped applying, and the two ways to keep a
ceiling - drop the bypass flag, or declare a version-3 capability manifest,
which now rides along with unrestricted as of the previous commit.

Once per session, not per dispatch: the resolver runs on every tool call.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 11:34:40 -07:00
Francesco Bonacci cdbea83e1c fix(computer-use): keep a v3 capability manifest on approval-bypassed runs
`--yolo` / `-z` route the session onto a private embedded daemon in
`unrestricted` mode. That daemon was constructed without the configured
capability manifest, and the serve command only attached
`--capability-manifest` when the mode was exactly `bounded`. So the moment a
run was bypassed, the user's declared ceiling was dropped:

    without -z:  --permission-mode bounded --capability-manifest ...
    with -z:     --permission-mode unrestricted --dangerously-bypass-approvals

No manifest, no warning. The most carefully configured run - a reviewed
ceiling, written by hand - became the least constrained one, silently, and
it failed open.

That was never a driver limitation. cua-driver documents the manifest as a
ceiling across modes ("A manifest can narrow a profile but never widen it";
its own authorization table calls it `optional_capability_manifest_ceiling`),
and accepts it alongside `--permission-mode unrestricted`.

The forwarding is version-aware, because the two manifest schemas differ
(cua-driver session_manifest.rs):

* v1/v2 are legacy and must declare `mode: bounded`. Handing one to an
  unrestricted runtime aborts startup with "legacy capability manifest mode
  must be bounded", so a naive forward would turn a working session into a
  hard failure. These are forwarded for bounded only, and a warning names
  the migration when one cannot apply.
* v3 must not declare a mode. It is the mode-independent ceiling, and it now
  rides along with unrestricted.

Unreadable or unparseable manifests are not forwarded outside bounded, on
the same fail-safe reasoning; bounded still forwards unconditionally and
lets the driver be the authority there.

Verified against cua-driver 0.20.0 on Windows. Launch args now carry
`--permission-mode unrestricted --dangerously-bypass-approvals
--capability-manifest <v3> --approve-capability-manifest`, and the ceiling
is enforced in the bypassed run - a tool outside the manifest is refused
("outside the capability manifest for this session ... blocked as a
protected resource") where the same config previously ran unbounded. A
legacy manifest was confirmed to abort driver startup when forwarded, which
is what the version gate prevents.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 11:34:40 -07:00
Francesco Bonacci 5f049b517b fix(computer-use): make the typed-browser bind/snapshot split discoverable
`cua_browser_state` has two branches, chosen implicitly: any call carrying
pid or window_id is a *binding* (browser_route.py:252), anything else is a
*snapshot*. A binding clears session state, mints fresh tab_ids, returns
binding metadata with no page content, and sets verification_required.

Nothing in the response says that. A caller that keeps passing pid/window_id
- the natural reading of "bind to this window, then read it" - re-binds
forever: the tab_id it just received is unbound by the next bind, so every
cua_browser_navigate comes back browser_verification_required, and the
refusal ("take a fresh snapshot") points at the same call that just re-bound.
Observed live as 11 consecutive refused navigates before the model gave up
and fell back to foreground SendInput on the address bar.

The same confusion silently swallowed include_screenshot: both calls that
requested one were bindings, which carry no page content, so the flag had
nothing to attach to and was dropped without comment.

A binding response now reports snapshot_required, next_step
(fresh_browser_state, matching the existing token convention) and a hint
naming the exact next call; requesting a screenshot on a binding reports
screenshot_deferred instead of dropping it. The verification refusal now
says to call cua_browser_state WITHOUT pid/window_id and why re-sending them
does not help. The schema documents that include_screenshot applies to
snapshots.

Behavior of the bind and snapshot branches themselves is unchanged - this is
purely about making the split legible to the caller.

Unit-tested. Not verified end to end on the reporting host: the driver
refuses the bind upstream there (`browser_requires_setup: no owned DevTools
endpoint`, and it does not accept a user-launched --remote-debugging-port),
so the typed route never reaches this branch. That attach failure is a
separate cua-driver issue.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 11:34:40 -07:00
Francesco Bonacci 9a96fdc5b8 fix(computer-use): enforce existing-profile grant, unblock the opt-in
Live-testing the Cua Driver 0.20 convergence on Windows 11 (session 2,
cua-driver 0.20.0) surfaced three defects in the existing-profile browser
path and in install status.

1. The config grant was silently nullified by an approval bypass.

`--yolo` / `-z` map onto a private unrestricted daemon, which answers every
browser_prepare. Because the host delegated the entire existing-profile
decision to the driver, that bypass also nullified
`computer_use.grant_existing_profile: false`: a plain `hermes -z` attached
to the user's real Chrome profile and read live page content over CDP, with
the driver reporting it as "the approved existing Chromium profile". It was
never approved.

An approval bypass is consent to skip prompts, not consent to read an
existing profile's pages, cookies, and storage. CuaTypedBrowserRoute.prepare
now enforces the key itself, regardless of permission mode. bounded stays
exempt - its reviewed capability manifest is the authorization boundary.
The authorization inputs are resolved in the backend from config and the
backend's immutable mode, never from model-supplied kwargs.

2. The grant, once set, still could not be used.

With `grant_existing_profile: true` the runtime is launched
`--grant existing-profile` correctly, but cua_browser_prepare then hit a
runtime approval prompt anyway - re-asking the user to authorize what the
config already authorized, and making the documented opt-in unusable on any
non-interactive run, where the prompt has nobody to answer it and the call
dies on approval timeout. The durable, file-backed grant now stands in for
that prompt. Scope is narrow: only the existing-profile prepare, only when
the grant is present; isolated launches still prompt and any resolution
failure falls closed to prompting.

3. `computer-use status` hid a custom override and spliced its output.

With HERMES_CUA_DRIVER_CMD pointed at cmd.exe, status printed the child's
multi-line banner and prompt inside the one-line version field, never
mentioned the override, and advised `hermes computer-use install` - which
install itself (correctly) refuses to run against an overridden path. It now
names the override and mirrors install's update-or-unset guidance, and
version output is reduced to one bounded line.

Verified on the reported host: `-z` existing-profile attach now refuses and
names the key; `grant: true` no longer prompts (33s vs a 300s approval
timeout); status names the override and prints one line. No change to the
reconciliation path - driver SHA256 unchanged end to end.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 11:34:40 -07:00
Francesco Bonacci 81af2ef013 fix(computer-use): reconcile existing cua-driver installs 2026-08-16 11:34:40 -07:00
Francesco Bonacci a403fe6f92 feat(computer-use): support Cua Driver 0.20 runtime contracts 2026-08-16 11:34:40 -07:00
Teknium c257e9196b fix: make every tool interruptible — sequential executor abandons on user interrupt
The sequential tool path only noticed a user interrupt after the running
tool returned: with the deadline disabled it ran the tool inline (fully
blocking), and with a deadline it waited in 5s slices without ever
checking agent._interrupt_requested. Any tool without cooperative
is_interrupted() polling (image_generate, tts, transcription, skills
sync, ...) held the whole turn hostage — the reported symptom was a
redirect queued ~40s behind a FAL image generation + upscale pass.

Executor backstop (class fix, covers ALL tools):
- _run_sequential_tool_execution_middleware always dispatches on the
  daemon worker (timeout None no longer means inline blocking) and polls
  the interrupt flag every 1s.
- On interrupt: 3s cooperative grace (mirrors the concurrent path), then
  synthesize a cancelled tool result (_ToolCancelledResult), emit the
  terminal post_tool_call with status=cancelled, and abandon the worker.
- _ToolCancelledResult suppresses downstream post-hook double emission
  exactly like _ToolTimeoutResult, so an abandoned worker finishing late
  cannot report success for a cancelled call.
- clarify (interactive, _NEVER_PARALLEL_TOOLS) keeps the inline path —
  it owns its own human wait.

Cooperative layer in the reported offender:
- image_generation_tool: blind handler.get() (generation + Clarity
  upscale) replaced with _wait_fal_result(), which polls is_interrupted()
  in 0.5s slices and raises ImageGenerationInterrupted immediately.
- _upscale_image propagates the interrupt instead of swallowing it into
  the "upscale failed, use original" fallback.

Message alternation is preserved: the cancelled result is a normal tool
result for the call_id. Sabotage-verified: with the old wait loop
restored, the new tests fail (tool blocks full runtime); with the fix
they pass in ~4s.
2026-08-16 11:32:02 -07:00
Teknium f06c41522e fix(image_gen): disable default-on upscaling everywhere — opt-in only
The Aug 8 default-on upscaling policy (66ea4e686) chained the Clarity
Upscaler after every sub-2MP generation. Clarity is an SD1.5 creative
tile-diffusion enhancer (creativity 0.35, "masterpiece" prompt prefix) —
it redraws content, which degraded output on 100% of generations for
models like GPT Image 2 and Ideogram whose value is precise text
rendering, CJK, and photorealistic detail.

Policy now: no model upscales by default, on FAL or Krea. The `upscale`
tool param remains as a per-call opt-in (`upscale: true`); explicit
requests still chain Clarity (FAL) / Krea Enhance as before.

- FAL catalog: all 17 default-on entries flipped to upscale=False
- Krea plugin: medium + medium-turbo per-model defaults flipped off
- Tool schema: upscale param described as opt-in with a fidelity warning
- Tests updated: catalog invariant now pins all-off; default-on cases
  now assert no upscaler call
- Docs (en + zh) updated to the opt-in policy
2026-08-16 11:12:05 -07:00
Teknium 19631b5292 fix(update): tighten gateway-side unit gates to exact/hyphenated shape
Mirror the strict unit-name shape from the hermes-serve gate (review on
PR #83595) on the gateway side too: the discovery gate and the SIGUSR1
eligibility helper now accept only `hermes-gateway.service` or the
`hermes-gateway-<profile>` family, so a near-prefix unit like
`hermes-gatewayd.service` can neither enter the restart path nor be sent
a SIGUSR1 it does not handle.
2026-08-16 10:51:50 -07:00
chelsealong 0b42cae068 fix(update): tighten hermes-serve unit gate, dedupe fleet/cleanup restarts
Review on #83595 flagged two service-lifecycle gaps in the hermes-serve
restart support:

- The unit-name gate accepted anything starting with "hermes-serve",
  which also matched the unrelated hermes-server.service. Require the
  exact base unit or the hyphenated profile family instead.
- The fleet-restart loop and _finish_dashboard_update_cleanup() could
  both restart the same hermes-serve unit — the loop restarts it
  directly, then cleanup's PID scan finds the fresh process and
  restarts its owning unit again. Thread the fleet loop's restarted
  unit names through to _kill_stale_dashboard_processes() so it skips
  units already handled.
2026-08-16 10:51:50 -07:00
chelsealong f97cc250da fix(update): restart hermes-serve systemd units alongside gateways
hermes update discovered and restarted hermes-gateway* systemd units but
never looked for hermes-serve* — the Desktop app's backend — so it kept
running stale pre-update code until the user restarted it by hand (#83438).

Extend the systemd unit discovery/restart loop to also match hermes-serve*
units. They don't wire SIGUSR1 to a graceful drain (only gateway/run.py
does), so restart eligibility for the graceful path is now gated on unit
name via a small, directly-tested helper; hermes-serve units fall straight
to the existing blunt systemctl restart path, matching the workaround the
issue already documents.
2026-08-16 10:51:50 -07:00
Teknium 250232ff91 Revert "fix(agent): harden canonical tool call deduplication"
This reverts commit 8fc4189edd.
2026-08-16 10:51:11 -07:00
Teknium 587ad8748f Revert "fix(agent): preserve local reasoning timeout opt-out"
This reverts commit 26b2b47593.
2026-08-16 10:51:11 -07:00
Teknium 06b9141109 fix(state): classify structural DB corruption as its own persistence cause
'database disk image is malformed' contains the word 'disk', so
classify_persistence_error bucketed SQLITE_CORRUPT / SQLITE_NOTADB
failures as 'disk' and the turn-completion explainer told users to
free disk space for a structurally damaged state.db (the #77386-family
misdiagnosis, reproduced in the v0.20.0 malformed-DB incident report).

- hermes_state: new 'corrupt' bucket in PERSISTENCE_ERROR_CAUSES,
  matched via _DB_CORRUPTION_MARKERS BEFORE the locked/disk buckets
- run_agent: explainer text for 'corrupt' points at hermes doctor and
  explicitly says freeing space will not help
- cron explainer-variant suppression picks the new variant up
  automatically (it iterates PERSISTENCE_ERROR_CAUSES)
2026-08-16 10:34:23 -07:00
Teknium ddc51ec109 fix(update): reload process-scan modules at the dashboard-cleanup entry point
Widen PR #87757 to cover the ZIP path: _update_via_zip() also calls
_finish_dashboard_update_cleanup() but never runs _reload_config_modules,
so the Windows git-broken fallback would still crash with the same
ImportError (cannot import name 'bounded_probe_run' from the stale cached
hermes_cli._subprocess_compat).

- new _reload_process_scan_modules() called inside
  _finish_dashboard_update_cleanup itself, so every current and future
  call site is covered; reloads dependency-first
  (_subprocess_compat, then dashboard_procs)
- reload failures log at warning (a miss surfaces seconds later as an
  ImportError in the same process)
- regression tests: reload-before-kill ordering, node-failure skip,
  stale-module symbol restoration (the exact #87134 boundary state),
  nonfatal reload failure, and the #87757 reload-list contract
2026-08-16 10:31:24 -07:00
Teknium d709d29f19 fix(agent): trim background_review to the enabled switch
Follow-up to #87400: drop the max_iterations and prompt_file knobs from
auxiliary.background_review. The aux model routing (provider/model/
base_url/...) predates #87400 and stays; the enabled switch and the
usage telemetry stay. The fork's iteration budget returns to the
historical hardcoded 16.
2026-08-16 10:27:52 -07:00
Ojas Sharma 7095e23eb2 fix(agent): attribute background-review usage and add cost controls
Persist fork token usage under session_model_usage task=background_review,
emit a per-fork completion log line, and expose enabled/max_iterations/
prompt_file so operators can see and bound the automatic review cost.

Address review feedback: load auxiliary.background_review once per spawn,
classify completion logs by summarize action prefixes, treat explicit
api_call_count=None as the documented default of 1, and WARNING on the
fail-open enabled-gate path.
2026-08-16 06:38:38 -07:00
Teknium 00ecb5d538 test(cli): retarget the wmic-encoding regression test at bounded_probe_run
The Windows-only test asserted encoding/errors kwargs on a mocked
subprocess.run, but the scan now routes through bounded_probe_run
(#87134), so subprocess.run is never invoked. Assert the probe call's
contract instead (errors='ignore', finite timeout), verify the parsed
PIDs, and add a fail-open case for probe failure. The test no longer
needs a Windows host once the probe is mocked, so the windows_only
gate is dropped.
2026-08-16 06:32:08 -07:00
kshitij 4e3de140c1 fix(cli): bound the Windows process-scan probes so a slow WMI scan cannot wedge hermes update (#87134)
subprocess.run(capture_output=True, timeout=N) is not hang-safe on
Windows: after the timeout fires, run()'s cleanup kills the direct child
and then joins the pipe reader threads with an UNBOUNDED communicate().
A descendant (conhost.exe under wmic/powershell) holding duplicated pipe
handles keeps the pipes from EOF and the join never returns.

_scan_gateway_pids() runs its wmic / Get-CimInstance Win32_Process scans
exactly that way, and on machines where the full process scan genuinely
exceeds its 10/15s budget (cold WMI on first boot, ARM VMs, heavy
Update/AV activity) hermes update wedged forever inside
_pause_windows_gateways_for_update() before printing a single line —
observed live on a fresh Windows 11 ARM64 VM with a faulthandler stack
pinning the main thread in subprocess._communicate and only a conhost.exe
child surviving. The single-flight update lock then blocks retries until
the wedged process is killed by hand.

This is the same deadlock class bounded_git_probe already fixed for git
probes (#68609 / #66037). Generalize that proven pattern into a shared
bounded_probe_run() — explicit communicate(timeout), kill_process_tree on
failure, bounded 1s drain, then abandon the daemonic readers — and
migrate the whole call-site class onto it:

- hermes_cli/gateway.py _scan_gateway_pids (the site that hung; reached
  from hermes update, cron, gateway restart/status, dashboard)
- hermes_cli/dashboard_procs.py wmic scan (same shape, reached on update)
- hermes_cli/claw.py tasklist + PowerShell probes (same shape; its
  try/except cannot catch a hang because a hang raises nothing)
- bounded_git_probe now delegates to bounded_probe_run (identical
  contract, one copy of the cleanup logic)

Unlike bounded_git_probe, bounded_probe_run returns the CompletedProcess
(or None) rather than collapsing to stdout, because the gateway scan
branches on returncode to trip its wmic -> powershell fallback.

Tests: tests/hermes_cli/test_bounded_probe_run.py covers success,
nonzero-exit passthrough, spawn failure, bounded timeout (fails against
the old unbounded semantics — verified by sabotage), errors= decoding,
DEVNULL stdin, POSIX process-group placement, and the bounded_git_probe
delegation contract. Existing test_git_probe_tree_kill.py passes
unchanged against the delegated implementation.

Closes #87134
2026-08-16 06:32:08 -07:00
Teknium 5b19c8ee55 fix(acp): make --acp probe tri-state, cached, and mock-safe
Salvage hardening on top of #87308 (thanks @Dudeman456):

- Tri-state verdict: inconclusive probes (binary missing, --help
  failed/timed out) return None and fall through to the normal spawn
  path, preserving the established 'Could not start Copilot ACP
  command' error instead of masking it. This also fixes the two
  test_copilot_acp_client HOME-env regressions that went red on the
  PR: their mocked-Popen path was intercepted by the new unmocked
  subprocess.run probe.
- Cache definitive verdicts per binary path so CLIs that DO support
  --acp pay the ~50ms --help cost once per process, not per prompt.
- Skip the probe entirely when custom ACP args don't include --acp.
- Fix the help-text regex: the old pattern never matched '[--acp]'
  (leading '[' is neither start-of-string nor whitespace) and \b
  after 'p' matched '--acpfoo'.
- Hermeticity: stub subprocess.run in the two HOME-env tests; add 6
  probe-specific tests (fast-fail, fall-through, caching, skip).
2026-08-16 06:31:46 -07:00
yflmq001 0816ba2f93 fix(cli): address review feedback on chat -c fail-loudly PR
Response to AI review (Enough1122) on #86812:

1. `_create_titled_session`: log the underlying exception before returning
   None so programmatic callers aren't left with an undebuggable "could
   not be created" — failures (DB lock, I/O, import) now land in errors.log
   via logger.exception.

2. Drop the source-reading `TestSourceGuard` — it violated the repo's
   "never read source code in tests" rule and was a change-detector
   (passed even if behavior regressed). The stderr routing is already
   covered by the real-path `test_missing_session_fails_on_stderr`.

3. Bare `-c` + `--create-if-missing` now prints a stderr note explaining
   the flag needs a session name, instead of silently ignoring it — makes
   the no-op self-evident to programmatic callers.

Tests: 6 targeted + 8 adjacent, all passing.
2026-08-16 06:31:26 -07:00
yflmq001 a2fcac087c fix(cli): chat -c fails loudly on stderr and gains --create-if-missing
`hermes chat -c "<title>" -q "<text>"` silently no-oped when no session
matched the title under quiet/programmatic use: the not-found message was
written to stdout (the channel quiet callers parse as the final response),
so a background send to a not-yet-existing named session vanished with no
error. Surfaces via Hermes-Bot-Mode bot-to-bot handoffs (#86794).

- not-found message now goes to stderr (exit 1 unchanged), so programmatic
  callers always see it even with -Q/--quiet
- new --create-if-missing: with `-c <title>` and no matching session, create
  a fresh session carrying the title and proceed — the deterministic
  "send to this named thread, making it if needed" primitive plugins asked for
- extract the -c resolution block into _resolve_continue_arg for testability

Tests: flag parsing, titled-session creation, stderr routing, source guard.
2026-08-16 06:31:26 -07:00
Adolanium fccf2b718e fix(gateway): complete /loop ticks after streamed already_sent turns
Streamed replies return None so the adapter does not send twice.
The /loop hook then saw empty text and never ran, so
awaiting_response stayed true and later ticks never fired.

Stash the delivered text on the event and use it for the post-turn
hooks. /goal uses the same path.

Tests: tests/gateway/test_loop_command.py
2026-08-16 06:31:07 -07:00
Teknium c2e94822d8 test(gitlock): pin the git-process guard in sweep tests (deflake slice 8)
The stale-lock removal tests asserted the sweep result while leaving
_git_proc_running() live: on CI the parallel per-file runner almost
always has a real git subprocess in flight, pgrep -x git hits, and
clear_stale_git_locks correctly refuses to sweep — failing the tests
for reasons unrelated to the code under test (surfaced on PR #86918,
which doesn't touch gitlock at all).

Monkeypatch the guard to False in the sweep tests and add an explicit
test pinning the guard's block-while-git-running behavior.
2026-08-16 06:29:53 -07:00
SHT 047c72eead test(auth): cover profile-scoped key resolution in resolve_provider (#86917)
Three regression tests: scoped DEEPSEEK_API_KEY is visible to auto
detection under multiplex (the #86917 failure); unscoped paths keep the
os.environ read; explicit config provider still wins.
2026-08-16 06:29:53 -07:00
SHT 795e035f60 fix(cron): coerce script_path to str in the NUL guard so it can never crash (#86829, review #86832)
"\x00" in script_path raises TypeError when a caller passes a non-str
(e.g. a pathlib.Path, which is not iterable) — the guard itself would
crash the scheduler. All current call sites pass plain str, but the
guard must be crash-proof: str() first, then check. Adds a regression
test running a real script through _run_job_script with a Path argument,
which fails with TypeError on the pre-fix guard.
2026-08-16 06:29:03 -07:00
SHT 5c290caa9c test(cron): pin the eager NUL rejection contract for _run_job_script (#86829) 2026-08-16 06:29:03 -07:00
Eva 089b437b35 fix(tui): stop queued dispatch after session close 2026-08-16 06:26:37 -07:00
Eva ad85feec43 fix(tui): settle session close against active turns 2026-08-16 06:26:37 -07:00
QDung210 b921fdd886 fix(gateway): improve Windows detach diagnostics 2026-08-16 06:26:12 -07:00
Darko Mesaroš a7253a6c02 fix(cli): deliver the Bedrock API key through a named provider
The Bedrock API-key flow stored the bearer token in OPENAI_API_KEY and set
a bare `provider: custom`. Since #28660 that variable is only honoured for
openai.com hosts, so for bedrock-mantle.*.api.aws the token was dropped and
requests went out with api_key="no-key-required", a 401 on every call.

Write a named `providers.bedrock-mantle` entry with
key_env: AWS_BEARER_TOKEN_BEDROCK instead. The named-provider branch in
runtime_provider.py resolves key_env; the bare-custom branch cannot.

Fixes authentication only. Per-model mantle route selection is separate.
2026-08-16 06:25:46 -07:00
Merge_Conflict - Pasi 309cf2c5e2 fix: honor JSON-array string forms for skills.disabled and agent.disabled_toolsets
`hermes config set` and JSON-mode editor saves store lists as quoted
strings (e.g. '["skill-a","skill-b"]' or "['memory']"). Both disable
filters treated such a string as a single name, so curated disable
lists silently filtered nothing with zero diagnostics.

Add parse_config_string_list() in agent.skill_utils and use it in
_normalize_string_set (skills.disabled / platform_disabled) and at
every agent.disabled_toolsets read site: tools_config resolve +
reconcile, CLI, gateway agent construction (both sites), cron
scheduler, and prompt_size. A scalar string still names a single
entry (#13026); malformed JSON falls back to the single-name
behavior instead of raising.

Fixes #86661
2026-08-16 06:24:56 -07:00
3Nya3 bd0586e062 fix(update): surface config mutations applied silently during version-bump-only updates 2026-08-16 06:24:31 -07:00
SHT 029a0d8c76 fix(cron): process .pth files for Windows uv-venv script jobs (#86567)
_windows_cron_python_invocation bypasses the uv venv launcher (to avoid
flashing a console window) and re-attaches the venv via PYTHONPATH — but
PYTHONPATH entries are plain sys.path additions and never get .pth
processing, so editable installs (pip install -e) were invisible to cron
script jobs (ModuleNotFoundError).

Bootstrap the script with site.addsitedir() on the venv site-packages,
then exec it as __main__ via runpy.run_path, preserving the script
directory on sys.path (python script.py semantics). Falls back to a
plain invocation when the venv layout is unresolvable.
2026-08-16 06:24:06 -07:00
kshitij 56526bc0d3 Merge pull request #87630 from kshitijk4poor/fix/kitty-extended-keys-proper
fix(cli): restore Kitty keyboard protocol push and complete the extended-key alias table
2026-08-16 15:57:33 +05:30
kshitij 03bf85d83b fix(cli): restore Kitty keyboard protocol push and complete the extended-key alias table
Commit 4c34eeb416 fixed dead Ctrl+C by removing the Kitty protocol push
(CSI >1u) from _EXTENDED_ENTER_KEYS_SEQ, keeping only modifyOtherKeys
level 2. That regressed kitty-the-terminal completely: kitty removed
xterm modifyOtherKeys support (kovidgoyal/kitty#4075) and only speaks
its own protocol, so after the removal kitty users lost Shift+Enter and
every other extended key — the CSI >4;2m we still pushed is a no-op
there (kitty even logs a PARSE ERROR for it).

The original reason for removing the push is obsolete: #87511 mapped
CSI-u control sequences, so Ctrl+C as ESC[99;5u now parses to
Keys.ControlC and fires the existing c-c binding. (The kernel-INTR
concern in that commit was moot — prompt_toolkit's raw mode clears
ISIG, so Ctrl+C is always handled by the binding, never the kernel.)

Restore the dual push (CSI >1u + CSI >4;2m), exactly mirroring the Ink
TUI, and complete the alias table for what the kitty disambiguate flag
actually emits — #87511 left real gaps, some of which its PR body
wrongly claimed were covered:

- Esc key: ESC[27u (+ modifiers) — previously leaked '[27u' as text
- Ctrl+Backspace -> backward-kill-word (#78285 was closed on the wrong
  claim that codepoint-127 mapping existed; it did not)
- Shift+Space -> space (#86866's second symptom; the Ctrl+Space
  mapping never covered modifier 2)
- Alt+Enter -> newline tuple; Shift+Tab -> BackTab; Ctrl+Tab -> Tab;
  Alt/Shift+Backspace
- Multi-modifier letters (Shift+Alt 4, Ctrl+Shift 6, Ctrl+Alt 7,
  Ctrl+Alt+Shift 8) normalized onto their Ctrl/Escape-prefix targets,
  both unshifted (kitty) and shifted (mok emitters) codepoints
- Kitty PUA functional keys: keypad -> non-keypad equivalents,
  F13-F24, and Ignore for lock/media/modifier-event keys so they are
  consumed instead of leaking (kitty emits these even in legacy mode)

Also: clear the VT100 parser's prefix cache after installing (stale
answers could misparse), and re-push extended keys after
_recover_terminal_input_modes' reset — the recovery previously popped
both modes mid-session and never re-enabled them, silently killing
Shift+Enter until restart.

Refs #87511, #87074, #56684, #56645, #78285, #86866, #87390.
2026-08-16 15:52:24 +05:30
Teknium 6c745e81ae fix(tools-config): stop reconfigure flow clobbering image_gen.use_gateway on managed FAL rows
The Nous Subscription image_gen row carries imagegen_backend="fal", so
_reconfigure_provider's post-model-picker step ran

    img_cfg["use_gateway"] = False

unconditionally at two sites, immediately after the managed branch had
written use_gateway=True. A user who picked Nous Subscription and then
re-entered `hermes tools` to change the model was silently flipped onto
their personal FAL_KEY.

Same bug class as fe63353cb, which fixed the plugin-provider selector
but missed these two legacy-backend sites in the reconfigure flow. Both
now write bool(managed_feature), matching the existing correct site in
_configure_provider.

Adds regression tests driven through the real TOOL_CATEGORIES managed
row; sabotage-verified (tests fail with the old behavior restored).
2026-08-16 02:44:18 -07:00
rainbowgits e14b095f20 fix(tui): map lineage edit ordinals past compression prefix
Desktop/TUI count full displayed lineage after compression, but
prompt.submit validated truncate ordinals against tip-only history.
Translate via display_history_prefix and recover stale 4018s on Desktop.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-16 02:24:40 -07:00
Teknium e22fa90769 feat(desktop): MCP fleet cost/usage overlay with schema token estimates and 30-day usage
Each configured server row on the MCP Capabilities page now shows what it
costs and whether it earns its keep:

- ~per-call token estimate of the server's tool schemas, summed over ENABLED
  tools only (ceil(schema_chars/4) via the existing include/exclude filter)
- 30-day usage count from getUsageAnalytics(30), cached per scope profile
  like the Toolsets tab's toolCallsCache, mapped to servers via the
  mcp__<server>__<tool> registry-name convention (tools/mcp_tool.py)
- a subtle muted "unused" pill on enabled, probed-ok servers with nonzero
  schema cost and zero 30-day uses — never a dialog

Backend: the /api/mcp/servers/{name}/test probe now fills an additive
per-tool `schema_chars` (length of the SAME converted registry schema the
agent registers). Older backends omit it → renderer shows counts only;
older renderers ignore the extra key. Display-only: nothing changes what
schemas are sent to models, no config knobs.

i18n keys (costTokens/usage30d/unusedPill) added to types/en/zh/zh-hant/ja
(ar inherits en via defineLocale overrides). Pure math lives in
lib/mcp-cost.ts with unit tests; Python wire shape pinned in
tests/hermes_cli/test_web_server_profile_unification.py.
2026-08-16 02:24:32 -07:00
yu-xin-c c3d74603aa fix(desktop): avoid PowerShell parent marker boot gate 2026-08-16 02:14:35 -07:00
Teknium 8033389662 fix(desktop): match custom provider aliases in model catalog menu
Fixes #87035
2026-08-16 02:08:53 -07:00
Teknium 25fabcf8eb fix(telegram): keep /loop and synthetic sends in the active DM topic
Fixes #87051
2026-08-16 02:07:35 -07:00
Teknium 1e6717baf2 fix(web_server): discover root user plugins under profile-scoped processes
When the backend is spawned profile-scoped (`--profile <name>` sets
HERMES_HOME=<root>/profiles/<name>), _discover_dashboard_plugins()
scanned only get_process_hermes_home()/plugins — the profile directory,
which has no plugins/ content. Pooled per-profile backends therefore
discovered zero user plugins, mounted no plugin API routes, and every
plugin REST call fell through to the SPA catch-all 404.

Also scan get_default_hermes_root()/plugins (which unwraps
<root>/profiles/<name> to <root> and leaves a custom HERMES_HOME
untouched when it is itself the root), matching how hermes_cli.plugins
resolves install locations. The profile home is scanned first, so a
profile-local plugin of the same name stays authoritative via the
existing seen_names dedupe.

Adds regression tests for root-plugin discovery under a profile-scoped
process and for profile-over-root precedence.

Fixes #87197 (plugin discovery half — the misleading /api/* catch-all
half is addressed separately in #87270).
2026-08-16 02:07:16 -07:00
kshitij 4ce0d64be4 test(caching): pin signal precedence and the openrouter-host opt-out
Review follow-up. Three coverage gaps in the LiteLLM matrix:

- The operator opt-out on a litellm-named provider pointed at an OpenRouter
  host. That route previously took the OpenRouter branch and ignored an
  explicit per-model `prompt_caching: false`; it is the only cell in the
  differential matrix where the salvage REMOVES caching, so pin it as
  intended rather than leaving it to be read as a regression.
- Signal precedence: an explicitly litellm-named provider grants even on a
  lookalike host, because the provider id is an independent signal and only
  the host-derived signal is token-gated. Intentional, now documented.
- A hyphen-delimited host label (`my-litellm-gw.internal.example.com`),
  which the token matcher handles but nothing exercised.

Traded the redundant `claude-3-7-sonnet` parametrize cell for the new host
case, so the matrix covers more shapes with the same cell count.

Tests: 83 passed. All three production fixes re-mutation-checked against
the final stack.
2026-08-16 14:34:30 +05:30
kshitij 435d6f30b5 fix(caching): match the litellm provider id token-wise too
Self-review follow-up. The previous commit fixed substring matching on the
HOST but left the provider-id side as a bare substring, so a user-named
provider like `custom:notlitellm` or `mylitellmthing` still matched and was
handed Anthropic markers — the same bug class, half-fixed.

Both signals now match `litellm` as a whole delimited token via a shared
helper. Real spellings (`litellm`, `custom:litellm`, `litellm-router`, and
the already-lowercased `LiteLLM`) still match; lookalikes no longer do.

Tests: 71 passed. Adds lookalike-provider and real-spelling guards; both
new guards mutation-checked. Differential matrix over 2688 configs vs
origin/main: 60 changes, every one a Claude model on a genuine LiteLLM
route getting the envelope layout, zero pre-existing routes altered.
2026-08-16 14:34:30 +05:30
kshitij 1b0e953b47 fix(caching): use the envelope layout for LiteLLM Claude on the OpenAI wire
Follow-up to the salvaged LiteLLM cache grant. The grant itself is right;
four things about how it was scoped were not.

1. Layout. The branch returned the native inner-block layout
   (use_native_layout=True) on api_mode == "chat_completions". That layout
   writes a TOP-LEVEL msg["cache_control"] on role:tool and empty-content
   messages and depends on the Anthropic adapter to relocate it into the
   block — but that adapter only runs for api_mode == "anthropic_messages"
   (agent/transports/anthropic.py registers there), and the
   chat_completions transport does no relocation. Measured on a 3-tool-turn
   transcript: 2 of the 4 available breakpoints landed on markers the
   provider never sees. Worse, when LiteLLM itself relocates a top-level
   marker for an OpenRouter-backed Claude route
   (OpenrouterConfig._move_cache_control_to_content), the marker lands on
   an empty assistant turn and produces a cache_control-marked empty text
   block — the HTTP 400 "text content blocks must contain" shape already
   guarded in agent/anthropic_adapter.py (#69512). Switched to the envelope
   layout, matching every other OpenAI-wire grant in this function:
   4 of 4 breakpoints honored, zero empty blocks.

2. Host matching. `"litellm" in base_url_hostname(...)` is the substring
   false-positive class base_url_hostname's own docstring warns against; it
   granted Anthropic markers to notlitellm.example.com,
   foolitellmbar.example and friends. Replaced with a label-token match in
   a named helper, so "litellm" must be a whole dot- or hyphen-delimited
   token. All three of the original test hosts still match; a "litellm"
   path segment on an unrelated host still does not.

3. Transport gate. `not is_anthropic_wire` also swept in codex_responses,
   bedrock_converse and codex_app_server. Gated on
   api_mode == "chat_completions" explicitly.

4. Operator override. The grant is inferred from a provider/host name, but
   the custom-provider capability lookup was gated on is_anthropic_wire, so
   an explicit `prompt_caching: false` for the route+model was honored on
   /v1/messages and silently ignored on /v1/chat/completions. The lookup
   now also runs for a LiteLLM route, and its layout follows the transport
   rather than the declaration (an explicit `true` must not promote a
   chat_completions request to the native layout).

Tests: 64 passed. Adds the wire-shape contract the original matrix was
missing (asserts no breakpoint sits on the message envelope, rather than
only checking the returned tuple), plus lookalike-host, other-transport,
and both operator-override directions. All five guards mutation-checked —
reverting each fix turns the corresponding test red.
2026-08-16 14:34:30 +05:30
Justin Bowes ff4df5e54d fix(caching): engage prompt caching for LiteLLM Claude on the OpenAI wire
anthropic_prompt_cache_policy() only granted Anthropic cache_control
markers to LiteLLM over the native Anthropic wire
(api_mode == "anthropic_messages"). A LiteLLM deployment exposing the
OpenAI-compatible surface instead (/v1/chat/completions, /v1/messages
-> 404) matched no grant branch and fell through to (False, False): no
cache_control injected, the system prompt sent as a plain string, and
the provider serving zero cache hits -- the entire prompt re-billed at
full price on every turn. Silent: no error, no warning, usage simply
shows 100% uncached input forever.

Add one branch after the is_anthropic_wire/is_claude case that grants
caching to Claude-family models on a LiteLLM endpoint regardless of
wire, with the native inner-block layout. Same failure class already
documented in-function for Qwen/DashScope.

Design:
- Gated on the Claude family only (is_claude); a Gemini/GPT/Qwen route
  through the same proxy must not receive markers (they may reject the
  cache_control block format -- cf. the DeepSeek/OpenCode exclusion).
- Matches on provider string OR base_url host, since provider naming
  varies per install (litellm, custom:litellm, or a bare custom alias
  pointed at a LiteLLM host).
- prompt_caching.cache_ttl: false still wins (the _cache_disabled early
  return is untouched).
- Generic strict OpenAI-wire custom providers (e.g. Fireworks) remain
  excluded -- verified by the existing over-reach regression test.

Tests: adds TestLiteLLMOpenAIWire covering the grant (several model
spellings x provider/host signals), no-over-reach (non-Claude on the
same proxy get nothing; operator disable wins), and adjacent behavior
(LiteLLM in Anthropic proxy mode still native layout). Full module:
43 passed.

Closes #84506. Original diagnosis, patch design, and measurements by
@ottosulin.
2026-08-16 14:34:30 +05:30