Commit Graph

12893 Commits

Author SHA1 Message Date
Ayush Nangia 60b93eb521 test(deadline): Phase 3a poisoned-state coverage
- async timeout marks once with a label-carrying rounded-timeout reason
- completion never marks
- sync flavor marks on timeout
- non-adopting backends keep the real timeout result
- a raising mark_suspect cannot eat the timeout or the label
2026-08-26 17:47:22 +05:30
Dimar Anez 9eb13d07b6 fix(terminal): tolerate macOS TCC PermissionError in _safe_getcwd
On macOS with TCC (Transparency, Consent, and Control), os.getcwd()
raises PermissionError: [Errno 1] Operation not permitted — not
FileNotFoundError — when the process CWD is under a protected location
(~/Documents, ~/Desktop, ~/Downloads) and the calling process lacks
Full Disk Access.

_safe_getcwd() only caught FileNotFoundError (deleted CWD), so the
terminal-tool cleanup thread, which calls _get_env_config() →
_safe_getcwd() every 60 s, logged a full stack trace on every tick.
This accumulated hundreds of MB of noise in mcp-stderr.log (observed
184 MB on a single-day session) without breaking functionality — the
cleanup thread's outer try/except swallowed the exception, but
exc_info=True kept emitting the traceback.

Fix: add PermissionError to the existing except clause so the fallback
chain (TERMINAL_CWD → $HOME) runs, matching the existing pattern for
deleted-CWD recovery (#17558). Complements #66306, which handles
PermissionError from subprocess.Popen(cwd=...) for an inaccessible
configured cwd on Linux; this handles the distinct case where the
live process CWD itself is TCC-blocked.

Tests cover: PermissionError fallback to $HOME, TERMINAL_CWD priority,
FileNotFoundError regression, happy path unchanged, and unrelated
OSError (NotADirectoryError) still propagating instead of being
swallowed.
2026-08-26 04:51:41 -07:00
Teknium bd134d0f30 test: loosen frozen bare-verdict dict in cua_0_9 sibling test to decision contract
The verify_fresh_state verdict now carries an optional human hint; assert
the decision + additive-field absence instead of the exact dict shape
(same contract loosening as test_computer_use_delivery_ladder.py).
2026-08-26 04:50:13 -07:00
Teknium 3da5897c39 refactor(computer_use): diet schema + delete prompt block (~1.4K tok/call); remove max_elements, ladder moves to response verdicts 2026-08-26 04:50:13 -07:00
Teknium 600d5166f0 test(gateway): prove delivered rows are never reclaimed by the reconnect sweep 2026-08-26 04:49:39 -07:00
milnerrad 8e1db41041 fix(gateway): redeliver transient failures after reconnect 2026-08-26 04:49:39 -07:00
Justin Johnson c19849cd02 fix(desktop): stop the status-stack poll storming a dead session with 4001s
The composer status stack polls `process.list` every 5s while a background
process row is on screen. `process.list` is session-scoped, so against a
runtime id the gateway no longer holds it returns 4001 "session not found".

`refreshBackgroundProcesses` swallowed *every* failure with a bare `catch {}`
commented "transient socket loss". A gone session is not transient: the poll
re-sent the same dead runtime id every 5 seconds for the lifetime of the
window. On one machine this produced 31,518 gateway rejections in a day
(vs 663 the day before), 18,614 of them against a single runtime id, and it
is what users see reported as "sessions stopped with a session not found
error" after an update.

The trigger is a reconnect, not the poll itself: anything that mints a fresh
runtime (gateway restart, the #94219 reconnect/replay work, an idle-reaped
pooled backend) strands the id the status stack is still holding, and nothing
in this path ever re-checked it.

Distinguish the two failure classes:

- 4001 / "session not found" is TERMINAL for that runtime id — latch the id
  and stop polling it.
- A timeout or transport error is transient — keep retrying, since the
  session may well still be alive. Misclassifying that direction would
  silently freeze the status stack on a healthy session.

The latch is cleared when the status stack (re)binds a session id, so a
session that comes back under a fresh runtime resumes polling normally
rather than staying dark for the life of the app.

Also name the method in the gateway's 4001 warning. That line was added in
c305839442 "for diagnosability", but without the RPC name it cannot say WHICH
client call is looping — the reason this storm could not be attributed from
the logs alone. A ContextVar set in `handle_request` carries it; it is
diagnostic only and never used for authorization.

Tests:
- composer-status: 4001 stops the poll, a timeout does not, one gone session
  never suppresses a healthy sibling, and a rebind resumes polling.
- tui_gateway: the rejection warning names the method.
2026-08-26 04:49:22 -07:00
codexbt 3dea11d703 fix(web_server): recheck WEB_DIST existence dynamically in mount_spa
When hermes dashboard --skip-build runs across agent updates, mount_spa checked WEB_DIST.exists() once at server startup and mounted an immutable 404 handler if the build was missing. As a result, subsequent builds while the server was running continued to serve 404 "Frontend not built".

- Removes early static return in mount_spa().
- Moves WEB_DIST.exists() check dynamically into _serve_index() and serve_spa().
- Mounts /assets StaticFiles with check_dir=False.
- Adds unit test test_mount_spa_dynamic_web_dist_recheck in tests/hermes_cli/test_web_server.py.

Closes #82614
2026-08-26 04:48:58 -07:00
Shakti Prasad Mohapatra 98e87ac886 fix(cli): preserve stale positive behind-count on fetch failure (#92578)
huklaa's review: a failed fetch makes origin/main stale, so a stale ref
cannot prove *currentness* (rev-list 0 is inconclusive), but a stale
positive count is still sound evidence an update exists. On fetch
failure, compute the stale behind-count and return it when > 0;
otherwise return None (inconclusive) and still skip the cache write.

Regression tests:
- fetch failure + stale rev-list 0 -> None (not 'up to date')
- fetch failure + stale rev-list 5 -> 5 (update evidence preserved)
- fetch failure + rev-list error -> None
2026-08-26 04:17:39 -07:00
Shakti Prasad Mohapatra 55d50d5c92 fix(cli): don't serve stale update-check results after fetch failure (#82166)
When _check_via_local_git's git fetch fails (timeout, offline, DNS),
the code silently fell through to compare HEAD against the stale
origin/main tracking ref, which can report 0 (up to date) even when
upstream has moved forward. Combined with the 6-hour cache in
check_for_updates, a single fetch failure could suppress update
notifications for days — the exact symptom in #82166 where the daily
cron reported 'up to date' for 4 days after v0.20.0 was released.

Two fixes:

1. _check_via_local_git now detects fetch failure (returncode != 0 or
   exception) and returns None instead of falling through to stale
   refs. The caller treats None as 'check could not run' rather than
   'up to date'.

2. check_for_updates no longer caches None results. Previously, a
   None from a failed check was cached for 6 hours, suppressing
   retries until the cache expired. Now only conclusive results
   (0 or >=1) are cached, so the next check attempt runs immediately
   on the next call.

Added regression tests:
- test_check_via_local_git_fetch_failure_returns_none
- test_check_for_updates_does_not_cache_none
2026-08-26 04:17:39 -07:00
kshitijk4poor d62a05e94c fix(checkpoints): surface skipped_oversize to users and stop misreporting failed deletes as restored
Follow-up to the salvaged #95207 fix, completing the misreport bug class:

- restore() now also drops delete_targets whose unlink failed (OSError
  swallowed) from restored_files — the sibling of the kept-oversize
  misreport the salvaged fix closed.
- /rollback output in the CLI (cli_commands_mixin) and gateway
  (slash_commands + gateway.rollback.kept_oversize locale key in all 17
  catalogs) now tells the user which files were kept because the size
  cap excluded them from every checkpoint; previously the file was
  correctly preserved but the user got no notice it was not reverted.
- Regression test for the failed-unlink misreport.
2026-08-26 16:44:43 +05:30
RickyYii 595b5ce68a refactor(checkpoints): call the size-cap predicate instead of restating it
Review feedback on #95207: `_exceeds_size_cap` and `_drop_oversize_from_index`
each computed the byte cap and compared against it. Both used `> cap`, so they
agreed, but only by coincidence of two independent expressions — nothing held
them together.

The coupling is the whole point of the fix. The checkpoint decides what to
store and safe restore decides what may be deleted; a threshold that drifted
between them would produce a file both absent from the checkpoint and not
recognised as capped at restore, which is exactly the deletion this branch
exists to prevent. `_drop_oversize_from_index` now calls the predicate.

Added a boundary case to TestSafeRestore that pins the round trip from both
ends: a file at exactly the cap is stored, so it must revert; one byte more is
excluded, so it must be kept. Mutation-checked — moving either side to `>=`
fails it, including the re-inlined-with-`>=` shape the reviewer described.

No behaviour change: the byte cap, the strict comparison and the
unstattable-path result are all as before.

Regression: the 10 test files covering checkpoint_manager / rollback, against
current main (1fe0f2f3a, 134 commits newer than the base measured on the first
commit) — 130 passed on main, 135 here (+5 new), zero failures either side.
2026-08-26 16:44:43 +05:30
RickyYii d28bf79927 fix(checkpoints): stop safe restore deleting files the size cap excluded
`/rollback <N>` runs `restore(..., safe=True)` — safe mode is the default,
`--all` opts out. Safe mode splits the changed files into two groups: those
present in the checkpoint are checked out, and those absent from it are treated
as files Hermes created during the turn and deleted, since deleting them is
what restores the pre-turn state.

Absence from the checkpoint is not proof of authorship. `max_file_size_mb`
(default 10) keeps large files out of every checkpoint via
`_drop_oversize_from_index`, so a file the agent appended to — a dataset, a
corpus, an export, a log — is absent for a completely different reason. Safe
mode deleted it. No checkpoint held a copy, so nothing could bring it back, and
`restored_files` listed the path, so the user was told it had been restored.

Reproduced on main with shipped defaults:

    corpus.jsonl (2 MB), agent appends to it, then /rollback 1
    safe_restore_plan restore=['corpus.jsonl', 'notes.py']
    restore ok=True restored_files=['corpus.jsonl', 'notes.py']
    notes.py     exists=True   content="v1 = 'original source'"
    corpus.jsonl exists=False  <- deleted, was in no checkpoint

Scope: this needs an agent write to the capped file. A large file Hermes never
touched is not in the ledger, lands in `skipped`, and was already safe.

The delete branch now asks whether the path is one the cap would have excluded,
using the same test `_drop_oversize_from_index` applies when building the
checkpoint, so "kept out of the checkpoint" and "refused deletion at restore"
share one definition. Such a path is reported under a new `skipped_oversize`
key and dropped from `restored_files`.

The classification keys on "absent from the checkpoint", not on "large now".
A file small enough to be checkpointed and later bloated past the cap does have
a stored version, and reverting to it is exactly what was asked for — it still
restores, and a test pins that.

The ledger records a content hash, not whether a write created or modified the
file, so an oversize path cannot be proven agent-created. Leaving one behind
costs a stale file the user can delete; removing it costs the file.

Tests: 4 cases in tests/tools/test_checkpoint_manager.py::TestSafeRestore. Two
fail on main — the deletion and the misreport. Two are guards: the
grew-past-the-cap revert, and the small agent-created file that must still be
removed.

Regression: the 10 test files covering checkpoint_manager / rollback —
130 passed on main, 134 with this change (+4 new), zero failures either side.
2026-08-26 16:44:43 +05:30
Teknium 9f8cdf89d6 fix(macos): keep the TCC anchor alive across CVE-repair rotations + sign anchor copies
Integration fixups so #82529's generation signing and #95131's interpreter
anchor cover each other's gaps (without these, each fix leaves the other's
rotation path broken):
- macos_tcc_anchor store detection now recognizes .hermes-runtime/python/
  generation-* stores: repair_vulnerable_runtime() rebuilds the venv against
  a generation interpreter, replacing the anchored bin/python with a fresh
  symlink — previously the anchor then read 'not uv-managed' and NEVER
  re-anchored, so every SQLite CVE repair silently orphaned terminal TCC
  grants (the exact #82427 scenario, path-keyed).
- _install_anchor signs the anchor copy with the same identifier-pinned DR
  (via managed_uv._macos_sign_managed_python) before it goes live: copy2
  carries the source build's cdhash-based signature, so an unsigned refresh
  would still change the stored csreq on every patch bump/repair despite
  the stable path. Best-effort, never blocks the anchor.
- Tests: generation-store recognition + repair-generation anchoring +
  sign-on-install call (sabotage-verified: dropping the generation root
  marker fails both new tests).
2026-08-26 04:14:16 -07:00
notkisk 8d6c0a3098 fix: preserve macOS TCC identity for managed Python 2026-08-26 04:14:16 -07:00
Teknium 19fde8a450 fix(dashboard): compare and spawn the venv interpreter by UNRESOLVED path
Follow-up on the cherry-picked #90030: candidate.resolve() breaks the fix
on the standard Linux venv layout, where venv/bin/python is a symlink to
the base interpreter. Resolving makes the venv python compare equal to
the dependency-less base (so the swap never happens), and returning the
resolved target would spawn the bare base interpreter, bypassing
pyvenv.cfg — the fix would silently not fix #90026 on the exact platform
it was reported from. Compare and return normalized UNRESOLVED paths:
the venv path IS the interpreter's identity. Adds the symlink-layout
regression test; live-E2E'd with a real dependency-less base runtime.
2026-08-26 04:03:21 -07:00
liuhao1024 ad8f995bcf fix(dashboard): spawn detached actions from the install's venv interpreter
Under an SSH remote backend the web server is launched by running the uv
BASE interpreter with the venv's site-packages injected into sys.path at
startup, so sys.executable is a dependency-less python. Detached
dashboard actions spawned from it (Update now, restart, anything routed
through _spawn_hermes_action) inherited neither the injected path nor a
PYTHONPATH and died on the first third-party import — 'Update now'
always failed instantly with ModuleNotFoundError: No module named 'yaml'
while 'hermes update' from the venv worked (#90026).

_dashboard_spawn_executable now prefers the install's own venv
interpreter (venv/bin/python, venv/Scripts/python.exe) when it differs
from sys.executable, resolving the same dependency set the venv launcher
provides. Same-interpreter launches return sys.executable unchanged,
preserving the Windows console-ownership behavior verbatim, and layouts
without an install venv keep the old fallback.
2026-08-26 04:03:21 -07:00
kshitijk4poor 365cbc242b test: assert omitted attach_to_session stays absent from formatted list output
Closes the gap flagged in review: the raw store was checked but not the
_format_job surface.
2026-08-26 16:06:41 +05:30
StanleyStetson 5e9adc9e4d fix(cron): forward attach_to_session through cronjob handler
The public schema and job store already support per-job
attach_to_session, but the registry adapter dropped the argument.
Create silently omitted the field; update reported "No updates provided."

Fixes #84802
2026-08-26 16:06:41 +05:30
kshitijk4poor 2f425872ef test: extend exact-shape assertion in relay-delivery guard for provenance tag
tests/cron/test_cron_relay_delivery_guards.py landed on main after #89329
branched; its exact-dict assertion needs the new _resolved_from field the
salvaged commit adds to origin-resolved targets.
2026-08-26 16:06:09 +05:30
Victor Kyriazakos 580daa7b96 fix(cron): mirror continuable-cron briefs for origin-fallback and opted-in explicit targets
A managed cron (created by a provisioning script, not from a live gateway
chat) never captures an origin. With cron.mirror_delivery: true and
deliver: origin, its brief was delivered to the home channel — the
user's own DM — but the transcript mirror and the in_channel session
seed were silently skipped: _target_matches_origin returns False for an
empty origin, and the whole continuable machinery keys off that check.
A user replying to the brief landed in a session with no record of it.
Field report 2026-08-17 (enterprise, Slack DM surface).

The June origin-scoping refactor (c06ceb3232) was written to exclude
broadcasts, and the exclusion is kept. What changes is the
classification: a home-channel FALLBACK for deliver=origin is the user's
primary conversation standing in for the origin, not a broadcast.

Changes:
- Delivery targets carry a resolution-provenance tag (_resolved_from:
  origin / origin_fallback / explicit; broadcast expansions untagged).
- _target_mirror_eligible replaces the bare origin check at the mirror
  gate: origin unchanged; origin_fallback eligible under the same flags
  as origin (per-job attach_to_session wins, else global
  cron.mirror_delivery); explicit platform:chat targets eligible ONLY
  under per-job attach_to_session — the global flag never activates
  them, so it cannot start writing transcript entries into arbitrary
  explicitly-addressed chats. 'all'/bare-platform stay never-eligible.
- Dedup OR-merges provenance so 'origin,all' resolving to the same chat
  keeps eligibility regardless of token order.
- _inchannel_seed_allowed guards the flat-session seed: group-channel
  session keys are user-isolated, so a seed without a user_id (origin-
  less job into a shared channel) would create an orphan session no
  reply resolves to — those targets fall back to the plain mirror. DM
  targets (keys don't embed user_id) always seed.
- cronjob tool schema text updated to describe the new attach scope.

Behavioral note: origin-less deliver=origin jobs under global
mirror_delivery now activate the full continuable path — on default
'thread' surface this opens a dedicated thread in the home channel
where the brief previously posted flat. That is the documented
continuable behavior; the silent flat post was the bug.

15 new tests (tests/cron/test_mirror_origin_fallback.py): eligibility
matrix (origin/fallback/explicit/all/bare/other-chat), dedup order
both ways, end-to-end mirror via _deliver_result for all four shapes,
origin regression control, seed user_id guard.
2026-08-26 16:06:09 +05:30
Teknium ad7b7255ab fix(state): renaming a bot's canonical Bot Chat is refused — the title IS the identity (#92473)
Bot Mode resolves the forever-chat by exact-title lookup on
(profile, 'Bot Chat'); no session-id pointer exists. A user rename
therefore orphaned the whole conversation: resolution missed, the next
click minted an empty replacement, and UNIQUE(title) then blocked ever
renaming back. Refuse the rename at SessionDB._set_session_title — the
single write path every surface funnels through (gateway session.title,
/title, CLI rename, REST). Hidden discriminates the registry row, so a
normal visible session a user happens to call 'Bot Chat' stays freely
renameable; re-asserting the same canonical title stays a no-op.
2026-08-26 03:21:42 -07:00
Teknium 45db70a80a fix(computer-use): fail closed on unverified CuaDriver.app + background launch
Hardening on top of the TCC daemon-identity salvage:
- _validate_cua_driver_app_signature: codesign -dv gate requiring EXACT
  Identifier=com.trycua.driver and the official team (4YEC26S9KF) before
  /usr/bin/open hands the bundle to LaunchServices — the identity fix must
  not double as a launcher for arbitrary/impostor bundles (suffixed
  identifiers and wrong teams rejected; unsigned dev builds only via
  computer_use.allow_unsigned_driver: true in config.yaml).
- _resolve_cua_driver_app_path: derive the bundle ONLY from the resolved
  driver binary — the /Applications fallback could launch a DIFFERENT
  install than the manifest resolved.
- open -n -g: don't activate/steal focus when launching the daemon.
- 7 new tests incl. sabotage-verified exact-match assertions.

Grafted from #76433's review direction (@Chadmc9889's original fail-closed
validation requirement).

Co-authored-by: Chadmc9889 <Chadmc9889@users.noreply.github.com>
2026-08-26 03:21:37 -07:00
projetsjsl 4746f614be fix(computer-use): preserve macOS TCC daemon identity
Launch private computer-use daemons through CuaDriver.app so Screen
Recording authorization remains attached to its stable bundle identity
instead of Hermes' ad-hoc signature.

Co-Authored-By: GPT-5.6 Codex <noreply@openai.com>
2026-08-26 03:21:37 -07:00
Teknium 1fe0f2f3ac feat(cron): import-error cron failures now name gateway code skew and the one-command fix (#95294 part 3)
When an agent cron job dies with an import-class error (cannot import
name / ModuleNotFoundError / ImportError), the failure summarizer — which
runs inside the gateway process — now consults gateway.code_skew: if the
process booted on a different revision than disk HEAD, the delivered
message appends 'gateway is running stale code (booted on X, disk is at
Y) — run hermes gateway restart'. Turns the reported two-day mystery
(15 missed jobs, identical ImportError, no explanation) into a one-line
fix instruction on the first failure.

Fail-safe by construction: skew detection returns None on non-git
installs and processes without a boot fingerprint, the probe seam
swallows every exception, and no_agent script jobs (fresh subprocess,
consistent imports) fall through to the generic cleaner — their
ImportErrors are the script's own problem, and blaming gateway skew
there would send the reader to the wrong place (same mode-gating as the
provider branches).

Reuses gateway/code_skew.py (the /model-switch skew detector) rather
than adding a second fingerprint reader.
2026-08-26 01:23:15 -07:00
Teknium f0c0c986c4 test: pin overlay policy off in embedded-daemon socket/ack contract test
The embedded spawn now consults the overlay policy (capability probe via
subprocess.run) when _cua_no_overlay() is true — which it is on headless
CI since the Linux X11 default flip. The fixed two-entry run side_effect
in this test didn't budget for the probe call; pin the policy off since
this test pins the socket/ack contract, not overlay behavior.
2026-08-26 00:54:36 -07:00
cvillarroel2 1a7f83a73b fix(computer_use): disable embedded daemon overlay 2026-08-26 00:54:36 -07:00
kshitijk4poor 4ba2608524 fix(compressor): widen empty-content abort to sibling no-response shapes + snapshot state field
Follow-up to PR #94531 salvage:
- classify the auxiliary boundary's terminal 'None response' /
  'invalid response' errors (#7264) into the same empty-content abort
  carve-out so those shapes also preserve the session (#94459's wider
  classification, sibling shapes from #94448)
- register _last_summary_empty_content_failure in
  _COMPRESSOR_ATTEMPT_STATE_FIELDS so pre-commit hard-cancel rollback
  restores the flag (conversation_compression snapshot allow-list)
- tests: cooldown re-entry keeps aborting; both sibling shapes abort
- attribution: map zhangyswx@163.com -> YusenZhang0601
2026-08-26 13:02:27 +05:30
TonyRainforest fa210e5a96 fix(compressor): abort compression on empty-content provider degradation to prevent context loss (#94448)
When an auxiliary or main summarizer LLM returns an HTTP 200 with an empty or whitespace-only response (e.g., degraded provider/channel), abort compression and preserve the full conversation context rather than falling through to the destructive static-fallback branch that drops the middle window.

- Track _last_summary_empty_content_failure across _generate_summary() and compress()
- Attempt fallback to the main model when an aux model returns empty content
- Abort compression and preserve all messages intact if no valid summary can be generated
- Record summary_empty_content_failure in telemetry and log actionable diagnostic guidance
- Add comprehensive unit tests in tests/agent/test_context_compressor.py

Fixes #94448
2026-08-26 13:02:27 +05:30
kshitijk4poor 635232ec4e fix(codex): canonicalize fc_-only tool-result ids to match the call side
The sweeper review on #49224 flagged that the assistant branch synthesizes
call_<suffix> from an fc_-only id while the tool-result branch kept the raw
fc_... string — so an oversized pair hashed to two DIFFERENT clamped
surrogates and the function_call_output arrived unmatched (HTTP 400).

Canonicalize the tool-result side to the same call_<suffix> before
clamping. Also fixes the pre-existing short-fc_ pairing mismatch
(call_short123 vs fc_short123). Regression test covers both lengths.
2026-08-26 12:58:35 +05:30
kshitijk4poor 31485d50ea fix(codex): sanitize replayed function_call.name to Responses API pattern (#31666)
A degenerate tool name stored in conversation history (dots, spaces,
unicode from an earlier model degeneration) bricks every subsequent
Codex Responses turn with a non-retryable HTTP 400:
  Invalid input[N].name: string does not match pattern '^[a-zA-Z0-9_-]+'

The 400 replays forever until the user manually starts a new session.

Add _sanitize_replayed_fn_name() — replaces invalid chars with '_'
(runs collapsed), degrades all-invalid names to 'fn' instead of empty
(an empty name would trade one 400 for a preflight ValueError).  Applied
at both replay sites: the chat-message converter and the preflight
choke-point.  Live tool-definition names are left untouched — they must
match the dispatch registry exactly.  Pairing is by call_id, so
renaming a replayed function_call is safe.

call_id overflow (the sibling half of #49224) was already fixed on main
by #73492 (_clamp_responses_call_id); this commit covers the remaining
invalid-name defect.

Credit: @Morad37 (#31678 — identified the bug, the replay sites, and
the regex contract), @lubosxyz (#49224 — replace-not-strip semantics
and 'fn' fallback to avoid the empty-name trap).

Fixes #31666
2026-08-26 12:58:35 +05:30
fangliquanflq 66186dc58f fix(desktop): keep bot reconciliation off inactive backends 2026-08-26 00:21:29 -07:00
kshitijk4poor fab534b503 fix: omit User-Agent from anonymous OpenViking identity probes
Anonymous probes (_anonymous_json) are designed to probe server identity
before disclosing credentials. Sending the Hermes version on these probes
would fingerprint the exact version to an untrusted/MITM endpoint.

Keep User-Agent on authenticated requests (_headers) and multipart uploads
(_multipart_headers), which already send credentials.
2026-08-26 12:43:17 +05:30
ehz0ah 3db5267008 feat(openviking): identify Hermes requests 2026-08-26 12:43:17 +05:30
David Metcalfe b2ed58c415 test(desktop): move NS*UsageDescription pin from pytest to Vitest (tests-js)
Address maintainer review feedback (PR #66215, comment by @teknium1):

> `tests/test_desktop_mac_entitlements.py:47` reads `apps/desktop/package.json`
> from pytest. `AGENTS.md:1319-1329` requires assertions about `package.json`
> and JS-side artifacts to be in the JS/Vitest suite; otherwise CI
> classification can skip the regression test on a JS-only change.

The CI change classifier (`scripts/ci/classify_changes.py`) marks
`apps/desktop/package.json` as `_FRONTEND` (in `_PY_SKIP`), so a Python test
that reads it would be skipped on a JS-only PR — regression goes green on
the PR, red on main.

Move the regression to `tests-js/desktop-mac-usage-descriptions.test.ts`,
following the same convention as commit dbf86b923 ("test: port macOS
entitlements test from Python to vitest"), which ports an earlier Python
entitlements regression into `tests-js/desktop-mac-entitlements.test.ts`
for the identical reason. The new file is a sibling of that one — both
pin Desktop macOS manifest contracts, but they assert against different
files (`entitlements.mac.plist` vs `build.mac.extendInfo` in package.json).

The Vitest port mirrors the original assertions 1:1: every
`NS*UsageDescription` key pinned (parametrized over key + required
substring + reason), no leading/trailing whitespace or newlines in any
`extendInfo` string, and a drift-protection assertion that fails when a
new privacy key is added to the build config without a matching row.

A runtime type guard on `extendInfo` ensures a non-string plist scalar
raises a clean assertion error here ("`X` in build.mac.extendInfo must
be a string (got boolean)") rather than crashing the test runner with
`value.trim is not a function` deep in the whitespace test — caught by
Flash + GPT-OSS cross-vendor review.

Verified:
- `cd tests-js && npm run check` → typecheck clean, 14/14 tests pass
  (4 files including the new one with 5 tests).
- Mutation: removing `NSAppleMusicUsageDescription` from
  `apps/desktop/package.json` flips 1 test red with the exact symptom
  ("Info.plist privacy usage description \`NSAppleMusicUsageDescription\`
  is missing"). Restore → 14/14 green.
- Mutation: adding an unpinned `NSSpeechRecognitionUsageDescription` with
  whitespace flips 2 tests red (drift-protection + whitespace).
- Mutation: adding a non-string `CFBundleBooleanTest: true` flips the
  whole file red with the clean "must be a string (got boolean)"
  assertion (no downstream crash).
- `apps/desktop` Electron Vitest project still passes (42 files,
  432 tests + 1 skipped).

Closes the maintainer comment thread on PR #66215.

Fixes #54551
2026-08-25 23:33:46 -07:00
David Metcalfe f0e9902664 fix(desktop): declare NSAppleMusicUsageDescription to disclaim MediaLibrary TCC prompt
The Hermes Desktop renderer initializes Chromium's audio stack on user
gesture (completion chimes via Web Audio API in completion-sound.ts,
voice TTS via voice-playback.ts, mic capture via use-mic-recorder.ts,
and an eager AudioContext prime in haptics-provider.tsx). On macOS 26+,
that initialization registers the helper with the MediaLibrary TCC
service (kTCCServiceMediaLibrary), which surfaces to the user as a
"Hermes wants to access Music" permission prompt even though Hermes
never reads or writes the Apple Music library.

The Info.plist (built from apps/desktop/package.json's build.mac.extendInfo)
already declares NSAudioCaptureUsageDescription and
NSMicrophoneUsageDescription, but NSAppleMusicUsageDescription was missing
from the desktop app entirely. macOS therefore shows a system-default or
generic prompt for the MediaLibrary bucket instead of an honest description
from the app.

Fix
---
Add NSAppleMusicUsageDescription to build.mac.extendInfo with copy that
disclaims Music library access while explaining the system audio stack
uses voice, TTS, and completion sounds.

Add tests/test_desktop_mac_entitlements.py to pin every NS*UsageDescription
key declared in the Desktop build config. The test:
- parametrized over a (key, required_substring, reason) table
- asserts no leading/trailing whitespace and no newline chars in any usage
  string (electron-builder passes them through verbatim; control chars
  render as broken prompt text)
- asserts drift-protection: a new NS*UsageDescription key added to the
  build config without a matching test row causes a hard failure

Pattern reference: PR #59486 ("fix(desktop): add macOS contacts privacy
strings") is the open canonical for the same shape of fix for Contacts;
PR #64582 / PR #65220 extend it for Reminders. The closed duplicate PRs

Related, not in this PR
-----------------------
- PR #62601 (sounddevice on macOS) is the gateway/CLI side of the same
  kTCCServiceMediaLibrary trigger.
- PR #45952 (macOS permission broker foundation) is architectural work
  for centralized TCC handling; this fix does not depend on it.
- PR #52839 (browser automation Chrome launch) mutes Chromium audio in
  a different surface; the same pattern is recorded there.

Fixes #54551
2026-08-25 23:33:46 -07:00
David Metcalfe c0b5a8e15d fix(desktop): return True when fallback sign + strict verification succeed
The legacy ad-hoc fallback signed and verified successfully but still
fell through to return False, contradicting the fixup's documented
contract. The success witness codified the contradiction. Return True
on the verified success path; the caller ignores the return value, so
no behavior change beyond the contract correction.
2026-08-25 23:23:11 -07:00
David Metcalfe 177688e31e fix(desktop): never delete safeStorage keychain item in the updater
Addresses round-2 review feedback on #90961. The previous commits
scoped the keychain deletion to the legacy ad-hoc fallback, but the
reviewer correctly held the blocker: the fallback ran codesign with
check=False, ignored the result, and unconditionally deleted 'Hermes
Safe Storage' — permanently orphaning gateway and native OAuth
credentials even when signing failed or a configured identity had
failed and routed into the fallback.

This commit removes the deletion entirely:
- _desktop_macos_reset_keychain_safe_storage is gone; no code path
  touches the keychain item anymore.
- The legacy fallback now checks the codesign result and runs
  codesign --verify --deep --strict; any failure leaves the item
  untouched and prints a warning.
- The keychain prompt after an ad-hoc re-sign is recoverable
  (Always Allow updates the ACL partition list and preserves the
  key); deletion is not. The durable proof-carrying migration
  belongs in Electron (safeStorage can read the old key) and is
  tracked as a follow-up.

Tests: 4 witnesses (stable path, default no-config success, fallback
failure, fallback success) all mutation-verified against both the
deletion regression and the ignored-codesign-result regression.
2026-08-25 23:23:11 -07:00
David Metcalfe 368ea2d88b test(desktop): mark keychain-reset scoping tests macos_only
The fixup no-ops on non-macOS (sys.platform guard), so the new
regression tests must carry the same @pytest.mark.macos_only marker
as their siblings (test_relaunchable_fixup_falls_back_to_legacy_adhoc_on_failure).
Without it the legacy-adhoc test failed on the Linux CI runner where
the fixup returns True before reaching the reset path.
2026-08-25 23:23:11 -07:00
David Metcalfe 91dcca9a9b fix(desktop): scope keychain reset to the legacy ad-hoc fallback only
The previous commit deleted the 'Hermes Safe Storage' keychain item after
every successful re-sign, including the stable certificate-anchored
identity path. On that path the designated requirement is stable across
rebuilds, so after the first launch under the new identity the keychain
ACL already matches; deleting the item on every update permanently
orphaned gateway-token and native-OAuth credentials that were working
fine (both are safeStorage-backed: electron/main.ts connection config
and native-oauth-tokens.json).

Addresses review feedback on #90961:
- Rename _desktop_macos_update_keychain_acl -> _desktop_macos_reset_keychain_safe_storage (it deletes, it does not update an ACL).
- Only invoke it on the legacy ad-hoc fallback path, where every rebuild
  produces a new cdhash so the ACL can never match and the alternative
  is a recurring prompt. The trade-off (re-enter credentials once per
  update) is documented; the durable fix is a stable signing identity.
- Add regression tests: stable path must NOT reset, ad-hoc fallback MUST.
2026-08-25 23:23:11 -07:00
Jeremy cab6c4f70b fix(desktop): fail closed when messaging DELETE cannot resolve an owner
A listed profile-less row in a multi-profile setup must not DELETE/archive against the primary backend. Unresolved ownership keeps the row, pins, and unread state and never calls the mutation.
2026-08-25 23:15:49 -07:00
Teknium 2664599644 test(gateway): prove the profile_route_rejected sentinel is observable
Review follow-up for the salvaged #94848: the ProfileRouteRejected marker
looked write-only inside the primary handler. Document that
_handle_message's ingress gate reads the same marker to drop the message
fail-closed, and add a regression test showing the rejected route is
stamped once, dispatch falls back to the default home, and routing is not
re-run on redelivery.
2026-08-25 23:15:15 -07:00
GarrettGlass 2afed50863 fix(gateway): authorize routed messages in transport scope 2026-08-25 23:15:15 -07:00
GarrettGlass 9ab748abb9 fix(gateway): scope routed history before session lookup 2026-08-25 23:15:15 -07:00
YappLeCunt c6ae9325ec fix(computer-use): default the cursor overlay off on Linux X11
cua-driver maps the agent-cursor overlay as a fullscreen, always-on-top,
all-workspaces X11 window (save-unders composited). When a computer-use
session ends uncleanly — an agent interrupted mid-capture, a stale target
window, a driver error — that window can be left stuck above every app on
every workspace, wedging desktop input until the app is restarted. This
is the same failure class as the HUD's transparent always-on-top window on
Mutter/X11 (#83473), and it bit a real user: an interrupted capture froze
the desktop, the app had to be force-restarted, and the overlay window was
still mapped fullscreen afterwards.

The overlay is cosmetic (a tinted cursor sprite); the driver, captures,
and synthetic input all work without it. `_cua_no_overlay()` already
defaulted it off on macOS (idle CPU redraw loop, #28152/#47032) and
headless/WSL2 Linux; this extends the same auto-detect to X11 desktop
sessions, where raw X11 stacking has no compositor-owned surface to tear
down with the driver's connection. Wayland keeps the overlay: the
compositor owns the layer-surface lifecycle, so a dead driver cannot leave
a stuck top window.

Behavior contract unchanged: an explicit `computer_use.no_overlay: false`
still restores the cursor on any platform, and `true` forces it off.

Tests: X11 (DISPLAY set, no Wayland env) and XDG_SESSION_TYPE=x11
auto-detect off; Wayland keeps the overlay; explicit false overrides
auto-detection on X11.
2026-08-25 23:14:34 -07:00
ruangraung dce4abe917 fix(gateway): log secondary startup-reconnect handoff failures instead of dropping them
Review follow-up to the AI code-review pass on PR #92074: the bridge task's
handoff into _schedule_secondary_profile_reconnect was unguarded at both call
sites inside the parked coroutine. The scheduler touches live registries
(_profile_failed_platforms slot creation, background-task registration), so an
unexpected raise there would kill the parked task as an unretrieved-task
exception — logged only at GC time via "Task exception was never retrieved",
where no operator ever looks. A fix whose entire purpose is to stop a platform
dying silently should not contain its own silent-death path; both handoff sites
now wrap the scheduler call with logger.exception so the failure lands in
gateway.log with profile and platform context.

The early-exit branch (gateway already _running when the bridge starts) had the
identical exposure and is guarded the same way — same bug class, fixed together.
Regression test drives a handoff raise end-to-end through the real bridge task:
the await completes cleanly, the error is captured in gateway.run's logger, and
no adapter or failed-platform slot leaks behind the failed handoff.
2026-08-25 22:55:07 -07:00
ruangraung 96489f3c1b fix(gateway): schedule secondary-profile reconnect when initial adapter connect fails
When gateway.multiplex_profiles is active, a secondary profile whose platform
adapter fails its initial connect at startup was silently given up on: the
failure branches in _start_one_profile_adapters() logged and disconnected, but
never scheduled recovery. One unlucky connect window during a Telegram API
outage left the profile permanently silent until manual restart (~80 min in
the observed incident), while the mid-run fatal path already recovers via
_handle_profile_adapter_fatal_error() -> _schedule_secondary_profile_reconnect().
The same gap hit both failure shapes: a clean False return from
_connect_initial_adapter_with_timeout() and an exception escaping it.

Fix: call _schedule_secondary_profile_startup_reconnect() from both startup
failure branches after _safe_adapter_disconnect(). Because secondary adapters
are started mid-start(), before self._running flips True, the regular
scheduler's not-self._running guard would silently drop the request — so the
new bridge parks a background task until startup completes (or shutdown
begins) and then hands off to _schedule_secondary_profile_reconnect()
verbatim: backoff, fresh-adapter rebuild under the profile runtime scope,
slot dedupe, and shutdown cancellation all come from the existing path.
Non-retryable failures are dropped at scheduling time exactly as the regular
scheduler drops them, keeping duplicate-credential/auth-failed startups dead
instead of looping.

Fixes #92064
2026-08-25 22:55:07 -07:00
webtecnica cc5ff96f9e fix(macos): stable TCC anchor for uv-managed python interpreter (#85345) 2026-08-25 22:10:12 -07:00
Teknium a329346f4f test: point drain-timeout assertion at _CUA_INSTALLER_DRAIN_GRACE
The #87196/#87720 conflict resolution kept the bounded-drain helper and
its constant; the windows_only kill-tree test still asserted the dropped
_CUA_INSTALLER_REAP_TIMEOUT name. Same 2-communicate contract, surviving
constant.
2026-08-25 21:59:52 -07:00
Teknium 23c3f5086e feat(update): unattended-safe cua-driver refresh with fail-fast preflights
Re-enables the routine confirmed-upgrade path on Windows that #95008
deferred wholesale, now that every unattended-hostile surface is closed:

- stdin=DEVNULL (salvaged #79871): upstream's Read-Host consent prompt
  can't block a hidden console.
- Bounded post-kill drain (salvaged #87720): a kill-surviving descendant
  holding the stdout pipe can't strand the update past its ceiling.
- 120s background ceiling (salvaged #87196): safe now that a legitimate
  600s lock wait can't occur on this path.
- NEW lock preflight: upstream's install lock held by a live process ->
  skip in ~0s instead of eating its 600s stale-lock window (the actual
  11-minute hang observed 2026-08-25; UAC was a red herring — base
  install is no-admin by upstream design).
- NEW 5s network preflight: github.com unreachable -> skip immediately.
- Windows unattended runs pass -NoAutoStart, skipping the ONLY
  install.ps1 branch that self-elevates (autostart task re-registration).
- Timeout diagnosability: partial installer output is logged on kill so
  the next hang names its stage instead of dying silently.

Contract repairs and fresh installs stay interactive-only (SmartScreen /
first-time elevation legitimately need a human).
2026-08-25 21:59:52 -07:00