Commit Graph

14240 Commits

Author SHA1 Message Date
kshitijk4poor 365cbc242b test: assert omitted attach_to_session stays absent from formatted list output
Closes the gap flagged in review: the raw store was checked but not the
_format_job surface.
2026-08-26 16:06:41 +05:30
StanleyStetson 5e9adc9e4d fix(cron): forward attach_to_session through cronjob handler
The public schema and job store already support per-job
attach_to_session, but the registry adapter dropped the argument.
Create silently omitted the field; update reported "No updates provided."

Fixes #84802
2026-08-26 16:06:41 +05:30
kshitijk4poor 2f425872ef test: extend exact-shape assertion in relay-delivery guard for provenance tag
tests/cron/test_cron_relay_delivery_guards.py landed on main after #89329
branched; its exact-dict assertion needs the new _resolved_from field the
salvaged commit adds to origin-resolved targets.
2026-08-26 16:06:09 +05:30
Victor Kyriazakos 580daa7b96 fix(cron): mirror continuable-cron briefs for origin-fallback and opted-in explicit targets
A managed cron (created by a provisioning script, not from a live gateway
chat) never captures an origin. With cron.mirror_delivery: true and
deliver: origin, its brief was delivered to the home channel — the
user's own DM — but the transcript mirror and the in_channel session
seed were silently skipped: _target_matches_origin returns False for an
empty origin, and the whole continuable machinery keys off that check.
A user replying to the brief landed in a session with no record of it.
Field report 2026-08-17 (enterprise, Slack DM surface).

The June origin-scoping refactor (c06ceb3232) was written to exclude
broadcasts, and the exclusion is kept. What changes is the
classification: a home-channel FALLBACK for deliver=origin is the user's
primary conversation standing in for the origin, not a broadcast.

Changes:
- Delivery targets carry a resolution-provenance tag (_resolved_from:
  origin / origin_fallback / explicit; broadcast expansions untagged).
- _target_mirror_eligible replaces the bare origin check at the mirror
  gate: origin unchanged; origin_fallback eligible under the same flags
  as origin (per-job attach_to_session wins, else global
  cron.mirror_delivery); explicit platform:chat targets eligible ONLY
  under per-job attach_to_session — the global flag never activates
  them, so it cannot start writing transcript entries into arbitrary
  explicitly-addressed chats. 'all'/bare-platform stay never-eligible.
- Dedup OR-merges provenance so 'origin,all' resolving to the same chat
  keeps eligibility regardless of token order.
- _inchannel_seed_allowed guards the flat-session seed: group-channel
  session keys are user-isolated, so a seed without a user_id (origin-
  less job into a shared channel) would create an orphan session no
  reply resolves to — those targets fall back to the plain mirror. DM
  targets (keys don't embed user_id) always seed.
- cronjob tool schema text updated to describe the new attach scope.

Behavioral note: origin-less deliver=origin jobs under global
mirror_delivery now activate the full continuable path — on default
'thread' surface this opens a dedicated thread in the home channel
where the brief previously posted flat. That is the documented
continuable behavior; the silent flat post was the bug.

15 new tests (tests/cron/test_mirror_origin_fallback.py): eligibility
matrix (origin/fallback/explicit/all/bare/other-chat), dedup order
both ways, end-to-end mirror via _deliver_result for all four shapes,
origin regression control, seed user_id guard.
2026-08-26 16:06:09 +05:30
Teknium ad7b7255ab fix(state): renaming a bot's canonical Bot Chat is refused — the title IS the identity (#92473)
Bot Mode resolves the forever-chat by exact-title lookup on
(profile, 'Bot Chat'); no session-id pointer exists. A user rename
therefore orphaned the whole conversation: resolution missed, the next
click minted an empty replacement, and UNIQUE(title) then blocked ever
renaming back. Refuse the rename at SessionDB._set_session_title — the
single write path every surface funnels through (gateway session.title,
/title, CLI rename, REST). Hidden discriminates the registry row, so a
normal visible session a user happens to call 'Bot Chat' stays freely
renameable; re-asserting the same canonical title stays a no-op.
2026-08-26 03:21:42 -07:00
Teknium 45db70a80a fix(computer-use): fail closed on unverified CuaDriver.app + background launch
Hardening on top of the TCC daemon-identity salvage:
- _validate_cua_driver_app_signature: codesign -dv gate requiring EXACT
  Identifier=com.trycua.driver and the official team (4YEC26S9KF) before
  /usr/bin/open hands the bundle to LaunchServices — the identity fix must
  not double as a launcher for arbitrary/impostor bundles (suffixed
  identifiers and wrong teams rejected; unsigned dev builds only via
  computer_use.allow_unsigned_driver: true in config.yaml).
- _resolve_cua_driver_app_path: derive the bundle ONLY from the resolved
  driver binary — the /Applications fallback could launch a DIFFERENT
  install than the manifest resolved.
- open -n -g: don't activate/steal focus when launching the daemon.
- 7 new tests incl. sabotage-verified exact-match assertions.

Grafted from #76433's review direction (@Chadmc9889's original fail-closed
validation requirement).

Co-authored-by: Chadmc9889 <Chadmc9889@users.noreply.github.com>
2026-08-26 03:21:37 -07:00
projetsjsl 4746f614be fix(computer-use): preserve macOS TCC daemon identity
Launch private computer-use daemons through CuaDriver.app so Screen
Recording authorization remains attached to its stable bundle identity
instead of Hermes' ad-hoc signature.

Co-Authored-By: GPT-5.6 Codex <noreply@openai.com>
2026-08-26 03:21:37 -07:00
Teknium 1fe0f2f3ac feat(cron): import-error cron failures now name gateway code skew and the one-command fix (#95294 part 3)
When an agent cron job dies with an import-class error (cannot import
name / ModuleNotFoundError / ImportError), the failure summarizer — which
runs inside the gateway process — now consults gateway.code_skew: if the
process booted on a different revision than disk HEAD, the delivered
message appends 'gateway is running stale code (booted on X, disk is at
Y) — run hermes gateway restart'. Turns the reported two-day mystery
(15 missed jobs, identical ImportError, no explanation) into a one-line
fix instruction on the first failure.

Fail-safe by construction: skew detection returns None on non-git
installs and processes without a boot fingerprint, the probe seam
swallows every exception, and no_agent script jobs (fresh subprocess,
consistent imports) fall through to the generic cleaner — their
ImportErrors are the script's own problem, and blaming gateway skew
there would send the reader to the wrong place (same mode-gating as the
provider branches).

Reuses gateway/code_skew.py (the /model-switch skew detector) rather
than adding a second fingerprint reader.
2026-08-26 01:23:15 -07:00
Teknium f0c0c986c4 test: pin overlay policy off in embedded-daemon socket/ack contract test
The embedded spawn now consults the overlay policy (capability probe via
subprocess.run) when _cua_no_overlay() is true — which it is on headless
CI since the Linux X11 default flip. The fixed two-entry run side_effect
in this test didn't budget for the probe call; pin the policy off since
this test pins the socket/ack contract, not overlay behavior.
2026-08-26 00:54:36 -07:00
cvillarroel2 1a7f83a73b fix(computer_use): disable embedded daemon overlay 2026-08-26 00:54:36 -07:00
kshitijk4poor 4ba2608524 fix(compressor): widen empty-content abort to sibling no-response shapes + snapshot state field
Follow-up to PR #94531 salvage:
- classify the auxiliary boundary's terminal 'None response' /
  'invalid response' errors (#7264) into the same empty-content abort
  carve-out so those shapes also preserve the session (#94459's wider
  classification, sibling shapes from #94448)
- register _last_summary_empty_content_failure in
  _COMPRESSOR_ATTEMPT_STATE_FIELDS so pre-commit hard-cancel rollback
  restores the flag (conversation_compression snapshot allow-list)
- tests: cooldown re-entry keeps aborting; both sibling shapes abort
- attribution: map zhangyswx@163.com -> YusenZhang0601
2026-08-26 13:02:27 +05:30
TonyRainforest fa210e5a96 fix(compressor): abort compression on empty-content provider degradation to prevent context loss (#94448)
When an auxiliary or main summarizer LLM returns an HTTP 200 with an empty or whitespace-only response (e.g., degraded provider/channel), abort compression and preserve the full conversation context rather than falling through to the destructive static-fallback branch that drops the middle window.

- Track _last_summary_empty_content_failure across _generate_summary() and compress()
- Attempt fallback to the main model when an aux model returns empty content
- Abort compression and preserve all messages intact if no valid summary can be generated
- Record summary_empty_content_failure in telemetry and log actionable diagnostic guidance
- Add comprehensive unit tests in tests/agent/test_context_compressor.py

Fixes #94448
2026-08-26 13:02:27 +05:30
kshitijk4poor 635232ec4e fix(codex): canonicalize fc_-only tool-result ids to match the call side
The sweeper review on #49224 flagged that the assistant branch synthesizes
call_<suffix> from an fc_-only id while the tool-result branch kept the raw
fc_... string — so an oversized pair hashed to two DIFFERENT clamped
surrogates and the function_call_output arrived unmatched (HTTP 400).

Canonicalize the tool-result side to the same call_<suffix> before
clamping. Also fixes the pre-existing short-fc_ pairing mismatch
(call_short123 vs fc_short123). Regression test covers both lengths.
2026-08-26 12:58:35 +05:30
kshitijk4poor 31485d50ea fix(codex): sanitize replayed function_call.name to Responses API pattern (#31666)
A degenerate tool name stored in conversation history (dots, spaces,
unicode from an earlier model degeneration) bricks every subsequent
Codex Responses turn with a non-retryable HTTP 400:
  Invalid input[N].name: string does not match pattern '^[a-zA-Z0-9_-]+'

The 400 replays forever until the user manually starts a new session.

Add _sanitize_replayed_fn_name() — replaces invalid chars with '_'
(runs collapsed), degrades all-invalid names to 'fn' instead of empty
(an empty name would trade one 400 for a preflight ValueError).  Applied
at both replay sites: the chat-message converter and the preflight
choke-point.  Live tool-definition names are left untouched — they must
match the dispatch registry exactly.  Pairing is by call_id, so
renaming a replayed function_call is safe.

call_id overflow (the sibling half of #49224) was already fixed on main
by #73492 (_clamp_responses_call_id); this commit covers the remaining
invalid-name defect.

Credit: @Morad37 (#31678 — identified the bug, the replay sites, and
the regex contract), @lubosxyz (#49224 — replace-not-strip semantics
and 'fn' fallback to avoid the empty-name trap).

Fixes #31666
2026-08-26 12:58:35 +05:30
fangliquanflq 66186dc58f fix(desktop): keep bot reconciliation off inactive backends 2026-08-26 00:21:29 -07:00
kshitijk4poor fab534b503 fix: omit User-Agent from anonymous OpenViking identity probes
Anonymous probes (_anonymous_json) are designed to probe server identity
before disclosing credentials. Sending the Hermes version on these probes
would fingerprint the exact version to an untrusted/MITM endpoint.

Keep User-Agent on authenticated requests (_headers) and multipart uploads
(_multipart_headers), which already send credentials.
2026-08-26 12:43:17 +05:30
ehz0ah 3db5267008 feat(openviking): identify Hermes requests 2026-08-26 12:43:17 +05:30
Ben Barclay d0a7144ba1 fix(telemetry): per-row claiming, mid-pass consent re-check, narrower 4xx
Second independent review found the lease fix incomplete. Reproduced
each finding before fixing.

BLOCKER — the batch lease expired mid-pass. _claim took up to 20 rows
under ONE shared lease, but a single package can legally consume ~96s
(three 30s timeouts plus 1s+5s backoff), so a full batch runs ~1900s
against a 180s lease. Later rows' leases expired while this pass still
held them, and another process re-sent them. Reproduced: 192s elapsed,
pkg-2 POSTed twice.

Packages are now claimed ONE AT A TIME, immediately before being sent,
so a lease only has to cover the package actually in flight. Verified:
same scenario now sends each package exactly once.

HIGH — revoking consent did not stop a running pass. The runtime read
send consent once before starting the thread, so a pass could keep
transmitting for minutes after a user set send: false, contradicting
the documented promise that it 'stops transmission immediately'.
Consent is now re-read before every package and fails CLOSED if it
cannot be established.

MEDIUM — all non-429 4xx were treated as permanent, discarding data.
403 is the ingest service's own origin guard: a Transform Rule or edge
misconfiguration would have permanently dropped every package sent
during the incident. Only 400 (malformed envelope) and 413 (over the
1 MiB cap) are terminal now; everything else retries.

MEDIUM — valid JSON that is not an object blocked the whole queue.
json.loads('["a"]') succeeds, then .get() raised AttributeError inside
the claim transaction, rolling it back and starving every healthy
package behind it. Payload shape and install_id are now validated, and
an unusable row is rejected individually.

LOW — the clock-rollback comment and test name claimed the opposite of
the code. The behaviour is right (a future issued_at means the recorded
age is untrustworthy, so reissue); the wording is now honest about it.

LOW — removed the stale HERMES_TELEMETRY_ENDPOINT reference left in
config_defaults after the override was deleted.

247 tests pass (was 234). Staging E2E re-run: both packages 202.
2026-08-26 17:10:43 +10:00
David Metcalfe b2ed58c415 test(desktop): move NS*UsageDescription pin from pytest to Vitest (tests-js)
Address maintainer review feedback (PR #66215, comment by @teknium1):

> `tests/test_desktop_mac_entitlements.py:47` reads `apps/desktop/package.json`
> from pytest. `AGENTS.md:1319-1329` requires assertions about `package.json`
> and JS-side artifacts to be in the JS/Vitest suite; otherwise CI
> classification can skip the regression test on a JS-only change.

The CI change classifier (`scripts/ci/classify_changes.py`) marks
`apps/desktop/package.json` as `_FRONTEND` (in `_PY_SKIP`), so a Python test
that reads it would be skipped on a JS-only PR — regression goes green on
the PR, red on main.

Move the regression to `tests-js/desktop-mac-usage-descriptions.test.ts`,
following the same convention as commit dbf86b923 ("test: port macOS
entitlements test from Python to vitest"), which ports an earlier Python
entitlements regression into `tests-js/desktop-mac-entitlements.test.ts`
for the identical reason. The new file is a sibling of that one — both
pin Desktop macOS manifest contracts, but they assert against different
files (`entitlements.mac.plist` vs `build.mac.extendInfo` in package.json).

The Vitest port mirrors the original assertions 1:1: every
`NS*UsageDescription` key pinned (parametrized over key + required
substring + reason), no leading/trailing whitespace or newlines in any
`extendInfo` string, and a drift-protection assertion that fails when a
new privacy key is added to the build config without a matching row.

A runtime type guard on `extendInfo` ensures a non-string plist scalar
raises a clean assertion error here ("`X` in build.mac.extendInfo must
be a string (got boolean)") rather than crashing the test runner with
`value.trim is not a function` deep in the whitespace test — caught by
Flash + GPT-OSS cross-vendor review.

Verified:
- `cd tests-js && npm run check` → typecheck clean, 14/14 tests pass
  (4 files including the new one with 5 tests).
- Mutation: removing `NSAppleMusicUsageDescription` from
  `apps/desktop/package.json` flips 1 test red with the exact symptom
  ("Info.plist privacy usage description \`NSAppleMusicUsageDescription\`
  is missing"). Restore → 14/14 green.
- Mutation: adding an unpinned `NSSpeechRecognitionUsageDescription` with
  whitespace flips 2 tests red (drift-protection + whitespace).
- Mutation: adding a non-string `CFBundleBooleanTest: true` flips the
  whole file red with the clean "must be a string (got boolean)"
  assertion (no downstream crash).
- `apps/desktop` Electron Vitest project still passes (42 files,
  432 tests + 1 skipped).

Closes the maintainer comment thread on PR #66215.

Fixes #54551
2026-08-25 23:33:46 -07:00
David Metcalfe f0e9902664 fix(desktop): declare NSAppleMusicUsageDescription to disclaim MediaLibrary TCC prompt
The Hermes Desktop renderer initializes Chromium's audio stack on user
gesture (completion chimes via Web Audio API in completion-sound.ts,
voice TTS via voice-playback.ts, mic capture via use-mic-recorder.ts,
and an eager AudioContext prime in haptics-provider.tsx). On macOS 26+,
that initialization registers the helper with the MediaLibrary TCC
service (kTCCServiceMediaLibrary), which surfaces to the user as a
"Hermes wants to access Music" permission prompt even though Hermes
never reads or writes the Apple Music library.

The Info.plist (built from apps/desktop/package.json's build.mac.extendInfo)
already declares NSAudioCaptureUsageDescription and
NSMicrophoneUsageDescription, but NSAppleMusicUsageDescription was missing
from the desktop app entirely. macOS therefore shows a system-default or
generic prompt for the MediaLibrary bucket instead of an honest description
from the app.

Fix
---
Add NSAppleMusicUsageDescription to build.mac.extendInfo with copy that
disclaims Music library access while explaining the system audio stack
uses voice, TTS, and completion sounds.

Add tests/test_desktop_mac_entitlements.py to pin every NS*UsageDescription
key declared in the Desktop build config. The test:
- parametrized over a (key, required_substring, reason) table
- asserts no leading/trailing whitespace and no newline chars in any usage
  string (electron-builder passes them through verbatim; control chars
  render as broken prompt text)
- asserts drift-protection: a new NS*UsageDescription key added to the
  build config without a matching test row causes a hard failure

Pattern reference: PR #59486 ("fix(desktop): add macOS contacts privacy
strings") is the open canonical for the same shape of fix for Contacts;
PR #64582 / PR #65220 extend it for Reminders. The closed duplicate PRs

Related, not in this PR
-----------------------
- PR #62601 (sounddevice on macOS) is the gateway/CLI side of the same
  kTCCServiceMediaLibrary trigger.
- PR #45952 (macOS permission broker foundation) is architectural work
  for centralized TCC handling; this fix does not depend on it.
- PR #52839 (browser automation Chrome launch) mutes Chromium audio in
  a different surface; the same pattern is recorded there.

Fixes #54551
2026-08-25 23:33:46 -07:00
Ben Barclay 49757d5e39 fix(telemetry): address review findings on the shared-metrics sender
Independent review found the claim mechanism did not work. Reproduced
against the real store: two senders POSTed the same package.

The claim wrote next_attempt_at = now, but selection requires
next_attempt_at <= now, so a concurrent pass matched the same row
immediately. It now writes a LEASE INTO THE FUTURE
(_CLAIM_LEASE_SECONDS), which is what actually excludes another pass,
and expires by itself if a process dies mid-send. _mark is additionally
guarded on send_state so a straggler whose lease lapsed cannot
overwrite a completed send back to pending.

The old concurrency test could not fail: it raised AssertionError from
inside a transport, and _send_one catches every exception as a
retryable transport error. It now records what the second pass saw.

Also from review:

- shutdown() never joined the send thread; the join was only wired into
  deactivate(). A short-lived CLI therefore killed an in-flight send at
  exit, on the only cadence this feature has.
- Removed HERMES_TELEMETRY_ENDPOINT. AGENTS.md reserves HERMES_* for
  secrets, and a behavioural override here was a consent hazard: an
  inherited variable could silently redirect telemetry a user agreed to
  send to Nous. The staging E2E writes the endpoint into its throwaway
  profile instead, which also exercises the real config path.
- Added the  shared-metrics toggle that AGENTS.md requires
  as the third opt-in surface, delegating to the setup prompt so the
  consent rules stay in one place.
- Non-429 4xx (401/403/404/413/422) are now permanent. Only 400 was,
  so a wrong path or oversized body retried every 15 minutes for 30
  days until retention pruned it.
- The opt-in day is stamped when the user consents, not on the first
  send pass, which silently dropped the opt-in day whenever the next
  export crossed midnight UTC.
- gzip now uses mtime=0. The embedded timestamp made two sends of one
  package differ on the wire, so the 'byte-identical retry' E2E was
  comparing parsed bodies and could not have caught it. It now compares
  raw request bytes.
- Reconciled the three stale claims in relay-shared-metrics.md that
  said no remote-delivery path exists.

233 tests pass (was 213). Staging E2E re-run through the config path:
both packages 202, and the service logged both objects written to S3.
2026-08-26 16:31:52 +10:00
David Metcalfe c0b5a8e15d fix(desktop): return True when fallback sign + strict verification succeed
The legacy ad-hoc fallback signed and verified successfully but still
fell through to return False, contradicting the fixup's documented
contract. The success witness codified the contradiction. Return True
on the verified success path; the caller ignores the return value, so
no behavior change beyond the contract correction.
2026-08-25 23:23:11 -07:00
David Metcalfe 177688e31e fix(desktop): never delete safeStorage keychain item in the updater
Addresses round-2 review feedback on #90961. The previous commits
scoped the keychain deletion to the legacy ad-hoc fallback, but the
reviewer correctly held the blocker: the fallback ran codesign with
check=False, ignored the result, and unconditionally deleted 'Hermes
Safe Storage' — permanently orphaning gateway and native OAuth
credentials even when signing failed or a configured identity had
failed and routed into the fallback.

This commit removes the deletion entirely:
- _desktop_macos_reset_keychain_safe_storage is gone; no code path
  touches the keychain item anymore.
- The legacy fallback now checks the codesign result and runs
  codesign --verify --deep --strict; any failure leaves the item
  untouched and prints a warning.
- The keychain prompt after an ad-hoc re-sign is recoverable
  (Always Allow updates the ACL partition list and preserves the
  key); deletion is not. The durable proof-carrying migration
  belongs in Electron (safeStorage can read the old key) and is
  tracked as a follow-up.

Tests: 4 witnesses (stable path, default no-config success, fallback
failure, fallback success) all mutation-verified against both the
deletion regression and the ignored-codesign-result regression.
2026-08-25 23:23:11 -07:00
David Metcalfe 368ea2d88b test(desktop): mark keychain-reset scoping tests macos_only
The fixup no-ops on non-macOS (sys.platform guard), so the new
regression tests must carry the same @pytest.mark.macos_only marker
as their siblings (test_relaunchable_fixup_falls_back_to_legacy_adhoc_on_failure).
Without it the legacy-adhoc test failed on the Linux CI runner where
the fixup returns True before reaching the reset path.
2026-08-25 23:23:11 -07:00
David Metcalfe 91dcca9a9b fix(desktop): scope keychain reset to the legacy ad-hoc fallback only
The previous commit deleted the 'Hermes Safe Storage' keychain item after
every successful re-sign, including the stable certificate-anchored
identity path. On that path the designated requirement is stable across
rebuilds, so after the first launch under the new identity the keychain
ACL already matches; deleting the item on every update permanently
orphaned gateway-token and native-OAuth credentials that were working
fine (both are safeStorage-backed: electron/main.ts connection config
and native-oauth-tokens.json).

Addresses review feedback on #90961:
- Rename _desktop_macos_update_keychain_acl -> _desktop_macos_reset_keychain_safe_storage (it deletes, it does not update an ACL).
- Only invoke it on the legacy ad-hoc fallback path, where every rebuild
  produces a new cdhash so the ACL can never match and the alternative
  is a recurring prompt. The trade-off (re-enter credentials once per
  update) is documented; the durable fix is a stable signing identity.
- Add regression tests: stable path must NOT reset, ad-hoc fallback MUST.
2026-08-25 23:23:11 -07:00
Jeremy cab6c4f70b fix(desktop): fail closed when messaging DELETE cannot resolve an owner
A listed profile-less row in a multi-profile setup must not DELETE/archive against the primary backend. Unresolved ownership keeps the row, pins, and unread state and never calls the mutation.
2026-08-25 23:15:49 -07:00
Teknium 2664599644 test(gateway): prove the profile_route_rejected sentinel is observable
Review follow-up for the salvaged #94848: the ProfileRouteRejected marker
looked write-only inside the primary handler. Document that
_handle_message's ingress gate reads the same marker to drop the message
fail-closed, and add a regression test showing the rejected route is
stamped once, dispatch falls back to the default home, and routing is not
re-run on redelivery.
2026-08-25 23:15:15 -07:00
GarrettGlass 2afed50863 fix(gateway): authorize routed messages in transport scope 2026-08-25 23:15:15 -07:00
GarrettGlass 9ab748abb9 fix(gateway): scope routed history before session lookup 2026-08-25 23:15:15 -07:00
YappLeCunt c6ae9325ec fix(computer-use): default the cursor overlay off on Linux X11
cua-driver maps the agent-cursor overlay as a fullscreen, always-on-top,
all-workspaces X11 window (save-unders composited). When a computer-use
session ends uncleanly — an agent interrupted mid-capture, a stale target
window, a driver error — that window can be left stuck above every app on
every workspace, wedging desktop input until the app is restarted. This
is the same failure class as the HUD's transparent always-on-top window on
Mutter/X11 (#83473), and it bit a real user: an interrupted capture froze
the desktop, the app had to be force-restarted, and the overlay window was
still mapped fullscreen afterwards.

The overlay is cosmetic (a tinted cursor sprite); the driver, captures,
and synthetic input all work without it. `_cua_no_overlay()` already
defaulted it off on macOS (idle CPU redraw loop, #28152/#47032) and
headless/WSL2 Linux; this extends the same auto-detect to X11 desktop
sessions, where raw X11 stacking has no compositor-owned surface to tear
down with the driver's connection. Wayland keeps the overlay: the
compositor owns the layer-surface lifecycle, so a dead driver cannot leave
a stuck top window.

Behavior contract unchanged: an explicit `computer_use.no_overlay: false`
still restores the cursor on any platform, and `true` forces it off.

Tests: X11 (DISPLAY set, no Wayland env) and XDG_SESSION_TYPE=x11
auto-detect off; Wayland keeps the overlay; explicit false overrides
auto-detection on X11.
2026-08-25 23:14:34 -07:00
ruangraung dce4abe917 fix(gateway): log secondary startup-reconnect handoff failures instead of dropping them
Review follow-up to the AI code-review pass on PR #92074: the bridge task's
handoff into _schedule_secondary_profile_reconnect was unguarded at both call
sites inside the parked coroutine. The scheduler touches live registries
(_profile_failed_platforms slot creation, background-task registration), so an
unexpected raise there would kill the parked task as an unretrieved-task
exception — logged only at GC time via "Task exception was never retrieved",
where no operator ever looks. A fix whose entire purpose is to stop a platform
dying silently should not contain its own silent-death path; both handoff sites
now wrap the scheduler call with logger.exception so the failure lands in
gateway.log with profile and platform context.

The early-exit branch (gateway already _running when the bridge starts) had the
identical exposure and is guarded the same way — same bug class, fixed together.
Regression test drives a handoff raise end-to-end through the real bridge task:
the await completes cleanly, the error is captured in gateway.run's logger, and
no adapter or failed-platform slot leaks behind the failed handoff.
2026-08-25 22:55:07 -07:00
ruangraung 96489f3c1b fix(gateway): schedule secondary-profile reconnect when initial adapter connect fails
When gateway.multiplex_profiles is active, a secondary profile whose platform
adapter fails its initial connect at startup was silently given up on: the
failure branches in _start_one_profile_adapters() logged and disconnected, but
never scheduled recovery. One unlucky connect window during a Telegram API
outage left the profile permanently silent until manual restart (~80 min in
the observed incident), while the mid-run fatal path already recovers via
_handle_profile_adapter_fatal_error() -> _schedule_secondary_profile_reconnect().
The same gap hit both failure shapes: a clean False return from
_connect_initial_adapter_with_timeout() and an exception escaping it.

Fix: call _schedule_secondary_profile_startup_reconnect() from both startup
failure branches after _safe_adapter_disconnect(). Because secondary adapters
are started mid-start(), before self._running flips True, the regular
scheduler's not-self._running guard would silently drop the request — so the
new bridge parks a background task until startup completes (or shutdown
begins) and then hands off to _schedule_secondary_profile_reconnect()
verbatim: backoff, fresh-adapter rebuild under the profile runtime scope,
slot dedupe, and shutdown cancellation all come from the existing path.
Non-retryable failures are dropped at scheduling time exactly as the regular
scheduler drops them, keeping duplicate-credential/auth-failed startups dead
instead of looping.

Fixes #92064
2026-08-25 22:55:07 -07:00
Ben Barclay 6fdf6f4d4a feat(telemetry): run the send pass off the export hook
Step 6 of the shared-metrics exporter, plus a loopback E2E.

_export now triggers an opt-in send pass on a daemon thread. The hook
runs on finish_task — the user's interactive path — so a 30s network
timeout there would be felt directly; the thread keeps that latency off
the caller. A test asserts _export returns in under a second while a
send is deliberately blocked.

At most one pass is in flight per process: a queued second pass would
add nothing, because the next hook fire picks up whatever is still
pending. Shutdown joins the thread for at most two seconds, then lets
it go — the packages remain in SQLite and go out on the next run, so
blocking a user's exit on a slow network is the wrong trade.

Sending is resolved per pass from the profile's own config, so turning
it off takes effect at the next hook fire without a restart.

E2E (tests/hermes_cli/test_shared_metrics_sender_e2e.py): the real
sender against a real HTTPServer on loopback — actual urllib, gzip,
headers and sockets rather than an injected fake. Covers delivery and
sent-state, 400/429/5xx handling, a retry sending byte-identical
bytes, gzip shrinking a realistic 120-metric package and the server
parsing it back, install_id never crossing the wire, the outbox file
staying untouched, and a dead server deferring without raising.

Wiring tests: 12, all negative-space properties — no send without
opt-in, no blocking, no pile-up, no crash propagation.
2026-08-26 15:49:25 +10:00
Ben Barclay 00c75cea33 feat(telemetry): send exported packages to the ingest service
Steps 4, 5 and 7 of the shared-metrics exporter: the send logic, the
consent gate, and backoff plus multi-process claiming. These arrive
together because the sender is not correct without all three.

Contract handling: 202 marks sent; 400 is permanent and never retried;
429 honours Retry-After (clamped to a day so a bogus value cannot park a
package); 5xx, timeouts and transport errors retry three times in-process
with 1s/5s/25s full-jitter backoff, then defer to a later pass.

Consent is gated on the package's PERIOD, not its creation time. A period
is split across packages created on different days, so a created-at gate
would send a period's tail while dropping its head and silently
undercount the opt-in day — data that looks complete and is wrong. The
opt-in day is recorded once and never moves, so toggling sending off and
on does not re-open the pre-consent backlog.

Rows are claimed in a write transaction, which is what stops two Hermes
processes sharing one database from sending the same package twice.
next_attempt_at persists backoff across restarts, so a hard-down service
is not retried on every task completion.

The body is recomputed from payload_json rather than stored a second
time: json.dumps is deterministic here (verified against the real outbox
— 11 of 11 files reproduce byte-for-byte), and the only mutable input,
the derived identity, is frozen on the row at first attempt. That keeps
retries byte-identical across a salt rotation for ~36 bytes instead of a
duplicate ~11 KB payload.

The outbox directory is never written to or deleted from. A 202 updates
SQLite only, because those files are the user's 30-day local history and
retention already owns their lifecycle.

Tests: 33. Two of them caught real defects in this commit — an
unreadable row aborted the claim transaction and blocked every package
behind it, and the compression assertions were passing through an
injected fake that bypassed the code under test.
2026-08-26 15:45:02 +10:00
Ben Barclay 7ffd454df6 feat(telemetry): derive the transmitted install identity via keyed HMAC
Step 3 of the shared-metrics exporter.

The shared-metrics doc commits that a remote exporter 'must not reuse
the persistent local identifier by default'. install_id is therefore
never transmitted: each package carries
HMAC-SHA256(local-only rotation salt, install_id) instead.

Within a 30-day rotation window the value is stable, so distinct
installs remain countable — the first question the data has to answer.
Across windows it changes, bounding long-term linkability. The
derivation is one-way, so the service cannot recover install_id.

The salt lives in telemetry_state next to install_id, so removing the
shared-metrics directory resets both together and the documented reset
behaviour keeps working with no second cleanup path.

Rotation is deliberately not a bare 'age > interval' check: a clock
that jumps backwards must not read as an expired salt, and an
unparseable issued-at reissues instead of raising.

substitute_install_id replaces exactly one field and copies rather than
mutating, so payload schema evolution stays a sender-side concern.

Tests: 19, including that install_id never survives substitution, that
no other field changes, and — the property that keeps retries
contract-compliant — that a package rebuilt from a FROZEN derived id is
byte-stable across a salt rotation while a fresh derivation is not.
2026-08-26 15:41:18 +10:00
Ben Barclay e5180ab3df feat(telemetry): add opt-in send config and send-state columns
Step 1+2 of the shared-metrics exporter.

Config: telemetry.shared_metrics.send (default false) and .endpoint
(default production), resolved by a new shared_metrics_send_config
module. Precedence is HERMES_TELEMETRY_ENDPOINT > config > default; the
env var exists so the live staging E2E never has to mutate a user's
config. send requires enabled and never implies it — that combination
is a misconfiguration the user believes is working, so it logs an ERROR
once per process rather than silently doing nothing. Plaintext
endpoints are refused unless the host is loopback, so a typo cannot
send telemetry in clear text.

Per AGENTS.md, outbound telemetry needs a user-facing opt-in, so
setup_telemetry now prompts for sending as a second, separate question
and force-disables send when collection is turned off.

Storage: six additive nullable columns on package_outbox for send
bookkeeping. The store schema version deliberately does NOT move —
_ensure_schema_in_transaction raises on any version it does not
recognise and has no forward-compatibility branch, so bumping it would
hard-fail an older Hermes, a second profile on an older build, or a
rollback, against the same file. Old readers select named columns and
never SELECT *, so the additions are invisible to them.

Also corrects the two places that promised telemetry is never uploaded
(config_defaults comment and cli-config.yaml.example); leaving them
would make them false privacy statements once sending ships.

Tests: 26 covering config precedence, the enabled/send relationship,
transport safety, fresh-database creation, upgrade from a pre-send
database (rows preserved, version pinned, idempotent), and that the
shipped export query still runs. Mutation-checked: bumping the schema
version fails 5 of them.
2026-08-26 15:39:47 +10:00
webtecnica cc5ff96f9e fix(macos): stable TCC anchor for uv-managed python interpreter (#85345) 2026-08-25 22:10:12 -07:00
Teknium a329346f4f test: point drain-timeout assertion at _CUA_INSTALLER_DRAIN_GRACE
The #87196/#87720 conflict resolution kept the bounded-drain helper and
its constant; the windows_only kill-tree test still asserted the dropped
_CUA_INSTALLER_REAP_TIMEOUT name. Same 2-communicate contract, surviving
constant.
2026-08-25 21:59:52 -07:00
Teknium 23c3f5086e feat(update): unattended-safe cua-driver refresh with fail-fast preflights
Re-enables the routine confirmed-upgrade path on Windows that #95008
deferred wholesale, now that every unattended-hostile surface is closed:

- stdin=DEVNULL (salvaged #79871): upstream's Read-Host consent prompt
  can't block a hidden console.
- Bounded post-kill drain (salvaged #87720): a kill-surviving descendant
  holding the stdout pipe can't strand the update past its ceiling.
- 120s background ceiling (salvaged #87196): safe now that a legitimate
  600s lock wait can't occur on this path.
- NEW lock preflight: upstream's install lock held by a live process ->
  skip in ~0s instead of eating its 600s stale-lock window (the actual
  11-minute hang observed 2026-08-25; UAC was a red herring — base
  install is no-admin by upstream design).
- NEW 5s network preflight: github.com unreachable -> skip immediately.
- Windows unattended runs pass -NoAutoStart, skipping the ONLY
  install.ps1 branch that self-elevates (autostart task re-registration).
- Timeout diagnosability: partial installer output is logged on kill so
  the next hang names its stage instead of dying silently.

Contract repairs and fresh installs stay interactive-only (SmartScreen /
first-time elevation legitimately need a human).
2026-08-25 21:59:52 -07:00
blunkjamie-dev 105cea64f7 fix(update): bound optional cua-driver refresh 2026-08-25 21:59:52 -07:00
Jack Lau 65fada0945 fix(update): bound the cua-driver installer drain after a failed kill
On Windows, `hermes update` can hang past its own 660s cua-driver timeout
until the user kills an orphaned PowerShell by hand. The timeout ceiling is
not the problem; the code that runs after it is.

`_run_cua_driver_installer` handles `TimeoutExpired` by killing the process
tree and then draining the pipes with a bare `proc.communicate()`. The kill
is best-effort by construction: every `psutil.Error` in `_kill_installer_tree`
is logged at debug level and stepped over, on the reasoning that a partly
killed tree beats none. That is the right call, but it means the drain has to
survive a partial kill, and an unbounded drain does not.

The concrete case is the one reported. `install.ps1` self-elevates through
`Start-Process -Verb RunAs`, so the descendant runs at High integrity and a
medium-integrity `child.kill()` raises `AccessDenied`. The per-child handler
logs it and continues. That survivor is still holding the `stdout=PIPE` write
handle it inherited, so the following `communicate()` waits for an EOF that
arrives only when somebody kills that process manually. A bounded 660s wait
becomes an unbounded one, after the warning has already printed.

Bound the drain instead. A kill that landed closes the pipe immediately, so
this costs nothing on the normal path; a kill that did not costs 15s rather
than forever. The original `TimeoutExpired` is re-raised either way, so the
existing manual re-run hint still prints and the update unwinds. Losing the
tail of a timed-out installer's log is the cheaper half of that trade, and it
is only lost in the case where the run already failed.

The drain deliberately does not close the pipe handles. `communicate()`'s
reader threads are still blocked on them and closing underneath them races;
they are daemon threads, so abandoning them does not hold the interpreter
open.

Both timeout handlers (streaming and captured) now go through one helper.
The streaming child inherits the console rather than a pipe, so it is much
harder to stall there, but the two branches should not drift on a rule this
small.

Tests: 5, in a new `TestInstallerTimeoutDrainIsBounded`. Two fail without the
fix, including the reported scenario end to end (a child kill refused with
`psutil.AccessDenied`, asserting the drain still carries a deadline). The
deadline is asserted as a kwarg rather than by timing, because a test that
proved the hang by hanging would be the same defect wearing a test's name.

Scope note: this does not touch the `stdin` inheritance that lets
`install.ps1`'s `Read-Host` block in the first place. That is #79684 and open
PR #79871 already carries the one-line `stdin=DEVNULL` fix; the two are
independent and neither subsumes the other, since `DEVNULL` cannot unblock a
UAC elevation dialog.

Fixes #87703
2026-08-25 21:59:52 -07:00
Kudakwashe Paradzayi 76faeb969b test: require Bot Chat preview under a live writer
The hang bound alone would pass if read-only open degraded to None.
2026-08-25 21:59:38 -07:00
Kudakwashe Paradzayi 458f26c0dc fix(desktop): stop the Bots roster hanging on a live profile
profiles.list opened every profile state.db as a writable SessionDB,
which waits out write-lock patience while that profile's backend is
mid-turn. The desktop RPC timed out and Bot Mode's infinite React
Query retry kept the sidebar on a spinner.

Inspect those DBs read-only and bound roster retries so names still
paint.
2026-08-25 21:59:38 -07:00
Teknium 6cb3a26353 fix(doctor): classify certificate-anchored DRs as stable + keep explicit env_map hermetic
Two follow-ups on top of the #86391 salvage:
- check_macos_tcc_grants: a certificate-anchored DR (hermes desktop
  --setup-tcc-identity, or a notarized release) now reports as stable in its
  own class instead of falling into the identifier-pinned message; the
  identifier-pinned message points at --setup-tcc-identity for the strongest
  anchor.
- collect_relay_plugin_cutover_findings: only merge process-level env vars
  when env_map is None (run_doctor's live path). An explicit env_map is a
  complete environment description — merging os.environ on top made
  report_deprecated_config_and_env non-hermetic on boxes exporting legacy
  relay vars (10 findings vs the expected 2 in
  test_report_does_not_count_as_blocking_issue).
2026-08-25 21:59:36 -07:00
David Metcalfe 2d37ed056e fix(macos): harden TCC check against codesign timeouts, clarify scope
Review feedback (AI review on #86391):
- guard _macos_desktop_dr subprocess.run against TimeoutExpired/FileNotFoundError
  so a hanging codesign degrades to the unreadable-DR warning, never crashing
  the doctor run (matches the file's existing subprocess guard pattern)
- select the desktop bundle by newest-mtime across release/mac-*/Hermes.app,
  matching _desktop_packaged_executable, instead of a fixed arch order
- note the cdhash-match proxy assumption at the classification site
- document why /Applications/Hermes.app (Hermes-Setup launcher,
  com.nousresearch.hermes.setup, certificate-anchored) is deliberately not probed
- extend the repair hint to cover per-service resets
- regression tests: codesign timeout and missing-codesign paths
2026-08-25 21:59:36 -07:00
David Metcalfe 8665b1e4db fix(macos): guard empty DR in TCC check, tighten platform-guard test
GPT-OSS review: an empty codesign output would fall through to the
'stable identity' branch and false-positive. Guard with  and
cover the empty-string case. Flash review: the non-macOS silence test
mocked the bundle to None, so it never exercised the platform guard;
mock a real path instead.
2026-08-25 21:59:36 -07:00
David Metcalfe 36c1755065 fix(macos): detect stale TCC grants and guide one-time re-grant
TCC keys permission grants to the app's code-signing requirement. Grants
made to pre-#73681 builds carry a cdhash-pinned requirement that no
longer matches the rebuilt bundle, so macOS re-prompts on every capture
even though the System Settings toggle shows ON — and the modern prompt
has no Allow button, so users cannot complete the one-time re-grant.

- hermes doctor: new check_macos_tcc_grants() reports the desktop
  bundle's DR class (cdhash-pinned → grants reset on every update;
  identifier-pinned → stable) and prints the exact stale-grant repair
  (tccutil reset, toggle ON, fully quit & relaunch).
- hermes update: after a successful update on macOS with a desktop app
  installed, print the one-line stale-grant guidance.
- docs: desktop.md no longer claims grants persist 'out of the box';
  documents the one-time re-grant for pre-fix grants.

Closes #86385
2026-08-25 21:59:36 -07:00
Teknium afee35700e fix(delegation): suppress subagent-owned process notifications in parent chat by default
Background processes started by subagents (task_id sa-*) route their
notify_on_complete / watch_pattern notifications to the parent
conversation (b95ec1cb5) because anything outliving the child needs a
durable consumer. In practice these 'npm ci finished' walls are noise
mid-conversation — the child's consolidated delegation result is the
deliverable.

- New config key delegation.surface_child_process_notifications
  (default false = suppress). Flag true restores the previous behavior
  exactly (delivery with subagent attribution line).
- drain_notifications drops (never requeues) completion/watch_match/
  watch_disabled events whose task_id starts with 'sa-' when the flag
  is false, logging at debug with session_id+task_id for diagnosis.
  Requeueing would pin them forever — children never drain notifies.
- async_delegation events are NEVER suppressed (they ARE the result).
- watch_disabled emitters now carry task_id so sa- sessions' safety
  events follow the same suppression as their other events.
- Config read errors fall back to the default (suppress) and never
  crash the drain loop.
- Docs: delegation.md + configuration.md.
2026-08-25 21:59:04 -07:00
Teknium f751a8c546 fix(update): also defer the missing-binary CUA install on Windows
Follow-up to the salvaged #94296: the two guards covered the repair and
confirmed-update branches, but when cua-driver is enabled yet not
installed at all, control still reached _run_cua_driver_installer() and
an automatic 'hermes update' would launch the interactive install.ps1
anyway. Add the same defer before the installer run, keep POSIX
behavior unchanged, and give the confirmed-update message a natural
fallback when latest_version is unknown.
2026-08-25 16:53:01 -07:00
Royalaid 0c23bf19af fix(update): defer interactive CUA installs on Windows 2026-08-25 16:53:01 -07:00