Commit Graph

28271 Commits

Author SHA1 Message Date
Ben Barclay 49757d5e39 fix(telemetry): address review findings on the shared-metrics sender
Independent review found the claim mechanism did not work. Reproduced
against the real store: two senders POSTed the same package.

The claim wrote next_attempt_at = now, but selection requires
next_attempt_at <= now, so a concurrent pass matched the same row
immediately. It now writes a LEASE INTO THE FUTURE
(_CLAIM_LEASE_SECONDS), which is what actually excludes another pass,
and expires by itself if a process dies mid-send. _mark is additionally
guarded on send_state so a straggler whose lease lapsed cannot
overwrite a completed send back to pending.

The old concurrency test could not fail: it raised AssertionError from
inside a transport, and _send_one catches every exception as a
retryable transport error. It now records what the second pass saw.

Also from review:

- shutdown() never joined the send thread; the join was only wired into
  deactivate(). A short-lived CLI therefore killed an in-flight send at
  exit, on the only cadence this feature has.
- Removed HERMES_TELEMETRY_ENDPOINT. AGENTS.md reserves HERMES_* for
  secrets, and a behavioural override here was a consent hazard: an
  inherited variable could silently redirect telemetry a user agreed to
  send to Nous. The staging E2E writes the endpoint into its throwaway
  profile instead, which also exercises the real config path.
- Added the  shared-metrics toggle that AGENTS.md requires
  as the third opt-in surface, delegating to the setup prompt so the
  consent rules stay in one place.
- Non-429 4xx (401/403/404/413/422) are now permanent. Only 400 was,
  so a wrong path or oversized body retried every 15 minutes for 30
  days until retention pruned it.
- The opt-in day is stamped when the user consents, not on the first
  send pass, which silently dropped the opt-in day whenever the next
  export crossed midnight UTC.
- gzip now uses mtime=0. The embedded timestamp made two sends of one
  package differ on the wire, so the 'byte-identical retry' E2E was
  comparing parsed bodies and could not have caught it. It now compares
  raw request bytes.
- Reconciled the three stale claims in relay-shared-metrics.md that
  said no remote-delivery path exists.

233 tests pass (was 213). Staging E2E re-run through the config path:
both packages 202, and the service logged both objects written to S3.
2026-08-26 16:31:52 +10:00
David Metcalfe c0b5a8e15d fix(desktop): return True when fallback sign + strict verification succeed
The legacy ad-hoc fallback signed and verified successfully but still
fell through to return False, contradicting the fixup's documented
contract. The success witness codified the contradiction. Return True
on the verified success path; the caller ignores the return value, so
no behavior change beyond the contract correction.
2026-08-25 23:23:11 -07:00
David Metcalfe 177688e31e fix(desktop): never delete safeStorage keychain item in the updater
Addresses round-2 review feedback on #90961. The previous commits
scoped the keychain deletion to the legacy ad-hoc fallback, but the
reviewer correctly held the blocker: the fallback ran codesign with
check=False, ignored the result, and unconditionally deleted 'Hermes
Safe Storage' — permanently orphaning gateway and native OAuth
credentials even when signing failed or a configured identity had
failed and routed into the fallback.

This commit removes the deletion entirely:
- _desktop_macos_reset_keychain_safe_storage is gone; no code path
  touches the keychain item anymore.
- The legacy fallback now checks the codesign result and runs
  codesign --verify --deep --strict; any failure leaves the item
  untouched and prints a warning.
- The keychain prompt after an ad-hoc re-sign is recoverable
  (Always Allow updates the ACL partition list and preserves the
  key); deletion is not. The durable proof-carrying migration
  belongs in Electron (safeStorage can read the old key) and is
  tracked as a follow-up.

Tests: 4 witnesses (stable path, default no-config success, fallback
failure, fallback success) all mutation-verified against both the
deletion regression and the ignored-codesign-result regression.
2026-08-25 23:23:11 -07:00
David Metcalfe 368ea2d88b test(desktop): mark keychain-reset scoping tests macos_only
The fixup no-ops on non-macOS (sys.platform guard), so the new
regression tests must carry the same @pytest.mark.macos_only marker
as their siblings (test_relaunchable_fixup_falls_back_to_legacy_adhoc_on_failure).
Without it the legacy-adhoc test failed on the Linux CI runner where
the fixup returns True before reaching the reset path.
2026-08-25 23:23:11 -07:00
David Metcalfe 91dcca9a9b fix(desktop): scope keychain reset to the legacy ad-hoc fallback only
The previous commit deleted the 'Hermes Safe Storage' keychain item after
every successful re-sign, including the stable certificate-anchored
identity path. On that path the designated requirement is stable across
rebuilds, so after the first launch under the new identity the keychain
ACL already matches; deleting the item on every update permanently
orphaned gateway-token and native-OAuth credentials that were working
fine (both are safeStorage-backed: electron/main.ts connection config
and native-oauth-tokens.json).

Addresses review feedback on #90961:
- Rename _desktop_macos_update_keychain_acl -> _desktop_macos_reset_keychain_safe_storage (it deletes, it does not update an ACL).
- Only invoke it on the legacy ad-hoc fallback path, where every rebuild
  produces a new cdhash so the ACL can never match and the alternative
  is a recurring prompt. The trade-off (re-enter credentials once per
  update) is documented; the durable fix is a stable signing identity.
- Add regression tests: stable path must NOT reset, ad-hoc fallback MUST.
2026-08-25 23:23:11 -07:00
David Metcalfe d1c6d5bab6 fix(desktop): reset keychain entry after macOS re-sign to prevent prompt on every launch
The self-updater rebuilds the desktop app locally via electron-builder after
every update (4aa9f738ce).  On macOS, the rebuilt app gets ad-hoc signed,
producing a different cdhash than the original CI-signed build.  macOS ties
the 'Hermes Safe Storage' keychain item's ACL to the code signature, so
the new signature doesn't match → macOS re-prompts for keychain access on
every launch.

After re-signing, delete the existing keychain item so Electron recreates
it with the correct ACL for the newly-signed app on next launch.  The
trade-off: previously encrypted tokens become unreadable (the user
re-enters the gateway token once), but the keychain prompt stops appearing
on every launch.

The delete-generic-password command doesn't require reading the secret
(no ACL check), so it runs without prompting.
2026-08-25 23:23:11 -07:00
hermes-seaeye[bot] bc943d5672 fmt(js): npm run fix on merge (#95299)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 06:20:59 +00:00
Teknium 71d804d005 fix(desktop): sort titlebar-overlay-width import before window-connection-route (lint) 2026-08-25 23:15:49 -07:00
Teknium a7e1410faf chore: map Kavyrocom attribution 2026-08-25 23:15:49 -07:00
Jeremy cab6c4f70b fix(desktop): fail closed when messaging DELETE cannot resolve an owner
A listed profile-less row in a multi-profile setup must not DELETE/archive against the primary backend. Unresolved ownership keeps the row, pins, and unread state and never calls the mutation.
2026-08-25 23:15:49 -07:00
Jeremy add24666d5 fix(desktop): prefer messaging/cron slice when restoring a dual-listed session 2026-08-25 23:15:49 -07:00
Jeremy 718654a9a2 fix(desktop): route messaging session DELETE/archive to owning profile 2026-08-25 23:15:49 -07:00
Mabolla 7944ad2168 fix(desktop): isolate SSH routing by window 2026-08-25 23:15:49 -07:00
Mabolla d77114d0d0 fix(desktop): align SSH terminal routing precedence 2026-08-25 23:15:49 -07:00
Mabolla a70e7b0ca7 fix(desktop): route terminals through registry SSH 2026-08-25 23:15:49 -07:00
Mabolla 24ab1acbe0 test(desktop): cover v2 registry SSH terminal scope 2026-08-25 23:15:49 -07:00
Mabolla d2c8727db5 fix(desktop): resolve active v2 registry SSH scope 2026-08-25 23:15:49 -07:00
Shaq Wu 120430be47 fix(desktop): retain routed socket through turn completion 2026-08-25 23:15:49 -07:00
chelsealong 6b3977b577 fix(desktop): decide image/file byte-upload by the session's own connection, not ambient
uploadComposerAttachment's remote/local decision (image.attach vs
image.attach_bytes) read $connection.get()?.mode — the window's ambient
connection — at all three call sites. Since the multi-connection
routing work landed, a session can be owned by a different, registered
connection than the ambient one (Bot Mode, the unified Sessions list):
the RPC itself already routes to that owner via requestForSessionProfile,
but the byte-vs-path decision did not, so a local-ambient window chatting
in a remote-owned session shipped a client-local composer-images path to
a backend that can't read it (#94640).

isSessionRemote() resolves the session's own owner route (falling back to
ambient only when no owner route is known, matching the existing RPC
routing behavior in session-states.ts) and replaces the three ambient
reads in use-prompt-actions/index.ts, session-tile-actions.ts, and
user-edit-composer.tsx.
2026-08-25 23:15:49 -07:00
Kavyrocom 76af834ddd fix(desktop): route ambient SSH profile deletion to remote 2026-08-25 23:15:49 -07:00
Teknium 2664599644 test(gateway): prove the profile_route_rejected sentinel is observable
Review follow-up for the salvaged #94848: the ProfileRouteRejected marker
looked write-only inside the primary handler. Document that
_handle_message's ingress gate reads the same marker to drop the message
fail-closed, and add a regression test showing the rejected route is
stamped once, dispatch falls back to the default home, and routing is not
re-run on redelivery.
2026-08-25 23:15:15 -07:00
GarrettGlass 2afed50863 fix(gateway): authorize routed messages in transport scope 2026-08-25 23:15:15 -07:00
GarrettGlass 9ab748abb9 fix(gateway): scope routed history before session lookup 2026-08-25 23:15:15 -07:00
Teknium 85bb6515c0 docs: document the no_overlay auto-detect and new Linux X11 default 2026-08-25 23:14:34 -07:00
YappLeCunt c6ae9325ec fix(computer-use): default the cursor overlay off on Linux X11
cua-driver maps the agent-cursor overlay as a fullscreen, always-on-top,
all-workspaces X11 window (save-unders composited). When a computer-use
session ends uncleanly — an agent interrupted mid-capture, a stale target
window, a driver error — that window can be left stuck above every app on
every workspace, wedging desktop input until the app is restarted. This
is the same failure class as the HUD's transparent always-on-top window on
Mutter/X11 (#83473), and it bit a real user: an interrupted capture froze
the desktop, the app had to be force-restarted, and the overlay window was
still mapped fullscreen afterwards.

The overlay is cosmetic (a tinted cursor sprite); the driver, captures,
and synthetic input all work without it. `_cua_no_overlay()` already
defaulted it off on macOS (idle CPU redraw loop, #28152/#47032) and
headless/WSL2 Linux; this extends the same auto-detect to X11 desktop
sessions, where raw X11 stacking has no compositor-owned surface to tear
down with the driver's connection. Wayland keeps the overlay: the
compositor owns the layer-surface lifecycle, so a dead driver cannot leave
a stuck top window.

Behavior contract unchanged: an explicit `computer_use.no_overlay: false`
still restores the cursor on any platform, and `true` forces it off.

Tests: X11 (DISPLAY set, no Wayland env) and XDG_SESSION_TYPE=x11
auto-detect off; Wayland keeps the overlay; explicit false overrides
auto-detection on X11.
2026-08-25 23:14:34 -07:00
ruangraung dce4abe917 fix(gateway): log secondary startup-reconnect handoff failures instead of dropping them
Review follow-up to the AI code-review pass on PR #92074: the bridge task's
handoff into _schedule_secondary_profile_reconnect was unguarded at both call
sites inside the parked coroutine. The scheduler touches live registries
(_profile_failed_platforms slot creation, background-task registration), so an
unexpected raise there would kill the parked task as an unretrieved-task
exception — logged only at GC time via "Task exception was never retrieved",
where no operator ever looks. A fix whose entire purpose is to stop a platform
dying silently should not contain its own silent-death path; both handoff sites
now wrap the scheduler call with logger.exception so the failure lands in
gateway.log with profile and platform context.

The early-exit branch (gateway already _running when the bridge starts) had the
identical exposure and is guarded the same way — same bug class, fixed together.
Regression test drives a handoff raise end-to-end through the real bridge task:
the await completes cleanly, the error is captured in gateway.run's logger, and
no adapter or failed-platform slot leaks behind the failed handoff.
2026-08-25 22:55:07 -07:00
ruangraung 96489f3c1b fix(gateway): schedule secondary-profile reconnect when initial adapter connect fails
When gateway.multiplex_profiles is active, a secondary profile whose platform
adapter fails its initial connect at startup was silently given up on: the
failure branches in _start_one_profile_adapters() logged and disconnected, but
never scheduled recovery. One unlucky connect window during a Telegram API
outage left the profile permanently silent until manual restart (~80 min in
the observed incident), while the mid-run fatal path already recovers via
_handle_profile_adapter_fatal_error() -> _schedule_secondary_profile_reconnect().
The same gap hit both failure shapes: a clean False return from
_connect_initial_adapter_with_timeout() and an exception escaping it.

Fix: call _schedule_secondary_profile_startup_reconnect() from both startup
failure branches after _safe_adapter_disconnect(). Because secondary adapters
are started mid-start(), before self._running flips True, the regular
scheduler's not-self._running guard would silently drop the request — so the
new bridge parks a background task until startup completes (or shutdown
begins) and then hands off to _schedule_secondary_profile_reconnect()
verbatim: backoff, fresh-adapter rebuild under the profile runtime scope,
slot dedupe, and shutdown cancellation all come from the existing path.
Non-retryable failures are dropped at scheduling time exactly as the regular
scheduler drops them, keeping duplicate-credential/auth-failed startups dead
instead of looping.

Fixes #92064
2026-08-25 22:55:07 -07:00
Ben Barclay 055d58ba33 test(telemetry): add the live staging E2E script
Sends real packages through the real sender to the real staging ingest
service and reports what came back. Uses a throwaway HERMES_HOME so an
operator's own telemetry state is never touched, and asserts the local
install_id did not cross the wire.

Kept as a script rather than a pytest case on purpose: it needs live
network and a deployed staging service, so it must not run in CI.
2026-08-26 15:55:06 +10:00
Teknium 682a34a7c6 test(desktop): isolate warm-cache mapping describe from prior gateway mock traffic
Follow-up for salvaged PR #93451: main grew branchStoredSession tests (#93444)
that drive the same hoisted requestGatewayForProfile/ForAgent mocks, so the
not-called assertions in the warm-cache describe saw stale recorded calls.
2026-08-25 22:51:31 -07:00
Teknium 1f78750280 chore: add contributor email mappings for salvaged PRs 2026-08-25 22:51:31 -07:00
Leonid Skorobogatyy 2a4460b58b fix(desktop): guard gateway-open resume against active fresh-draft transition (#68594)
When switching profiles, useRouteResume() could treat the target profile
gateway opening as a resume command for the old routed session before
React Router commits the /new pathname. The gatewayBecameOpen trigger
was not guarded by the existing freshDraftReady discriminator, unlike
the stuckOnRoutedSession guard on line 137 which already uses it.

Fix: add && !freshDraftReady to the gatewayBecameOpen condition in
shouldResume (line 143). This is consistent with the existing pattern
and preserves reconnect behavior (freshDraftReady=false during normal
reconnects).

Adds a regression test that simulates the exact race: profile switch
clears refs and sets freshDraftReady=true, gateway closes, then profile
B gateway opens while pathname is still /session-a. Asserts no resume
fires.
2026-08-25 22:51:31 -07:00
echoes666 4b21f970aa fix(desktop): release profile activation lease on dial failure
Clear the profile-door activation lease in a finally block whether the
secondary dial succeeds or rejects. Preserve reconnect scheduling and the
original error propagation.

Add a regression proving a rejected activation can be pruned immediately.
2026-08-25 22:51:31 -07:00
echoes666 c06d3c0bd9 fix(desktop): surface profile-switch dial failures instead of silently activating
Rebase of #81165 onto post-#87600 main, slimmed to the error-surfacing UX
half as requested - the silent-misroute half already landed via #87600.

- ensureGatewayForProfile: keep scheduleReconnect() on a failed secondary
  dial (transient failures still self-heal) but rethrow so the caller can
  surface the failure instead of falling through to setActive() with a
  closed socket (#81094). openSecondary logs the dial target and rethrows
  the ORIGINAL error so reconnectSecondary's message-based fail-stop
  classification ("No connection with id", "no longer exists") keeps working.
- profile.ensureGatewayProfile: propagate the rejection to await-callers
  (session actions, slash commands surface it in their own flows); the three
  fire-and-forget call sites (selectProfile, newSessionInProfile, voice
  wiring) get explicit .catch(notifyError) so the failure is always visible.
- Tests: gateway.test.ts pins rethrow-without-activation plus backoff
  self-heal; profile-switch-failure.test.ts pins the profile-door
  propagation; gateway-shared-remote's pooled-reconnect test updated to
  the new contract (reject first, publish once the backend returns).

Verified: tsc --noEmit clean; 65 tests across the gateway*/profile*
suites pass; eslint and prettier clean on touched files.
2026-08-25 22:51:31 -07:00
cmoiccool f3ae1a0c3a refactor(desktop): use shared local connection id 2026-08-25 22:51:31 -07:00
cmoiccool 7d6bf56c7b fix(desktop): preserve profile overrides on local source picks
Treat the explicit local registry source like the legacy profile-only path so per-profile remote overrides still resolve. Keep non-local registry sources connection-scoped and cover both profile selection entry points.
2026-08-25 22:51:31 -07:00
Tom Reinelt 613822afff fix(desktop): invalidate stale runtimes before profile reopen
Remember secondary scopes across renderer pool pruning and clear their process-local runtime bindings before a later socket can publish open. This prevents eager session RPCs from racing a respawned profile backend with an ID minted by the previous process.
2026-08-25 22:51:31 -07:00
Luke Roberts 893c8b1fdd fix(desktop): preserve remote session routing 2026-08-25 22:51:31 -07:00
arya abe584288b fix(desktop): preserve registry owner for session sends 2026-08-25 22:51:31 -07:00
arya 26932ea0bf fix(desktop): retain registry session ownership 2026-08-25 22:51:31 -07:00
arya 1ec32e7385 fix(desktop): reuse primary during owned route activation 2026-08-25 22:51:31 -07:00
arya 6c0ddecf66 fix(desktop): reuse registry primary for owned session RPCs 2026-08-25 22:51:31 -07:00
Zeus-Deus 502298ef62 fix(desktop): preserve recovery refresh ownership 2026-08-25 22:51:31 -07:00
Zeus-Deus 693dd5d042 fix(desktop): guard stale refresh ownership 2026-08-25 22:51:31 -07:00
Zeus-Deus b3bdf0c816 fix(desktop): guard async switch publications 2026-08-25 22:51:31 -07:00
Zeus-Deus ee3dc554f5 fix(desktop): enforce gateway switch publication ownership 2026-08-25 22:51:31 -07:00
Zeus-Deus f23b3c805c fix(desktop): preserve switch loading ownership 2026-08-25 22:51:31 -07:00
Zeus-Deus 559a56c360 fix(desktop): preserve gateway switch recovery ownership 2026-08-25 22:51:31 -07:00
Zeus-Deus c2bce09d48 fix(desktop): recover failed gateway switch setup 2026-08-25 22:51:31 -07:00
Zeus-Deus b687c6b2c1 fix(desktop): cancel timed-out gateway activations 2026-08-25 22:51:31 -07:00
Zeus-Deus 300fcdbfbd fix(desktop): harden gateway-switch commit ordering for stalled and overlapping switches (#93937)
Review follow-up. The two-phase switch closed the runtime-id leak for a
single switch, but two shapes still broke the wipe→publish ordering:

Overlapping: click A commits (wipes) and waits for its activation behind
the profile-store mutex; click B finishes its dial and wipes too; then A's
activation lands and publishes A AFTER B's destructive wipe, and A's
endGatewaySwitch() dropped the (boolean) barrier while B was mid-commit.
Now the wipe runs INSIDE the serialized section via a `beforeActivate`
commit hook on ensureGatewayAgent, synchronously right before the socket
is activated; the hook re-checks the click revision and declines when
superseded, so a queued-then-superseded switch neither wipes nor
activates. The barrier is token-owned (latest switch wins), and the mutex
wait is a loop so waiters that wake together can't run interleaved.

Stalled: a wedged spawn / ticket mint / IPC (the #93454 class) left the
spinner up, swallowed later clicks on the same source, and — inside
ensureGatewayAgent — latched the mutex and the barrier. Every await is now
bounded (dial, activation, descriptor lookups, setLastUsed) via a shared
withTimeout helper extracted from use-gateway-boot. A commit that timed
out after the new source was already published counts as committed; one
that stalls after the wipe lowers the barrier and repaints the source
that is still active.

Tests: queued-then-superseded switch, dial that never answers (and retry
not swallowed), activation stalled after/before publication, barrier
ownership, commit-hook ordering inside the mutex, bounded descriptor
lookup releasing the mutex. All fail against the previous revision.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 22:51:31 -07:00