Independent review found the claim mechanism did not work. Reproduced
against the real store: two senders POSTed the same package.
The claim wrote next_attempt_at = now, but selection requires
next_attempt_at <= now, so a concurrent pass matched the same row
immediately. It now writes a LEASE INTO THE FUTURE
(_CLAIM_LEASE_SECONDS), which is what actually excludes another pass,
and expires by itself if a process dies mid-send. _mark is additionally
guarded on send_state so a straggler whose lease lapsed cannot
overwrite a completed send back to pending.
The old concurrency test could not fail: it raised AssertionError from
inside a transport, and _send_one catches every exception as a
retryable transport error. It now records what the second pass saw.
Also from review:
- shutdown() never joined the send thread; the join was only wired into
deactivate(). A short-lived CLI therefore killed an in-flight send at
exit, on the only cadence this feature has.
- Removed HERMES_TELEMETRY_ENDPOINT. AGENTS.md reserves HERMES_* for
secrets, and a behavioural override here was a consent hazard: an
inherited variable could silently redirect telemetry a user agreed to
send to Nous. The staging E2E writes the endpoint into its throwaway
profile instead, which also exercises the real config path.
- Added the shared-metrics toggle that AGENTS.md requires
as the third opt-in surface, delegating to the setup prompt so the
consent rules stay in one place.
- Non-429 4xx (401/403/404/413/422) are now permanent. Only 400 was,
so a wrong path or oversized body retried every 15 minutes for 30
days until retention pruned it.
- The opt-in day is stamped when the user consents, not on the first
send pass, which silently dropped the opt-in day whenever the next
export crossed midnight UTC.
- gzip now uses mtime=0. The embedded timestamp made two sends of one
package differ on the wire, so the 'byte-identical retry' E2E was
comparing parsed bodies and could not have caught it. It now compares
raw request bytes.
- Reconciled the three stale claims in relay-shared-metrics.md that
said no remote-delivery path exists.
233 tests pass (was 213). Staging E2E re-run through the config path:
both packages 202, and the service logged both objects written to S3.
The legacy ad-hoc fallback signed and verified successfully but still
fell through to return False, contradicting the fixup's documented
contract. The success witness codified the contradiction. Return True
on the verified success path; the caller ignores the return value, so
no behavior change beyond the contract correction.
Addresses round-2 review feedback on #90961. The previous commits
scoped the keychain deletion to the legacy ad-hoc fallback, but the
reviewer correctly held the blocker: the fallback ran codesign with
check=False, ignored the result, and unconditionally deleted 'Hermes
Safe Storage' — permanently orphaning gateway and native OAuth
credentials even when signing failed or a configured identity had
failed and routed into the fallback.
This commit removes the deletion entirely:
- _desktop_macos_reset_keychain_safe_storage is gone; no code path
touches the keychain item anymore.
- The legacy fallback now checks the codesign result and runs
codesign --verify --deep --strict; any failure leaves the item
untouched and prints a warning.
- The keychain prompt after an ad-hoc re-sign is recoverable
(Always Allow updates the ACL partition list and preserves the
key); deletion is not. The durable proof-carrying migration
belongs in Electron (safeStorage can read the old key) and is
tracked as a follow-up.
Tests: 4 witnesses (stable path, default no-config success, fallback
failure, fallback success) all mutation-verified against both the
deletion regression and the ignored-codesign-result regression.
The fixup no-ops on non-macOS (sys.platform guard), so the new
regression tests must carry the same @pytest.mark.macos_only marker
as their siblings (test_relaunchable_fixup_falls_back_to_legacy_adhoc_on_failure).
Without it the legacy-adhoc test failed on the Linux CI runner where
the fixup returns True before reaching the reset path.
The previous commit deleted the 'Hermes Safe Storage' keychain item after
every successful re-sign, including the stable certificate-anchored
identity path. On that path the designated requirement is stable across
rebuilds, so after the first launch under the new identity the keychain
ACL already matches; deleting the item on every update permanently
orphaned gateway-token and native-OAuth credentials that were working
fine (both are safeStorage-backed: electron/main.ts connection config
and native-oauth-tokens.json).
Addresses review feedback on #90961:
- Rename _desktop_macos_update_keychain_acl -> _desktop_macos_reset_keychain_safe_storage (it deletes, it does not update an ACL).
- Only invoke it on the legacy ad-hoc fallback path, where every rebuild
produces a new cdhash so the ACL can never match and the alternative
is a recurring prompt. The trade-off (re-enter credentials once per
update) is documented; the durable fix is a stable signing identity.
- Add regression tests: stable path must NOT reset, ad-hoc fallback MUST.
The self-updater rebuilds the desktop app locally via electron-builder after
every update (4aa9f738ce). On macOS, the rebuilt app gets ad-hoc signed,
producing a different cdhash than the original CI-signed build. macOS ties
the 'Hermes Safe Storage' keychain item's ACL to the code signature, so
the new signature doesn't match → macOS re-prompts for keychain access on
every launch.
After re-signing, delete the existing keychain item so Electron recreates
it with the correct ACL for the newly-signed app on next launch. The
trade-off: previously encrypted tokens become unreadable (the user
re-enters the gateway token once), but the keychain prompt stops appearing
on every launch.
The delete-generic-password command doesn't require reading the secret
(no ACL check), so it runs without prompting.
A listed profile-less row in a multi-profile setup must not DELETE/archive against the primary backend. Unresolved ownership keeps the row, pins, and unread state and never calls the mutation.
uploadComposerAttachment's remote/local decision (image.attach vs
image.attach_bytes) read $connection.get()?.mode — the window's ambient
connection — at all three call sites. Since the multi-connection
routing work landed, a session can be owned by a different, registered
connection than the ambient one (Bot Mode, the unified Sessions list):
the RPC itself already routes to that owner via requestForSessionProfile,
but the byte-vs-path decision did not, so a local-ambient window chatting
in a remote-owned session shipped a client-local composer-images path to
a backend that can't read it (#94640).
isSessionRemote() resolves the session's own owner route (falling back to
ambient only when no owner route is known, matching the existing RPC
routing behavior in session-states.ts) and replaces the three ambient
reads in use-prompt-actions/index.ts, session-tile-actions.ts, and
user-edit-composer.tsx.
Review follow-up for the salvaged #94848: the ProfileRouteRejected marker
looked write-only inside the primary handler. Document that
_handle_message's ingress gate reads the same marker to drop the message
fail-closed, and add a regression test showing the rejected route is
stamped once, dispatch falls back to the default home, and routing is not
re-run on redelivery.
cua-driver maps the agent-cursor overlay as a fullscreen, always-on-top,
all-workspaces X11 window (save-unders composited). When a computer-use
session ends uncleanly — an agent interrupted mid-capture, a stale target
window, a driver error — that window can be left stuck above every app on
every workspace, wedging desktop input until the app is restarted. This
is the same failure class as the HUD's transparent always-on-top window on
Mutter/X11 (#83473), and it bit a real user: an interrupted capture froze
the desktop, the app had to be force-restarted, and the overlay window was
still mapped fullscreen afterwards.
The overlay is cosmetic (a tinted cursor sprite); the driver, captures,
and synthetic input all work without it. `_cua_no_overlay()` already
defaulted it off on macOS (idle CPU redraw loop, #28152/#47032) and
headless/WSL2 Linux; this extends the same auto-detect to X11 desktop
sessions, where raw X11 stacking has no compositor-owned surface to tear
down with the driver's connection. Wayland keeps the overlay: the
compositor owns the layer-surface lifecycle, so a dead driver cannot leave
a stuck top window.
Behavior contract unchanged: an explicit `computer_use.no_overlay: false`
still restores the cursor on any platform, and `true` forces it off.
Tests: X11 (DISPLAY set, no Wayland env) and XDG_SESSION_TYPE=x11
auto-detect off; Wayland keeps the overlay; explicit false overrides
auto-detection on X11.
Review follow-up to the AI code-review pass on PR #92074: the bridge task's
handoff into _schedule_secondary_profile_reconnect was unguarded at both call
sites inside the parked coroutine. The scheduler touches live registries
(_profile_failed_platforms slot creation, background-task registration), so an
unexpected raise there would kill the parked task as an unretrieved-task
exception — logged only at GC time via "Task exception was never retrieved",
where no operator ever looks. A fix whose entire purpose is to stop a platform
dying silently should not contain its own silent-death path; both handoff sites
now wrap the scheduler call with logger.exception so the failure lands in
gateway.log with profile and platform context.
The early-exit branch (gateway already _running when the bridge starts) had the
identical exposure and is guarded the same way — same bug class, fixed together.
Regression test drives a handoff raise end-to-end through the real bridge task:
the await completes cleanly, the error is captured in gateway.run's logger, and
no adapter or failed-platform slot leaks behind the failed handoff.
When gateway.multiplex_profiles is active, a secondary profile whose platform
adapter fails its initial connect at startup was silently given up on: the
failure branches in _start_one_profile_adapters() logged and disconnected, but
never scheduled recovery. One unlucky connect window during a Telegram API
outage left the profile permanently silent until manual restart (~80 min in
the observed incident), while the mid-run fatal path already recovers via
_handle_profile_adapter_fatal_error() -> _schedule_secondary_profile_reconnect().
The same gap hit both failure shapes: a clean False return from
_connect_initial_adapter_with_timeout() and an exception escaping it.
Fix: call _schedule_secondary_profile_startup_reconnect() from both startup
failure branches after _safe_adapter_disconnect(). Because secondary adapters
are started mid-start(), before self._running flips True, the regular
scheduler's not-self._running guard would silently drop the request — so the
new bridge parks a background task until startup completes (or shutdown
begins) and then hands off to _schedule_secondary_profile_reconnect()
verbatim: backoff, fresh-adapter rebuild under the profile runtime scope,
slot dedupe, and shutdown cancellation all come from the existing path.
Non-retryable failures are dropped at scheduling time exactly as the regular
scheduler drops them, keeping duplicate-credential/auth-failed startups dead
instead of looping.
Fixes#92064
Sends real packages through the real sender to the real staging ingest
service and reports what came back. Uses a throwaway HERMES_HOME so an
operator's own telemetry state is never touched, and asserts the local
install_id did not cross the wire.
Kept as a script rather than a pytest case on purpose: it needs live
network and a deployed staging service, so it must not run in CI.
Follow-up for salvaged PR #93451: main grew branchStoredSession tests (#93444)
that drive the same hoisted requestGatewayForProfile/ForAgent mocks, so the
not-called assertions in the warm-cache describe saw stale recorded calls.
When switching profiles, useRouteResume() could treat the target profile
gateway opening as a resume command for the old routed session before
React Router commits the /new pathname. The gatewayBecameOpen trigger
was not guarded by the existing freshDraftReady discriminator, unlike
the stuckOnRoutedSession guard on line 137 which already uses it.
Fix: add && !freshDraftReady to the gatewayBecameOpen condition in
shouldResume (line 143). This is consistent with the existing pattern
and preserves reconnect behavior (freshDraftReady=false during normal
reconnects).
Adds a regression test that simulates the exact race: profile switch
clears refs and sets freshDraftReady=true, gateway closes, then profile
B gateway opens while pathname is still /session-a. Asserts no resume
fires.
Clear the profile-door activation lease in a finally block whether the
secondary dial succeeds or rejects. Preserve reconnect scheduling and the
original error propagation.
Add a regression proving a rejected activation can be pruned immediately.
Rebase of #81165 onto post-#87600 main, slimmed to the error-surfacing UX
half as requested - the silent-misroute half already landed via #87600.
- ensureGatewayForProfile: keep scheduleReconnect() on a failed secondary
dial (transient failures still self-heal) but rethrow so the caller can
surface the failure instead of falling through to setActive() with a
closed socket (#81094). openSecondary logs the dial target and rethrows
the ORIGINAL error so reconnectSecondary's message-based fail-stop
classification ("No connection with id", "no longer exists") keeps working.
- profile.ensureGatewayProfile: propagate the rejection to await-callers
(session actions, slash commands surface it in their own flows); the three
fire-and-forget call sites (selectProfile, newSessionInProfile, voice
wiring) get explicit .catch(notifyError) so the failure is always visible.
- Tests: gateway.test.ts pins rethrow-without-activation plus backoff
self-heal; profile-switch-failure.test.ts pins the profile-door
propagation; gateway-shared-remote's pooled-reconnect test updated to
the new contract (reject first, publish once the backend returns).
Verified: tsc --noEmit clean; 65 tests across the gateway*/profile*
suites pass; eslint and prettier clean on touched files.
Treat the explicit local registry source like the legacy profile-only path so per-profile remote overrides still resolve. Keep non-local registry sources connection-scoped and cover both profile selection entry points.
Remember secondary scopes across renderer pool pruning and clear their process-local runtime bindings before a later socket can publish open. This prevents eager session RPCs from racing a respawned profile backend with an ID minted by the previous process.
Review follow-up. The two-phase switch closed the runtime-id leak for a
single switch, but two shapes still broke the wipe→publish ordering:
Overlapping: click A commits (wipes) and waits for its activation behind
the profile-store mutex; click B finishes its dial and wipes too; then A's
activation lands and publishes A AFTER B's destructive wipe, and A's
endGatewaySwitch() dropped the (boolean) barrier while B was mid-commit.
Now the wipe runs INSIDE the serialized section via a `beforeActivate`
commit hook on ensureGatewayAgent, synchronously right before the socket
is activated; the hook re-checks the click revision and declines when
superseded, so a queued-then-superseded switch neither wipes nor
activates. The barrier is token-owned (latest switch wins), and the mutex
wait is a loop so waiters that wake together can't run interleaved.
Stalled: a wedged spawn / ticket mint / IPC (the #93454 class) left the
spinner up, swallowed later clicks on the same source, and — inside
ensureGatewayAgent — latched the mutex and the barrier. Every await is now
bounded (dial, activation, descriptor lookups, setLastUsed) via a shared
withTimeout helper extracted from use-gateway-boot. A commit that timed
out after the new source was already published counts as committed; one
that stalls after the wipe lowers the barrier and repaints the source
that is still active.
Tests: queued-then-superseded switch, dial that never answers (and retry
not swallowed), activation stalled after/before publication, barrier
ownership, commit-hook ordering inside the mutex, bounded descriptor
lookup releasing the mutex. All fail against the previous revision.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>