Commit Graph

4143 Commits

Author SHA1 Message Date
Deus a0578f4ef3 test(desktop): model registry topology in dispatcher coverage 2026-08-26 03:22:37 -07:00
Deus f912657602 test(desktop): prove fresh chat socket continuity 2026-08-26 03:22:37 -07:00
Deus c662e1be7b fix(desktop): retain primary gateway registry identity after reload 2026-08-26 03:22:37 -07:00
Deus 6fdf873464 fix(desktop): serialize profile switches and drain edit redials 2026-08-26 03:22:37 -07:00
Deus 962b308be1 fix(desktop): require exact owners in registry topology 2026-08-26 03:22:37 -07:00
Zeus-Deus 07b87f1470 fix(desktop): profile-rail fresh chats keep one exact session owner from create through every later RPC
selectProfile(name) / newSessionInProfile(name) keep only $newChatProfile and
clear $newChatRoute. #94147 taught the send path to capture the (registry
source, profile) pair as the draft's exact owner, but three gaps still let a
session created on the composite gateway conn:local::omar degrade to the bare
string "omar" — and requestGatewayForProfile("omar") is a DIFFERENT socket
than the one that minted the runtime, so the next session-scoped RPC 4001'd
"session not found" while the runtime was ws-orphan-reaped:

1. Ownership persistence. The exact owner lived only in the bounded,
   in-memory owner-hint map. The primary aggregate serves a `local` registry
   source's rows WITHOUT connection_id (the unified-list splice tags non-local
   sources only), and mergeSessionPage replaced the optimistic row with that
   untagged row on the first sidebar refresh. After a hint eviction or a
   relaunch nothing exact was left.
   - One canonical SessionOwnerRoute type (store/session-request-router);
     AgentProfileRoute / SessionProfileRoute / SessionRpcOwnerRoute alias it.
   - Every owner ladder gains the connection-tagged ROW rung
     (knownSessionOwner / sessionOwnerRouteFromRow): the session-RPC
     dispatcher, knownOwnerForSession, the tile delegate, foregroundSessionScopes
     and the async probe (resolveSessionOwner) all yield the exact route when
     the row carries its connection.
   - mergeSessionPage carries connection_id onto a row that comes back
     untagged for the same profile (merge, don't clobber).
   - Owner hints are persisted (bounded LRU, hermes.desktop.sessionOwnerHints.v1)
     and rehydrated in LRU order; a removed registry connection drops its
     hints (use-gateway-boot onChanged) so fail-closed can't pin sessions to a
     dead source.

2. Fail closed. A request carrying session_id whose owner no rung could name
   silently fell to the ambient presentation gateway, turning missing metadata
   into a misleading backend "session not found". createSessionRpcDispatcher
   and requestForOwnedSession now reject with an explicit
   SessionOwnerResolutionError (store/session-owner-resolution). The ONE case
   where ambient is the owner by construction stays ambient: no registry
   source live AND at most one profile (legacy single-backend Desktop, whose
   older backends omit `profile` on rows). Main-pane runtime ids (native
   approval.respond, queued sends) now translate to their stored id through
   the per-runtime state mirror (storedSessionIdForRuntimeId), so they resolve
   an owner instead of tripping the gate.

3. Lifecycle. Between session.create returning and the foreground publication
   ($selectedStoredSessionId via navigate → route effect, or $sessionTiles),
   the owner entry had no active request and was not yet foreground-pinned:
   a prune recompute or a refcount-0 lease release could close the socket
   holding the just-minted runtime before the first prompt.submit. Both create
   paths now hold retainGatewayForAgent across the create RPC and hand off to
   holdSessionOwnerUntilForeground (session-states), which names the owner in
   foregroundSessionScopes — every registry dispose path honors it — until the
   session becomes selected/tiled, the caller releases it (failed create,
   mid-create drift close), or a 60s TTL expires. Nothing latches.

Regressions:
- profile-rail-fresh-chat-owner.test.tsx: new case evicts the hint AND
  merges an untagged refresh row after turn one; turn two still rides the
  same conn:local::omar socket, no probe, no session.close, no v1 socket;
  the existing case also asserts the owner is foreground-pinned from create.
- session-rpc-dispatcher.test.ts (new): fail-closed error + ambient never
  called; legacy single-backend stays ambient; tagged-row rung routes by
  runtime id; hint outranks untagged row; probe result routes exactly.
- session.test.ts: persisted hints survive a simulated relaunch in LRU order,
  malformed storage is ignored, per-connection forget; knownSessionOwner;
  mergeSessionPage carries / drops / preserves identity on connection_id.
- session-states-foreground-scopes.test.ts: tagged-row scope for the selected
  thread; hold pins from create, retires on selection / tile mount, explicit
  release, TTL expiry.
- session-states-runtime-map.test.ts: state-mirror rung; knownOwnerForSession
  through mirror + hint / tagged row; requestForOwnedSession fail-closed vs
  legacy ambient.
- wiring-routing.test.ts: connection-tagged row rung; hint still outranks it.

Stacked on #94145, #94147 and #94178 (merged as the integration base);
review from this commit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 03:22:37 -07:00
Zeus-Deus cb75983abf fix(desktop): profile-rail fresh chats keep their registry source as the exact owner
The Sessions / profile-rail path (selectProfile, newSessionInProfile, a
connection switch, `/profile`) sets $newChatProfile and deliberately clears
$newChatRoute, so a fresh chat had no explicit owner. The session was created
on the active registry gateway (conn:local::omar) but its durable owner
degraded to the bare string "omar": follow-up RPCs dialed
requestGatewayForProfile("omar") — a different socket than the one that
minted the WebSocket-scoped runtime — and 4001'd "session not found" while
the runtime was ws-orphan-reaped.

- store/profile: capture the active registry source together with the
  new-chat profile intent ($newChatConnectionId / captureNewChatSource) in
  selectProfile, newSessionInProfile, newSessionInAgent, connection switches
  and `/profile`; resolveNewChatOwnerRoute() derives the exact
  { connectionId, profile } route whenever a registry source is live, even
  with $newChatRoute null (legacy v1 primary still yields null).
- use-session-actions: session.create, the owner hint, the optimistic row's
  profile + connection_id, and the failed-create cleanup all use that
  effective owner (main chat and tile paths).
- use-prompt-actions/submit: re-pin targetStoredSessionId after a fresh
  create. It was captured before the create (null) and seedOptimistic handed
  it to updateSessionState, which the state cache read as a DETACH — the
  fresh stored↔runtime binding was severed the moment the chat existed, so
  every later session-scoped RPC failed to translate the runtime id, never
  saw the tile route / owner hint / row, probed REST by runtime id and fell
  to the ambient socket.
- contrib: the session-RPC dispatcher is factored out of wiring.tsx
  (createSessionRpcDispatcher) so the exact production routing is what the
  integration test drives.

Regression (profile-rail-fresh-chat-owner.test.tsx) drives the real path:
mocked sockets under the real registry store, primary = remote default,
active source = local, selectProfile("omar") ($newChatProfile = "omar",
$newChatRoute = null), real useSessionStateCache / useSessionActions /
usePromptActions and the production dispatcher; asserts session.create and
BOTH prompt.submit calls hit the same conn:local::omar gateway object, no
session-scoped RPC reached the primary or a v1 "omar" socket, no
session.close, the binding survives both turns, no REST probe.

Verified with the packaged Linux Desktop against the real ~/.hermes
(primary = remote OAuth gateway, "This device" as registry source, omar via
the profile rail, two prompts): both prompts persisted on one session in
profiles/omar/state.db, no ws_orphan_reap.

Refs #94071

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 03:22:37 -07:00
Zeus-Deus 77edf1b4df fix(desktop): routed fresh chat keeps its exact owner after session.create
A fresh chat created through $newChatRoute lost its owner the moment
session.create returned. The create RPC rode the captured route
(requestGatewayForAgent), but the optimistic row was stamped from
$activeGatewayProfile — still `default` in All-profiles / Bot routing —
and no owner hint was recorded. The first turn ran on the routed
backend (e.g. local::omar); every later session-scoped RPC resolved the
row as `default` and 4001'd "session not found", leaving the routed
runtime to be ws-orphan-reaped.

Make the ownership transition atomic with the create:
- record capturedRoute as the stored session's exact owner hint the
  moment a routed create returns a stored id (main chat and tile paths);
- upsertOptimisticSession accepts an explicit owner and stamps
  profile = targetProfile || profile plus the owning connection_id,
  falling back to the ambient profile only for an unrouted create;
- contrib/wiring resolves session RPC owners as: persisted tile owner
  route → exact unique owner hint → session-row profile → cross-profile
  probe (resolveSessionRpcOwner, pure + unit-tested), so prompt.submit,
  session.resume, attachments, interrupt, redirect and recovery all use
  the exact owner; knownOwnerForSession follows the same ladder;
- the mid-create drift-abort session.close rides capturedRoute too.

Regressions: an integration test (ambient default, route local::omar,
two turns → both prompt.submit hit local::omar, no session-not-found,
no session.close) and a unit test asserting a routed fresh create never
yields { followupOwner: 'default', foregroundScope: 'conn:local::omar' }.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 95e2ef66becb641d3bdedf016dab6c104a777ae7)
2026-08-26 03:22:37 -07:00
Teknium 4faa721d7d fix(desktop): clicking a bot no longer burns a model turn on a fake user prompt
The intro kickoff ('Hey, tell me about yourself!') now fires ONLY from
genuine New Agent creation. The bot-click canonical resolution path mints
silently: the eager session.title write already persists the lazy row on
modern gateways, so the kickoff's session-persistence job is obsolete
there. A resolution miss (retitled row, hidden-listing gap, post-update
skew) previously re-fired the kickoff on EVERY click — a burned model
turn plus a user-attributed prompt the user never typed (ScottFive
report). Older gateways that reject the eager title keep a narrow compat
kickoff, else the pruner reaps the empty lazy session.
2026-08-26 00:52:54 -07:00
hermes-seaeye[bot] 1a19c52dd2 fmt(js): npm run fix on merge (#95365)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 07:45:59 +00:00
Teknium 25d46c7887 fix(desktop): drop unused registryGatewayWsUrl import left by the header-binding refactor rebase 2026-08-26 00:40:35 -07:00
Victor Nogueira 21f34794be fix(desktop): recover remote sessions after gateway restart 2026-08-26 00:40:35 -07:00
686f6c61 8085614f6a fix(desktop): log when Windows remote SSH skip teardown
POSIX disconnect does not apply to connectWindowsRemote. Quit still
closes the tunnel; log the skipped serve kill so it is not a silent
no-op.
2026-08-26 00:40:35 -07:00
686f6c61 0c69cac48a fix(desktop): terminate owned SSH serve backends on quit
teardownSshConnection closed the tunnel and SSH transport but never
killed the detached serve --isolated process. Spawn uses setsid/nohup,
so the backend reparents to pid 1, keeps state.db open, and accumulates
across Cmd+Q. Reuse cleanupStale via disconnect while SSH can still
exec, sequence remote kill before close, and seal the bootstrap
coordinator so reconnect during a prevented first quit cannot respawn.
The quit race is 6s to cover cleanupStale's 5s wait-for-exit loop.
2026-08-26 00:40:35 -07:00
Jaime Marques 14d16c2578 fix(desktop): clarify primary SSH reuse failures 2026-08-26 00:40:35 -07:00
Jaime Marques 3263ca2af6 fix(desktop): reuse migrated primary SSH backend
Avoid opening a second SSH lifecycle when a migrated registry request targets the same primary/default backend already booted through the legacy route. Compare effective SSH configuration for representation-only drift, treat empty and default as the same root profile, and keep named profiles isolated.
2026-08-26 00:40:35 -07:00
Ahmett101 1b8f3eda38 fix(desktop): scope registered ssh primary gateway 2026-08-26 00:40:35 -07:00
Ravi Tharuma 949f5169de fix(desktop): treat ticket 401 as sign-in when native tokens are unreadable 2026-08-26 00:40:35 -07:00
Marco Fernstaedt 024b9c0545 fix(desktop): refresh remote WebSocket header cache recency 2026-08-26 00:40:35 -07:00
Marco Fernstaedt b4162de333 fix(desktop): bind headers to scoped WebSocket URL 2026-08-26 00:40:35 -07:00
Teknium 5a285d3436 fix(desktop): restore stale-branch-reverted main.ts/update files; detach post-switch profile refresh from switch completion
The rebase re-landed pre-#74805 versions of the backend release gate,
venv-blocker rescan, mac entitlements/usage tests, and package.json from
the stale branch base — restored to main's versions (only the salvaged
enumeration/profileMetadata/profile:remember hunks kept in main.ts).
refreshActiveProfile's new bounded retry chain (#70679) is no longer
awaited inside the switch-completion barrier, so a slow/unhealthy backend
cannot hold $gatewaySwitching past the switch-ownership deadline; also
drop an unused $connection import from the earlier conflict compose.
2026-08-26 00:22:27 -07:00
Teknium b722177fcd fix(desktop): drop duplicate knownSessionOwner re-landed by rebase (main's richer variant wins) 2026-08-26 00:22:27 -07:00
tachi 317ae240fb fix(desktop): SSH/reconnect owner continuity, attachment routing, transport-error recovery
Salvage of #94192's unique work (owner-hardening portions that overlap
the class-1 branch — #94824/#93451 seams — and out-of-cluster #94864 are
intentionally excluded):

- use-gateway-request: recognize the full transport-error family
  (ECONNRESET & friends, including error.code and error.cause.code) so a
  reset SSH/remote socket triggers the connection-owned reconnect instead
  of surfacing as a request failure; background profiles keep the
  registry reconnect path for composite remote/SSH sources.
- session-tile-actions: tile attachment uploads and session RPCs follow
  the tile's composite owner (connectionId+profile) even when the active
  gateway moved to a same-named profile on another source.
- knownSessionOwner: sessions expose their complete owner (registry
  connection + profile) instead of a bare profile name that silently
  collapsed the route back to the local path; delegate/wiring resolve
  owners through it.

Fixes the SSH-reconnect share of #91365-adjacent routing gaps.
Salvaged (partial) from #94192.
2026-08-26 00:22:27 -07:00
Tom 2ed39365d6 fix(desktop): thread eager profile metadata through registry enumeration
Never-interacted remote bots painted as bare handles because roster rows
carried only profile names: display_name/title/ui_meta/has_avatar were
fetched lazily on first interaction (#91365). Thread credential-free
profile metadata from the enumeration-time /api/profiles body through
enumerateRegistryAgentSources (main.ts) and buildAgentRoster
(connection-registry.ts), keeping it attached to the connection-qualified
row across the same-install collapse. The plugin.js botRosterMeta half of
the original PR is dropped — superseded by landed #92731.

Fixes #91365
Salvaged (partial) from #92708.
2026-08-26 00:22:27 -07:00
Tilly-YL 2952119bce fix(desktop): remember selected profile across restarts
The profile rail's live workspace switch never persisted the selection,
so the Desktop always booted back into the previous startup profile
(#79886). Route the successful primary-backend activation through a new
persistence-only hermes:profile:remember IPC (validated
writeActiveDesktopProfile) that records the choice WITHOUT tearing down
the backend or reloading the window like hermes:profile:set does.
Registry-source picks name another source's profiles and do not touch
the startup preference. Reapplied semantically over three weeks of
main.ts/preload.ts drift (selectProfile now routes through
activateOnCurrentSource, #91349/#91365 seams).

Fixes #79886
Salvaged from #79888.
2026-08-26 00:22:27 -07:00
David Metcalfe 57043c2bc0 fix(desktop): single-flight refreshProfiles with retry recovery in global remote mode
Global remote mode fires refreshProfiles while the remote HTTP proxy is
still routing: the one-shot fetch failed silently and the rail stayed
empty until a manual refresh. Retry with 500ms/1000ms backoff, surface
terminal failures on the console, and dedupe concurrent callers into a
single retry chain (gateway open fires useBackgroundSync and the
activeGatewayProfile effect at once). Reapplied semantically on top of
the #85731 epoch guard: a stranded epoch stops the retry chain and
invalidation detaches the single-flight slot.

Fixes #70679
Salvaged from #74500.
2026-08-26 00:22:27 -07:00
chelsealong a928596758 fix(desktop): document connect-on-demand origin, add fallback-profiles integration test
Addresses AI-review feedback on #94653: note where the 'connect-on-demand'
sentinel is produced, and cover the interaction between
isLocalEnumerationFailure and localRouteFallbackProfiles directly (not just
the helper in isolation).
2026-08-26 00:22:27 -07:00
chelsealong c475484f63 fix(desktop): do not treat deferred local enumeration as a failure
'connect-on-demand' means local roster enumeration was intentionally
skipped to avoid spawning a local backend on a remote-only workspace,
not that it failed. The plugin-profile-routes IPC handler passed
Boolean(error) straight through, so that deferral was treated as a
genuine failure and Bot Mode re-synthesized cached local profile rows
even though local was never dialed.

Fixes #94648
2026-08-26 00:22:27 -07:00
Jeremy McKeehen 4ca1f532be test(desktop): pass the pin-write fence into the Show-all order assertion
resolvePinnedSessions requires unconfirmedPinWrites; the reconnection-scope test omitted it and would fail strict tsc.
2026-08-26 00:22:27 -07:00
Jeremy McKeehen 3751b04550 test(desktop): lock pin upgrade to server-authoritative pull
Old per-profile pin caches caused the stale unpin resurrection. Prove they are ignored and that sessions.pinned repopulates the gateway-wide key without a migration PATCH.
2026-08-26 00:22:27 -07:00
Jeremy McKeehen ff57f173d8 fix(desktop): keep pin list identity gateway-wide
Pin localStorage was keyed per connection and profile, so an unpin
reloaded a stale copy on switch and re-asserted pinned=true.
Scope pins by connection only so they survive rescope and stay isolated per gateway.
2026-08-26 00:22:27 -07:00
Kolton Jacobs 9577d66317 fix(desktop): probe a cached pooled remote backend before dispatching to it
A pooled remote backend (Bot Mode, group chat) keeps its descriptor and SSH
forward cached in the backend pool. When the remote Desktop relaunches, the
remote process dies but the local forward stays LISTENing, so
ensureRegistryBackend() keeps returning the dead descriptor and every dispatch
to that machine fails until the app is restarted.

The background sweep cannot cover this: revalidatePooledRemoteBackends() only
runs from the renderer reconnect IPC, which never fires while the primary
connection stays healthy.

Validate the exact cached descriptor at dispatch time with a short /api/status
probe (2.5 s). On failure, retire the pool entry and its SSH forward, then
reconnect on demand. Concurrent dispatches share one retire/reconnect sequence
through a RemoteRevalidationCoordinator keyed on the cached promise, and
identity checks make a late failure from an old descriptor unable to tear
down a replacement another caller already installed.

Verified on a two-Mac setup (MacBook + Mac mini over SSH): after relaunching
the Mac mini's Desktop, a group-chat turn from the MacBook now reaches the
mini's backend and its reply lands, where it previously failed forever.
2026-08-26 00:21:49 -07:00
RayCharlizard 616d6c5432 fix(desktop): retire the composer busy latch on gateway reconnect (#93059)
reconcileBusyStatesOnReconnect downgraded stale busy/awaiting claims by
writing the $sessionStates mirror directly. The claim has four holders —
the wiring cache, that mirror, the focused view's draft $busy /
$awaitingResponse, and busyRef — and only the write path (the delegate's
updateSessionState) keeps them in lockstep. After a reconnect that orphans
a mid-turn runtime (a respawned backend re-mints runtime ids, so the
terminal busy:false never arrives) the mirror cleared but the composer
stayed latched: Send failed isTargetSessionBusy and silently no-oped until
restart, and warm resume could OR the stale cache copy back over the
backend's running:false.

- SessionTileDelegate.retireBusyClaim?: optional twin of
  invalidateRuntimeBindings; writes through updateSessionState, returns
  false (and writes nothing) for a runtime the cache never held.
- reconcileBusyStatesOnReconnect routes each in-scope downgrade through it,
  keeps the mirror publish as the fallback, and on a primary reconcile also
  clears the focused draft latches. Scoped reconciles leave the composer
  alone.

Tests: hook (real useGatewayBoot + fake socket), store (write-path route,
miss fallback, primary vs scoped), cache (real updateSessionState) and
delegate (hit/miss) — RED on main, GREEN here. Full desktop UI suite,
typecheck and lint pass.

Written with LLMs under human direction: initial report and diagnosis by
GPT-5.6 (OpenAI Codex); root-cause refinement, design and review by
Claude Fable 5; implementation and tests by Claude Opus 5.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 00:21:49 -07:00
Flownium 2e75f9dc48 fix(desktop): preserve terminal state during reconnect hydration 2026-08-26 00:21:49 -07:00
Flownium 19251ac9e3 fix(desktop): hydrate transcript after reconnect attach 2026-08-26 00:21:49 -07:00
Teknium b52b05c1de test(desktop): profile-door dial failure now rejects (post-#81165 contract) — #92265 invariant unchanged 2026-08-26 00:21:29 -07:00
Teknium 57876f4b71 fix(desktop): typecheck fixes for salvaged tests (afterEach import, routed-request mock typing) 2026-08-26 00:21:29 -07:00
Teknium 9730bc78ff fix(desktop): feature-detect ctx.onDispose in the hide-sweep scheduler
Direct-file plugin hosts don't provide onDispose; every other call site in
plugin.js already guards it. Follow-up to the #94915 salvage.
2026-08-26 00:21:29 -07:00
Teknium 0366eacae3 fix(desktop): re-home the active key when the primary gateway re-homes
Follow-up to the #93892 keep-set salvage (#93916): the new
"remote tile keep-set must not pin a local same-named secondary" test
exposed a real scoping defect — route identity in the prune keep-set must
be full composite scope (connectionId + profile), never a bare profile
name inherited by accident.

setPrimaryGateway() moved g.primaryProfile without moving g.activeKey
when the active route WAS the primary. The stale bare-name activeKey
(e.g. 'default') then matched a later, unrelated LOCAL 'default'
secondary in pruneSecondaryGateways' `key === g.activeKey` spare, so a
keep-set of composite scopes like 'conn:homelab::default' appeared to
pin the local socket forever. Now the active key follows the primary
re-home, keeping the exact-scope identity contract intact.
2026-08-26 00:21:29 -07:00
fangliquanflq ef6532c25c test(desktop): cover bot reconciliation runtime lifecycle 2026-08-26 00:21:29 -07:00
fangliquanflq 66186dc58f fix(desktop): keep bot reconciliation off inactive backends 2026-08-26 00:21:29 -07:00
joaomarcos 3cf4f7abf7 fix(desktop): retain open pane gateway owners
Follow-up to the mode-switch teardown split: classify legacy secondaries
via an explicit isLegacySecondary() helper, keep open-pane gateway owners
in the boot keep-set, and cover the explicit `local` registry source not
being classified as legacy.

Salvaged from PR #94370. The PR's off-topic edit-composer changes
(user-edit-composer.tsx, user-message-edit.test.tsx — an unrelated edit
submit-cooldown tweak) were dropped from this cherry-pick.

Dropped-files: apps/desktop/src/components/assistant-ui/thread/user-edit-composer.tsx, apps/desktop/src/components/assistant-ui/thread/user-message-edit.test.tsx
2026-08-26 00:21:29 -07:00
joaomarcos 700ff78097 fix(desktop): preserve registered gateways during mode switches 2026-08-26 00:21:29 -07:00
LovePlayCode 1808d33a24 fix(desktop): keep owner-routed tile gateways out of idle prune
Bot chats stay on a secondary while chrome stays on the launch profile.
The keep-set only counted busy sessions, so idle prune closed the tile
socket and resume spun forever. Keep open tiles, route catalog reads to
the owner, and hydrate model/provider from resume.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-26 00:21:29 -07:00
ygd58 7e0b535610 fix(desktop): require an open socket before publishing a secondary gateway route
Fixes #92265 (proposed fix #2; #1 and #4 are separate follow-ups, see below).

ensureGatewayForAgent() and ensureGatewayForProfile() both decided
whether a secondary activation "succeeded" by checking Boolean(entry.connection)
alone. entry.connection is set in openSecondary() BEFORE the WebSocket
dial completes (`entry.connection = conn` happens ahead of
`await entry.gateway.connect(wsUrl)`), so a transient first-dial
failure -- caught by the surrounding try/catch and left for
scheduleReconnect's backoff retry -- still left entry.connection
truthy. Both functions then treated this as a successful activation:
applyActive() switched g.activeKey and published $gateway to the
closed socket, and publishActiveConnection() pushed the connection
descriptor to the UI. The next chat RPC then failed with "Hermes
gateway is not connected" against a route the user/desktop believed
was live.

Added an isOpen(entry.gateway) check alongside the existing
Boolean(entry.connection) check in both functions' activation/publish
conditions, gating BOTH applyActive() (which switches g.activeKey and
publishes $gateway) and publishActiveConnection() (which pushes the
connection descriptor) on the socket having actually reached 'open'.
A failed first dial now correctly returns false / leaves the previous
active route untouched, matching option 3 from the issue's own
proposed fix ("if both bounded attempts fail, keep the existing
active route") -- the existing scheduleReconnect backoff still owns
recovery for that entry going forward.

Not implemented in this PR (separate, lower-priority follow-ups):
- Proposed fix #1 (one immediate bounded reconnect attempt before
  returning activation status) -- a larger behavioral change with its
  own retry/timing tradeoffs; left to a separate PR.
- Proposed fix #4 (Bot Mode's own connection-ID-only guard in
  plugins/hermes-bots/plugin.js) -- host.ensureAgent() calls into the
  now-fixed gateway.ts functions, so this class of bug is already
  closed at the root; Bot Mode's own additional profile/state
  verification may still be worth adding but is a separate, narrower
  hardening pass on top of this fix.

Found and fixed a genuine test-suite inconsistency while verifying:
the existing "refreshes the active connection after a pooled profile
reconnect succeeds" test in gateway-shared-remote.test.ts asserted
setConnection was called once after a SINGLE ensureGatewayForProfile()
call whose first dial failed -- i.e. it encoded the exact bug this
issue reports as the EXPECTED, correct behavior. Rewrote it to assert
the corrected contract: the failed first attempt does not call
setConnection at all, and a realistic retry (calling
ensureGatewayForProfile() again, since g.activeKey correctly never
left the primary after the failed attempt -- ensureActiveGatewayOpen()
is for reconnecting an already-active gateway that went stale, not
retrying an activation that never succeeded) succeeds and publishes
once the second dial goes through.

Added a new test file (gateway-secondary-open-check.test.ts) following
the established mocking pattern from gateway-agent-scope.test.ts,
covering both ensureGatewayForAgent and ensureGatewayForProfile: a
transient first-dial failure does not activate/publish (the exact
reported symptom), and a successful dial still activates/publishes
normally (sanity, no regression to the happy path). Verified as
genuine regressions by reverting both isOpen() checks and confirming
2 of 4 new tests fail with exactly the reported symptom (activated
resolves true / the primary gets replaced despite the failed dial).

44/44 pass across all 9 gateway-related test files (no regression).

Dupe-swarm winner for issue #92265; Biotrioo (PR #92307) was the earliest
submitter of the swarm and deserves first-report credit.
2026-08-26 00:21:29 -07:00
hermes-seaeye[bot] d1627d5133 fmt(js): npm run fix on merge (#95329)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 06:51:21 +00:00
3x3xX3N0N b04f8578eb fix(desktop): first Windows update attempt no longer fails on dying backend processes (#74805)
taskkill /T /F returns when termination is INITIATED, not completed, and
the pre-handoff unlock gate only probed the venv hermes.exe shim — which
the 'python.exe -m hermes_cli.main serve' backend need not hold at all.
The gate could therefore pass on its first iteration with zero dwell
while the killed pythons were still unmapping .pyd files; the
venv-blocker scan (no liveness filter) then reported those dying
processes as holders and aborted the hand-off. Every first update
attempt from the footbar failed; the manual retry succeeded because the
process table had settled by then.

The unlock gate now lives in backend-release-gate.ts (dependency-free,
backend-child.ts pattern) and requires BOTH the shim unlocked AND every
signalled PID to have actually left the process table; stragglers
collected per-pass are killed and join the watch set. On deadline the
old shim-only criterion survives as the escape hatch — lingering PIDs
past 15s are the venv-blocker re-scan's job. applyUpdates additionally
re-scans up to 2x with a 1.5s settle before aborting on 'blocked', so
untracked grandchildren an AV driver holds in teardown stop failing the
update while a REAL holder still aborts on the third scan.

Surgical reapply of PR #78037 fix 1 by @3x3xX3N0N onto the post-#87599
code shape (stopBackendTreesForUpdate extraction, stopSafeBlockers
re-scan path). The re-scan settle idea was first submitted by
@MaheshBhushan (#74831); the killed-PID tracking seam matches
@webtecnica's #74956.

Co-authored-by: MaheshBhushan <128616744+MaheshBhushan@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Co-authored-by: Hermes <hermes@nousresearch.com>
2026-08-25 23:46:33 -07:00
Teknium a7eee2a7a7 fix(desktop): pin the complete macOS usage-description set + add reminders entitlement
Batch follow-ups on top of the salvaged privacy declarations:
- tests-js/desktop-mac-usage-descriptions.test.ts: EXPECTED_USAGE_DESCRIPTIONS
  now pins the FINAL key set (camera + calendar x2 + reminders x2 + screen
  capture + local network) so the drift-protection assertion locks the whole
  batch as permanent regression coverage.
- entitlements.mac.plist: add com.apple.security.personal-information.reminders
  alongside the calendars entitlement #65220 added — the reminders usage
  descriptions need the matching entitlement under hardened runtime (sibling
  site the original PR missed).
2026-08-25 23:33:46 -07:00
David Metcalfe f0e9902664 fix(desktop): declare NSAppleMusicUsageDescription to disclaim MediaLibrary TCC prompt
The Hermes Desktop renderer initializes Chromium's audio stack on user
gesture (completion chimes via Web Audio API in completion-sound.ts,
voice TTS via voice-playback.ts, mic capture via use-mic-recorder.ts,
and an eager AudioContext prime in haptics-provider.tsx). On macOS 26+,
that initialization registers the helper with the MediaLibrary TCC
service (kTCCServiceMediaLibrary), which surfaces to the user as a
"Hermes wants to access Music" permission prompt even though Hermes
never reads or writes the Apple Music library.

The Info.plist (built from apps/desktop/package.json's build.mac.extendInfo)
already declares NSAudioCaptureUsageDescription and
NSMicrophoneUsageDescription, but NSAppleMusicUsageDescription was missing
from the desktop app entirely. macOS therefore shows a system-default or
generic prompt for the MediaLibrary bucket instead of an honest description
from the app.

Fix
---
Add NSAppleMusicUsageDescription to build.mac.extendInfo with copy that
disclaims Music library access while explaining the system audio stack
uses voice, TTS, and completion sounds.

Add tests/test_desktop_mac_entitlements.py to pin every NS*UsageDescription
key declared in the Desktop build config. The test:
- parametrized over a (key, required_substring, reason) table
- asserts no leading/trailing whitespace and no newline chars in any usage
  string (electron-builder passes them through verbatim; control chars
  render as broken prompt text)
- asserts drift-protection: a new NS*UsageDescription key added to the
  build config without a matching test row causes a hard failure

Pattern reference: PR #59486 ("fix(desktop): add macOS contacts privacy
strings") is the open canonical for the same shape of fix for Contacts;
PR #64582 / PR #65220 extend it for Reminders. The closed duplicate PRs

Related, not in this PR
-----------------------
- PR #62601 (sounddevice on macOS) is the gateway/CLI side of the same
  kTCCServiceMediaLibrary trigger.
- PR #45952 (macOS permission broker foundation) is architectural work
  for centralized TCC handling; this fix does not depend on it.
- PR #52839 (browser automation Chrome launch) mutes Chromium audio in
  a different surface; the same pattern is recorded there.

Fixes #54551
2026-08-25 23:33:46 -07:00
Chen Jin 5d28665410 fix(desktop): declare NSLocalNetworkUsageDescription for macOS 15+ (#81563)
Since macOS 15 (Sequoia), an app that accesses the local network without
declaring NSLocalNetworkUsageDescription in its Info.plist is denied
silently: no prompt, no entry in System Settings → Privacy & Security →
Local Network, and LAN connections are dropped at the network layer
(manifesting as 'No route to host').

The desktop's build.mac.extendInfo declares camera/microphone/audio
usage descriptions but not the local-network one, so the terminal and
SSH features inside the app could never reach LAN hosts on macOS 15+.
Add the missing declaration.
2026-08-25 23:33:46 -07:00