Commit Graph

3759 Commits

Author SHA1 Message Date
hermes-seaeye[bot] 86ae906e88 fmt(js): npm run fix on merge (#95511)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 11:57:06 +00:00
miha 1bd5da3ac6 fix(desktop): skip macOS TCC-protected media dirs in git repo scan
The sidebar's home-dir repo crawl descends into ~/Pictures, ~/Music,
~/Movies and ~/Public — and into Photos/Music library packages — which
triggers Photos, Media Library and Files & Folders permission prompts
attributed to Hermes.app. Because the app is ad-hoc signed and re-signed
on every self-update (#49110), macOS drops all TCC grants after each
update, so these prompts re-fire every time.

Skip the media folders as direct children of a search root (nested dirs
like ~/dev/Music are ordinary and still scanned; an explicitly passed
root is still walked), and skip Apple library packages
(*.photoslibrary/*.musiclibrary/*.tvlibrary/*.aplibrary) at any depth.

Partial mitigation for #49110 / #52010: removes the Photos and Media
Library prompts entirely; the identity reset itself needs Developer ID
signed release artifacts (tracked in #49110).
2026-08-26 04:51:48 -07:00
Teknium bb3421bf25 fix(desktop): single-owner backend dial claim in Electron main (#90812)
reconnectGateway()'s in-flight lock lives at renderer module scope, so it
only dedupes reconnects inside ONE window. Two windows racing the same
wake both invoke the main-process backend ensure IPC, and for a pooled
SSH connection the loser of the pool-entry race could bootstrap a
duplicate remote backend (two tunnels, two remote serve processes).

Electron main is the single owner of backend lifecycles, so the claim
now lives there: BackendDialClaims keys in-flight dials by the pool
scope key from backendScopeKey(connectionId, profile) — the composite
identity seam wave-1 #93189 established for effective-identity reuse.
'hermes:connection' and 'hermes:connection:for' route through
backendDialClaims.run(), so concurrent renderer dials for one scope
coalesce onto one spawn and the second caller receives the first's
result. A claim exists only while its dial promise is unsettled: both
outcomes release it, a failed dial is never cached (fail closed, not
latched), and a synchronously-throwing dial rejects the claim instead
of escaping the seam.

The #93910 resume rebuild re-dials retired pool keys through the same
claim (redialPoolBackendAfterResume + new parseBackendScopeKey), so a
resume-driven rebuild and a concurrent renderer reconnect also coalesce
instead of racing.
2026-08-26 04:49:56 -07:00
Teknium 3123624c07 fix(desktop): revalidate pooled remote/SSH backends on power resume (#93910)
After macOS sleep/resume, pooled remote SSH descriptors kept serving dead
tunnels: a remote entry has no child 'exit' to clear it, the renderer
keepalive spares it from the idle reaper, and the wake-path nudges from
b90289b04/febed060a only re-drive the PRIMARY renderer socket. The
background failure-streak policy needs several probe rounds before it
drops a descriptor, so the Bots pane showed 'Gateway offline' long after
the network was back.

New revalidateSuspectPooledRemoteBackends(): on resume every pooled
remote is suspect — probe each once (bounded by
REMOTE_LIVENESS_TIMEOUT_MS), retire the dead ones immediately (pool
entry + SSH bootstrap + tunnel/master teardown) and rebuild them through
the caller's dial path, while healthy descriptors are left untouched. A
failed retire skips the rebuild (never dial on top of an installed
descriptor); a failed rebuild is logged and left to the renderer's
normal reconnect — the sweep never throws.

attachPowerResumeRemoteRevalidation() wires the sweep to the Electron
powerMonitor 'resume'/'unlock-screen' seam with a 15s holdoff so the
near-simultaneous macOS wake signals coalesce into one sweep and can
never form a hot loop; overlapping kicks additionally join the one
in-flight sweep via the existing RemoteRevalidationCoordinator.
2026-08-26 04:49:56 -07:00
Teknium b455abe0b3 fix(desktop): poll-guard reset is fire-and-forget off the redial path
composer-status imports $gateway from this module (cycle forces the
dynamic import), and awaiting the module load inside openSecondary sat on
the timed redial path — under CI load that pushed cold-start redials past
waitFor budgets in the lifecycle suite. The reset needs no ordering
guarantee relative to the dial; detach it.
2026-08-26 04:49:22 -07:00
Teknium 62534e2b5a fix(desktop): isolate the poll-guard reset import + sort-imports lint
The combined dynamic import meant a failed composer-status import (mocked
test graphs) silently skipped resetTileRuntimeBindings too — the exact
lifecycle regression CI caught. Separate best-effort trys per module.
2026-08-26 04:49:22 -07:00
Teknium fe615a0099 fix(desktop): republish the connections registry to renderers after every successful save (#95393)
Live-confirmed on the Phase B build: hermesDesktop.connections.save()
succeeds and the registry on disk gains the row, but the switcher menu
(fed by the renderer $connectionsRegistry snapshot) keeps painting the
stale list until reload. remove() already broadcasts
hermes:connections:changed; save() only did so on the dial-material-edit
branch, so a brand-new connection or a label rename never reached the
switcher's onChanged re-pull (or any other window).

Fix at the publish seam only: saveRegistryConnection now broadcasts a new
'saved' reason for every successful save that isn't a dial-material edit.
'saved' is a pure registry-refresh signal — the use-gateway-boot listener
explicitly ignores it (nothing moved, so no dispose/redial/forget), while
the switcher's existing onChanged listener re-pulls the snapshot.

Tests:
- electron/hardening.test.ts pins both broadcast branches in
  saveRegistryConnection (source-assertion pattern; main.ts has no exports).
- connection-switcher.test.tsx mirrors the live repro scenario
  (/tmp/mg-ab/w2_95393.py): menu before save lacks the row, Electron's
  'saved' push arrives, menu after — without reload — shows it.
2026-08-26 04:49:22 -07:00
Bruno Bza 06be6cffbc fix(desktop): release reconnect-orphaned warm transcripts once their authoritative state settles
A gateway connection that dies mid-turn leaves cached session snapshots
whose busy/awaitingResponse flags can never settle: the respawned
backend re-mints runtime ids, so no terminal publish ever reaches the
orphaned snapshot again. #isWarmSettled treated those frozen flags as
live work, so every orphan pinned its full warm transcript until app
restart — roughly 5MB per reconnect cycle, which turned the restart
loop in #95189 into renderer OOM.

SessionStateCache now accepts an optional isAuthoritativelyActive
probe. When wired, in-flight flags only block eviction while the
authoritative $sessionStates record still claims work for the same
runtime id; without the probe the legacy always-block behavior is
preserved byte-for-byte. Eviction remains gated on needsInput, pending
drafts, and active references, so a genuinely running turn (which
re-asserts busy on every publish) is never a casualty.

The useSessionStateCache hook wires the probe to the store it already
imports. Reconnect reconciliation (reconcileBusyStatesOnReconnect)
settles the authoritative record, and the next prune drains the
orphaned cache entry through the normal LRU path, ownership included.
2026-08-26 04:49:22 -07:00
Teknium a7ea156470 fix(desktop): harden the dead-session poll guard per #94950 review
Two review-thread deltas on the salvaged #94950 latch:

- Match the gateway's structured 4001 code, not a message substring, when
  the rejection carries one (JsonRpcGatewayError). A coded error that
  merely mentions 'session not found' in wrapped text (e.g. a 5007 tool
  failure) must not latch the guard and freeze the status stack on a
  healthy session. The substring fallback survives only for codeless
  legacy errors.

- Reset the latch on runtime re-mint, not only on status-stack rebind:
  wire resetBackgroundPollingGuard() at both reconnect seams that already
  drop stale runtime bindings (use-gateway-boot's post-reconnect
  resetTileRuntimeBindings and gateway.ts's reopening path), so ids the
  dead runtime 4001'd resume polling once a respawned backend re-mints
  them.

Tests: code-specific match both directions; full-reset resumes every
latched session.
2026-08-26 04:49:22 -07:00
Justin Johnson c19849cd02 fix(desktop): stop the status-stack poll storming a dead session with 4001s
The composer status stack polls `process.list` every 5s while a background
process row is on screen. `process.list` is session-scoped, so against a
runtime id the gateway no longer holds it returns 4001 "session not found".

`refreshBackgroundProcesses` swallowed *every* failure with a bare `catch {}`
commented "transient socket loss". A gone session is not transient: the poll
re-sent the same dead runtime id every 5 seconds for the lifetime of the
window. On one machine this produced 31,518 gateway rejections in a day
(vs 663 the day before), 18,614 of them against a single runtime id, and it
is what users see reported as "sessions stopped with a session not found
error" after an update.

The trigger is a reconnect, not the poll itself: anything that mints a fresh
runtime (gateway restart, the #94219 reconnect/replay work, an idle-reaped
pooled backend) strands the id the status stack is still holding, and nothing
in this path ever re-checked it.

Distinguish the two failure classes:

- 4001 / "session not found" is TERMINAL for that runtime id — latch the id
  and stop polling it.
- A timeout or transport error is transient — keep retrying, since the
  session may well still be alive. Misclassifying that direction would
  silently freeze the status stack on a healthy session.

The latch is cleared when the status stack (re)binds a session id, so a
session that comes back under a fresh runtime resumes polling normally
rather than staying dark for the life of the app.

Also name the method in the gateway's 4001 warning. That line was added in
c305839442 "for diagnosability", but without the RPC name it cannot say WHICH
client call is looping — the reason this storm could not be attributed from
the logs alone. A ContextVar set in `handle_request` carries it; it is
diagnostic only and never used for authorization.

Tests:
- composer-status: 4001 stops the poll, a timeout does not, one gone session
  never suppresses a healthy sibling, and a rebind resumes polling.
- tui_gateway: the rejection warning names the method.
2026-08-26 04:49:22 -07:00
Teknium 574bd7175c fix(desktop): unify boot-class getConnection() budgets on one shared 45s constant
Follow-up to the #95039 salvage: the cherry-picked bound used the 20s
RECONNECT_ATTEMPT_TIMEOUT_MS on boot()/softSwitch() getConnection(), but a
reviewer note (and the Phase A registry-restore work) established that
boot-class awaits must ride out a full backend cold spawn — main's spawn
budget is 45s (DEFAULT_BACKEND_READY_TIMEOUT_MS). A 20s renderer bound would
latch boot errors on healthy-but-slow cold boots.

Introduce BACKEND_BOOT_WAIT_TIMEOUT_MS (45s) in lib/with-timeout.ts as the
single shared boot-class budget, point boot()/softSwitch() getConnection()
and connections.ts BOOT_DESCRIPTOR_WAIT_TIMEOUT_MS at it, and keep the 20s
reconnect budget only for reconnect-class awaits against an already-spawned
backend. No magic-number drift: 45_000 now appears once in renderer code.
2026-08-26 04:49:22 -07:00
nftpoetrist 31f3de1f06 fix(desktop): bound getConnection() on the boot and soft-switch paths (#93454)
resolveGatewayWsUrl() in attemptReconnect() because a wedged IPC
round-trip into the main process (e.g. a stuck revalidation after a
liveness-probe trip) can hang these awaits forever. A later fix
(e8d5660bae) extended the bound to resolveGatewayWsUrl() in boot() and
softSwitch() too, but left the getConnection() call immediately above
it in both functions unbounded.

If that call wedges: during initial boot the 'Starting Hermes...'
screen never resolves (bootCompleted never flips, nothing hits catch),
and during a soft gateway/profile switch  latches
true forever since the try block's finally never runs.

Wrap both with the same withTimeout()/RECONNECT_ATTEMPT_TIMEOUT_MS
pattern already used for the sibling calls.
2026-08-26 04:49:22 -07:00
Teknium d22e2b9f6e fix(desktop): a canonical-title race adopts the winner instead of forking the forever chat (#92473, part 2)
Between the registry miss and the eager session.title write, another
writer can take the canonical title (peer dm minting server-side, a
second machine, cross-connection sync). UNIQUE(title) rejects our write
with 'already in use' — which the compat path previously read as 'old
gateway' and prompted into OUR stray lazy session, forking the forever
chat. A uniqueness rejection now re-consults the registry and adopts the
winner; the zero-message stray is abandoned to the gateway pruner.
Genuine old-gateway failures (unknown method) keep the compat kickoff.
2026-08-26 04:03:35 -07:00
hermes-seaeye[bot] c427367938 fmt(js): npm run fix on merge (#95457)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 10:27:25 +00:00
Teknium 8e9459c97f style(desktop): sort registrySourceOwnsPrimaryBackend imports (lint) 2026-08-26 03:22:37 -07:00
Teknium 2947272233 refactor(desktop): one canonical write shape for connection_id row stamping
Reconcile #94901's API-layer row stamping with #94656's durable-owner
persistence: extract lib/session-owner-stamp.ts as THE canonical
stamp-untagged-rows write path (never clobbers an explicit owner,
never stamps `local`) and re-express api/sessions'
stampActiveConnectionOwner through it. #94656's writers (optimistic
row from the captured owner route, mergeSessionPage carry, cache
patch) are exact-owner writers and stay as-is; the helper's contract
documents why it must not overwrite them.

Credit: row-stamping concept from PR #94901 (joe-rodgers) and
PR #95007 (weismanfamily); persistence shape from PR #94656
(Zeus-Deus).

Co-authored-by: joe-rodgers <25499388+joe-rodgers@users.noreply.github.com>
2026-08-26 03:22:37 -07:00
joe-rodgers fb393ee08b fix(desktop): stamp remote list rows with their owning connection; retry one transient projects.tree loss
Partial cherry-pick of PR #94901 (joe-rodgers). Surviving scope:
- api/sessions: stampActiveConnectionOwner — rows returned by the
  active non-local gateway are stamped with its registry connection_id
  (explicit owners from multi-source responses preserved), so a later
  resume cannot fall back to a same-named local profile.
- store/projects: one-shot projects.tree retry when a remote source
  switch leaves the first read RPC on a newly-opened socket without a
  response (request timed out / gateway connection closed), only while
  the same gateway/profile is still foreground. Component fix for the
  live-confirmed #92352 sidebar-never-paints gap.

Dropped scope (superseded on main / by the #94656 anchor landed just
below): knownSessionOwner+SessionOwnerScope rewiring in session.ts,
session-states.ts, wiring.tsx (main and #94656 carry richer variants),
and the $connection-derived optimistic-row stamp in
use-session-actions/utils.ts (#94656 stamps the optimistic row from
the captured exact owner route instead of ambient state).

Original-PR: #94901
Dropped-scope: routing half of 2cb5bdbf1 (session.ts, session-states.ts, wiring.tsx, use-session-actions/utils.ts hunks)
2026-08-26 03:22:37 -07:00
Teknium b2a58dbb39 fix(desktop): bound the boot descriptor wait so a dead primary cannot strand the registry restore
Follow-up to the #95007 partial cherry: waitForInitialConnection() was
an unbounded listen on $connection — a primary that never publishes its
descriptor (spawn failure, dead SSH target) would strand
initializeConnectionsRegistry() forever and the last-used source would
never be restored. Bound it with the codebase's withTimeout helper
(same pattern as the sibling SWITCH_* call sites in this file): after
45s (the primary spawn budget) the restore proceeds exactly as it did
before the wait existed, and the listener is torn down either way.

Regression test: boot restore proceeds after the deadline with the
descriptor never arriving.

Original-PR: #95007
2026-08-26 03:22:37 -07:00
Michael Weisman fd031488aa fix(desktop): prove registry-primary backend ownership electron-side; wait for the primary descriptor before boot restore
Partial cherry-pick of PR #95007 (weismanfamily). Surviving scope:
- electron connection-registry: registrySourceOwnsPrimaryBackend() —
  descriptor-level proof that a registry-scoped request names the
  already-running primary backend, wired into ensureRegistryBackend as
  the generic (non-SSH-fingerprint) primary-owns short-circuit so a
  cloud/url registry primary cannot spawn a second isolated server.
- store/connections: waitForInitialConnection() before the boot-time
  source restore, so the sidebar registry cannot dial the preferred
  source a second time while the identical primary backend is still
  publishing its connection identity.

Dropped scope (superseded on main): the renderer routing half —
primaryConnectionId plumbing, primaryOwnsAgent short-circuits in
requestGatewayForAgent/openGatewayForAgent/ensureGatewayForAgent
(main has isPrimaryRegistryRoute via 1ec32e738), the 3-arg
setPrimaryGateway boot wiring (main has setPrimaryGatewayConnection),
the use-session-list-actions stampConnectionOwner (row stamping lands
via #94656/#94901), and the wiring.tsx bare-profile promotion commit
65d106e41 (main's knownSessionOwner covers it).

Original-PR: #95007
Dropped-scope: renderer routing half of d1c0fb093; all of 65d106e41
2026-08-26 03:22:37 -07:00
Deus d6e323bd63 fix(desktop): fail closed through registry boot and drain owner holds 2026-08-26 03:22:37 -07:00
Deus a0578f4ef3 test(desktop): model registry topology in dispatcher coverage 2026-08-26 03:22:37 -07:00
Deus f912657602 test(desktop): prove fresh chat socket continuity 2026-08-26 03:22:37 -07:00
Deus c662e1be7b fix(desktop): retain primary gateway registry identity after reload 2026-08-26 03:22:37 -07:00
Deus 6fdf873464 fix(desktop): serialize profile switches and drain edit redials 2026-08-26 03:22:37 -07:00
Deus 962b308be1 fix(desktop): require exact owners in registry topology 2026-08-26 03:22:37 -07:00
Zeus-Deus 07b87f1470 fix(desktop): profile-rail fresh chats keep one exact session owner from create through every later RPC
selectProfile(name) / newSessionInProfile(name) keep only $newChatProfile and
clear $newChatRoute. #94147 taught the send path to capture the (registry
source, profile) pair as the draft's exact owner, but three gaps still let a
session created on the composite gateway conn:local::omar degrade to the bare
string "omar" — and requestGatewayForProfile("omar") is a DIFFERENT socket
than the one that minted the runtime, so the next session-scoped RPC 4001'd
"session not found" while the runtime was ws-orphan-reaped:

1. Ownership persistence. The exact owner lived only in the bounded,
   in-memory owner-hint map. The primary aggregate serves a `local` registry
   source's rows WITHOUT connection_id (the unified-list splice tags non-local
   sources only), and mergeSessionPage replaced the optimistic row with that
   untagged row on the first sidebar refresh. After a hint eviction or a
   relaunch nothing exact was left.
   - One canonical SessionOwnerRoute type (store/session-request-router);
     AgentProfileRoute / SessionProfileRoute / SessionRpcOwnerRoute alias it.
   - Every owner ladder gains the connection-tagged ROW rung
     (knownSessionOwner / sessionOwnerRouteFromRow): the session-RPC
     dispatcher, knownOwnerForSession, the tile delegate, foregroundSessionScopes
     and the async probe (resolveSessionOwner) all yield the exact route when
     the row carries its connection.
   - mergeSessionPage carries connection_id onto a row that comes back
     untagged for the same profile (merge, don't clobber).
   - Owner hints are persisted (bounded LRU, hermes.desktop.sessionOwnerHints.v1)
     and rehydrated in LRU order; a removed registry connection drops its
     hints (use-gateway-boot onChanged) so fail-closed can't pin sessions to a
     dead source.

2. Fail closed. A request carrying session_id whose owner no rung could name
   silently fell to the ambient presentation gateway, turning missing metadata
   into a misleading backend "session not found". createSessionRpcDispatcher
   and requestForOwnedSession now reject with an explicit
   SessionOwnerResolutionError (store/session-owner-resolution). The ONE case
   where ambient is the owner by construction stays ambient: no registry
   source live AND at most one profile (legacy single-backend Desktop, whose
   older backends omit `profile` on rows). Main-pane runtime ids (native
   approval.respond, queued sends) now translate to their stored id through
   the per-runtime state mirror (storedSessionIdForRuntimeId), so they resolve
   an owner instead of tripping the gate.

3. Lifecycle. Between session.create returning and the foreground publication
   ($selectedStoredSessionId via navigate → route effect, or $sessionTiles),
   the owner entry had no active request and was not yet foreground-pinned:
   a prune recompute or a refcount-0 lease release could close the socket
   holding the just-minted runtime before the first prompt.submit. Both create
   paths now hold retainGatewayForAgent across the create RPC and hand off to
   holdSessionOwnerUntilForeground (session-states), which names the owner in
   foregroundSessionScopes — every registry dispose path honors it — until the
   session becomes selected/tiled, the caller releases it (failed create,
   mid-create drift close), or a 60s TTL expires. Nothing latches.

Regressions:
- profile-rail-fresh-chat-owner.test.tsx: new case evicts the hint AND
  merges an untagged refresh row after turn one; turn two still rides the
  same conn:local::omar socket, no probe, no session.close, no v1 socket;
  the existing case also asserts the owner is foreground-pinned from create.
- session-rpc-dispatcher.test.ts (new): fail-closed error + ambient never
  called; legacy single-backend stays ambient; tagged-row rung routes by
  runtime id; hint outranks untagged row; probe result routes exactly.
- session.test.ts: persisted hints survive a simulated relaunch in LRU order,
  malformed storage is ignored, per-connection forget; knownSessionOwner;
  mergeSessionPage carries / drops / preserves identity on connection_id.
- session-states-foreground-scopes.test.ts: tagged-row scope for the selected
  thread; hold pins from create, retires on selection / tile mount, explicit
  release, TTL expiry.
- session-states-runtime-map.test.ts: state-mirror rung; knownOwnerForSession
  through mirror + hint / tagged row; requestForOwnedSession fail-closed vs
  legacy ambient.
- wiring-routing.test.ts: connection-tagged row rung; hint still outranks it.

Stacked on #94145, #94147 and #94178 (merged as the integration base);
review from this commit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 03:22:37 -07:00
Zeus-Deus cb75983abf fix(desktop): profile-rail fresh chats keep their registry source as the exact owner
The Sessions / profile-rail path (selectProfile, newSessionInProfile, a
connection switch, `/profile`) sets $newChatProfile and deliberately clears
$newChatRoute, so a fresh chat had no explicit owner. The session was created
on the active registry gateway (conn:local::omar) but its durable owner
degraded to the bare string "omar": follow-up RPCs dialed
requestGatewayForProfile("omar") — a different socket than the one that
minted the WebSocket-scoped runtime — and 4001'd "session not found" while
the runtime was ws-orphan-reaped.

- store/profile: capture the active registry source together with the
  new-chat profile intent ($newChatConnectionId / captureNewChatSource) in
  selectProfile, newSessionInProfile, newSessionInAgent, connection switches
  and `/profile`; resolveNewChatOwnerRoute() derives the exact
  { connectionId, profile } route whenever a registry source is live, even
  with $newChatRoute null (legacy v1 primary still yields null).
- use-session-actions: session.create, the owner hint, the optimistic row's
  profile + connection_id, and the failed-create cleanup all use that
  effective owner (main chat and tile paths).
- use-prompt-actions/submit: re-pin targetStoredSessionId after a fresh
  create. It was captured before the create (null) and seedOptimistic handed
  it to updateSessionState, which the state cache read as a DETACH — the
  fresh stored↔runtime binding was severed the moment the chat existed, so
  every later session-scoped RPC failed to translate the runtime id, never
  saw the tile route / owner hint / row, probed REST by runtime id and fell
  to the ambient socket.
- contrib: the session-RPC dispatcher is factored out of wiring.tsx
  (createSessionRpcDispatcher) so the exact production routing is what the
  integration test drives.

Regression (profile-rail-fresh-chat-owner.test.tsx) drives the real path:
mocked sockets under the real registry store, primary = remote default,
active source = local, selectProfile("omar") ($newChatProfile = "omar",
$newChatRoute = null), real useSessionStateCache / useSessionActions /
usePromptActions and the production dispatcher; asserts session.create and
BOTH prompt.submit calls hit the same conn:local::omar gateway object, no
session-scoped RPC reached the primary or a v1 "omar" socket, no
session.close, the binding survives both turns, no REST probe.

Verified with the packaged Linux Desktop against the real ~/.hermes
(primary = remote OAuth gateway, "This device" as registry source, omar via
the profile rail, two prompts): both prompts persisted on one session in
profiles/omar/state.db, no ws_orphan_reap.

Refs #94071

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 03:22:37 -07:00
Zeus-Deus 77edf1b4df fix(desktop): routed fresh chat keeps its exact owner after session.create
A fresh chat created through $newChatRoute lost its owner the moment
session.create returned. The create RPC rode the captured route
(requestGatewayForAgent), but the optimistic row was stamped from
$activeGatewayProfile — still `default` in All-profiles / Bot routing —
and no owner hint was recorded. The first turn ran on the routed
backend (e.g. local::omar); every later session-scoped RPC resolved the
row as `default` and 4001'd "session not found", leaving the routed
runtime to be ws-orphan-reaped.

Make the ownership transition atomic with the create:
- record capturedRoute as the stored session's exact owner hint the
  moment a routed create returns a stored id (main chat and tile paths);
- upsertOptimisticSession accepts an explicit owner and stamps
  profile = targetProfile || profile plus the owning connection_id,
  falling back to the ambient profile only for an unrouted create;
- contrib/wiring resolves session RPC owners as: persisted tile owner
  route → exact unique owner hint → session-row profile → cross-profile
  probe (resolveSessionRpcOwner, pure + unit-tested), so prompt.submit,
  session.resume, attachments, interrupt, redirect and recovery all use
  the exact owner; knownOwnerForSession follows the same ladder;
- the mid-create drift-abort session.close rides capturedRoute too.

Regressions: an integration test (ambient default, route local::omar,
two turns → both prompt.submit hit local::omar, no session-not-found,
no session.close) and a unit test asserting a routed fresh create never
yields { followupOwner: 'default', foregroundScope: 'conn:local::omar' }.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 95e2ef66becb641d3bdedf016dab6c104a777ae7)
2026-08-26 03:22:37 -07:00
Teknium 4faa721d7d fix(desktop): clicking a bot no longer burns a model turn on a fake user prompt
The intro kickoff ('Hey, tell me about yourself!') now fires ONLY from
genuine New Agent creation. The bot-click canonical resolution path mints
silently: the eager session.title write already persists the lazy row on
modern gateways, so the kickoff's session-persistence job is obsolete
there. A resolution miss (retitled row, hidden-listing gap, post-update
skew) previously re-fired the kickoff on EVERY click — a burned model
turn plus a user-attributed prompt the user never typed (ScottFive
report). Older gateways that reject the eager title keep a narrow compat
kickoff, else the pruner reaps the empty lazy session.
2026-08-26 00:52:54 -07:00
hermes-seaeye[bot] 1a19c52dd2 fmt(js): npm run fix on merge (#95365)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 07:45:59 +00:00
Teknium 25d46c7887 fix(desktop): drop unused registryGatewayWsUrl import left by the header-binding refactor rebase 2026-08-26 00:40:35 -07:00
Victor Nogueira 21f34794be fix(desktop): recover remote sessions after gateway restart 2026-08-26 00:40:35 -07:00
686f6c61 8085614f6a fix(desktop): log when Windows remote SSH skip teardown
POSIX disconnect does not apply to connectWindowsRemote. Quit still
closes the tunnel; log the skipped serve kill so it is not a silent
no-op.
2026-08-26 00:40:35 -07:00
686f6c61 0c69cac48a fix(desktop): terminate owned SSH serve backends on quit
teardownSshConnection closed the tunnel and SSH transport but never
killed the detached serve --isolated process. Spawn uses setsid/nohup,
so the backend reparents to pid 1, keeps state.db open, and accumulates
across Cmd+Q. Reuse cleanupStale via disconnect while SSH can still
exec, sequence remote kill before close, and seal the bootstrap
coordinator so reconnect during a prevented first quit cannot respawn.
The quit race is 6s to cover cleanupStale's 5s wait-for-exit loop.
2026-08-26 00:40:35 -07:00
Jaime Marques 14d16c2578 fix(desktop): clarify primary SSH reuse failures 2026-08-26 00:40:35 -07:00
Jaime Marques 3263ca2af6 fix(desktop): reuse migrated primary SSH backend
Avoid opening a second SSH lifecycle when a migrated registry request targets the same primary/default backend already booted through the legacy route. Compare effective SSH configuration for representation-only drift, treat empty and default as the same root profile, and keep named profiles isolated.
2026-08-26 00:40:35 -07:00
Ahmett101 1b8f3eda38 fix(desktop): scope registered ssh primary gateway 2026-08-26 00:40:35 -07:00
Ravi Tharuma 949f5169de fix(desktop): treat ticket 401 as sign-in when native tokens are unreadable 2026-08-26 00:40:35 -07:00
Marco Fernstaedt 024b9c0545 fix(desktop): refresh remote WebSocket header cache recency 2026-08-26 00:40:35 -07:00
Marco Fernstaedt b4162de333 fix(desktop): bind headers to scoped WebSocket URL 2026-08-26 00:40:35 -07:00
Teknium 5a285d3436 fix(desktop): restore stale-branch-reverted main.ts/update files; detach post-switch profile refresh from switch completion
The rebase re-landed pre-#74805 versions of the backend release gate,
venv-blocker rescan, mac entitlements/usage tests, and package.json from
the stale branch base — restored to main's versions (only the salvaged
enumeration/profileMetadata/profile:remember hunks kept in main.ts).
refreshActiveProfile's new bounded retry chain (#70679) is no longer
awaited inside the switch-completion barrier, so a slow/unhealthy backend
cannot hold $gatewaySwitching past the switch-ownership deadline; also
drop an unused $connection import from the earlier conflict compose.
2026-08-26 00:22:27 -07:00
Teknium b722177fcd fix(desktop): drop duplicate knownSessionOwner re-landed by rebase (main's richer variant wins) 2026-08-26 00:22:27 -07:00
tachi 317ae240fb fix(desktop): SSH/reconnect owner continuity, attachment routing, transport-error recovery
Salvage of #94192's unique work (owner-hardening portions that overlap
the class-1 branch — #94824/#93451 seams — and out-of-cluster #94864 are
intentionally excluded):

- use-gateway-request: recognize the full transport-error family
  (ECONNRESET & friends, including error.code and error.cause.code) so a
  reset SSH/remote socket triggers the connection-owned reconnect instead
  of surfacing as a request failure; background profiles keep the
  registry reconnect path for composite remote/SSH sources.
- session-tile-actions: tile attachment uploads and session RPCs follow
  the tile's composite owner (connectionId+profile) even when the active
  gateway moved to a same-named profile on another source.
- knownSessionOwner: sessions expose their complete owner (registry
  connection + profile) instead of a bare profile name that silently
  collapsed the route back to the local path; delegate/wiring resolve
  owners through it.

Fixes the SSH-reconnect share of #91365-adjacent routing gaps.
Salvaged (partial) from #94192.
2026-08-26 00:22:27 -07:00
Tom 2ed39365d6 fix(desktop): thread eager profile metadata through registry enumeration
Never-interacted remote bots painted as bare handles because roster rows
carried only profile names: display_name/title/ui_meta/has_avatar were
fetched lazily on first interaction (#91365). Thread credential-free
profile metadata from the enumeration-time /api/profiles body through
enumerateRegistryAgentSources (main.ts) and buildAgentRoster
(connection-registry.ts), keeping it attached to the connection-qualified
row across the same-install collapse. The plugin.js botRosterMeta half of
the original PR is dropped — superseded by landed #92731.

Fixes #91365
Salvaged (partial) from #92708.
2026-08-26 00:22:27 -07:00
Tilly-YL 2952119bce fix(desktop): remember selected profile across restarts
The profile rail's live workspace switch never persisted the selection,
so the Desktop always booted back into the previous startup profile
(#79886). Route the successful primary-backend activation through a new
persistence-only hermes:profile:remember IPC (validated
writeActiveDesktopProfile) that records the choice WITHOUT tearing down
the backend or reloading the window like hermes:profile:set does.
Registry-source picks name another source's profiles and do not touch
the startup preference. Reapplied semantically over three weeks of
main.ts/preload.ts drift (selectProfile now routes through
activateOnCurrentSource, #91349/#91365 seams).

Fixes #79886
Salvaged from #79888.
2026-08-26 00:22:27 -07:00
David Metcalfe 57043c2bc0 fix(desktop): single-flight refreshProfiles with retry recovery in global remote mode
Global remote mode fires refreshProfiles while the remote HTTP proxy is
still routing: the one-shot fetch failed silently and the rail stayed
empty until a manual refresh. Retry with 500ms/1000ms backoff, surface
terminal failures on the console, and dedupe concurrent callers into a
single retry chain (gateway open fires useBackgroundSync and the
activeGatewayProfile effect at once). Reapplied semantically on top of
the #85731 epoch guard: a stranded epoch stops the retry chain and
invalidation detaches the single-flight slot.

Fixes #70679
Salvaged from #74500.
2026-08-26 00:22:27 -07:00
chelsealong a928596758 fix(desktop): document connect-on-demand origin, add fallback-profiles integration test
Addresses AI-review feedback on #94653: note where the 'connect-on-demand'
sentinel is produced, and cover the interaction between
isLocalEnumerationFailure and localRouteFallbackProfiles directly (not just
the helper in isolation).
2026-08-26 00:22:27 -07:00
chelsealong c475484f63 fix(desktop): do not treat deferred local enumeration as a failure
'connect-on-demand' means local roster enumeration was intentionally
skipped to avoid spawning a local backend on a remote-only workspace,
not that it failed. The plugin-profile-routes IPC handler passed
Boolean(error) straight through, so that deferral was treated as a
genuine failure and Bot Mode re-synthesized cached local profile rows
even though local was never dialed.

Fixes #94648
2026-08-26 00:22:27 -07:00
Jeremy McKeehen 4ca1f532be test(desktop): pass the pin-write fence into the Show-all order assertion
resolvePinnedSessions requires unconfirmedPinWrites; the reconnection-scope test omitted it and would fail strict tsc.
2026-08-26 00:22:27 -07:00
Jeremy McKeehen 3751b04550 test(desktop): lock pin upgrade to server-authoritative pull
Old per-profile pin caches caused the stale unpin resurrection. Prove they are ignored and that sessions.pinned repopulates the gateway-wide key without a migration PATCH.
2026-08-26 00:22:27 -07:00