Between the registry miss and the eager session.title write, another
writer can take the canonical title (peer dm minting server-side, a
second machine, cross-connection sync). UNIQUE(title) rejects our write
with 'already in use' — which the compat path previously read as 'old
gateway' and prompted into OUR stray lazy session, forking the forever
chat. A uniqueness rejection now re-consults the registry and adopts the
winner; the zero-message stray is abandoned to the gateway pruner.
Genuine old-gateway failures (unknown method) keep the compat kickoff.
Reconcile #94901's API-layer row stamping with #94656's durable-owner
persistence: extract lib/session-owner-stamp.ts as THE canonical
stamp-untagged-rows write path (never clobbers an explicit owner,
never stamps `local`) and re-express api/sessions'
stampActiveConnectionOwner through it. #94656's writers (optimistic
row from the captured owner route, mergeSessionPage carry, cache
patch) are exact-owner writers and stay as-is; the helper's contract
documents why it must not overwrite them.
Credit: row-stamping concept from PR #94901 (joe-rodgers) and
PR #95007 (weismanfamily); persistence shape from PR #94656
(Zeus-Deus).
Co-authored-by: joe-rodgers <25499388+joe-rodgers@users.noreply.github.com>
Partial cherry-pick of PR #94901 (joe-rodgers). Surviving scope:
- api/sessions: stampActiveConnectionOwner — rows returned by the
active non-local gateway are stamped with its registry connection_id
(explicit owners from multi-source responses preserved), so a later
resume cannot fall back to a same-named local profile.
- store/projects: one-shot projects.tree retry when a remote source
switch leaves the first read RPC on a newly-opened socket without a
response (request timed out / gateway connection closed), only while
the same gateway/profile is still foreground. Component fix for the
live-confirmed #92352 sidebar-never-paints gap.
Dropped scope (superseded on main / by the #94656 anchor landed just
below): knownSessionOwner+SessionOwnerScope rewiring in session.ts,
session-states.ts, wiring.tsx (main and #94656 carry richer variants),
and the $connection-derived optimistic-row stamp in
use-session-actions/utils.ts (#94656 stamps the optimistic row from
the captured exact owner route instead of ambient state).
Original-PR: #94901
Dropped-scope: routing half of 2cb5bdbf1 (session.ts, session-states.ts, wiring.tsx, use-session-actions/utils.ts hunks)
Follow-up to the #95007 partial cherry: waitForInitialConnection() was
an unbounded listen on $connection — a primary that never publishes its
descriptor (spawn failure, dead SSH target) would strand
initializeConnectionsRegistry() forever and the last-used source would
never be restored. Bound it with the codebase's withTimeout helper
(same pattern as the sibling SWITCH_* call sites in this file): after
45s (the primary spawn budget) the restore proceeds exactly as it did
before the wait existed, and the listener is torn down either way.
Regression test: boot restore proceeds after the deadline with the
descriptor never arriving.
Original-PR: #95007
Partial cherry-pick of PR #95007 (weismanfamily). Surviving scope:
- electron connection-registry: registrySourceOwnsPrimaryBackend() —
descriptor-level proof that a registry-scoped request names the
already-running primary backend, wired into ensureRegistryBackend as
the generic (non-SSH-fingerprint) primary-owns short-circuit so a
cloud/url registry primary cannot spawn a second isolated server.
- store/connections: waitForInitialConnection() before the boot-time
source restore, so the sidebar registry cannot dial the preferred
source a second time while the identical primary backend is still
publishing its connection identity.
Dropped scope (superseded on main): the renderer routing half —
primaryConnectionId plumbing, primaryOwnsAgent short-circuits in
requestGatewayForAgent/openGatewayForAgent/ensureGatewayForAgent
(main has isPrimaryRegistryRoute via 1ec32e738), the 3-arg
setPrimaryGateway boot wiring (main has setPrimaryGatewayConnection),
the use-session-list-actions stampConnectionOwner (row stamping lands
via #94656/#94901), and the wiring.tsx bare-profile promotion commit
65d106e41 (main's knownSessionOwner covers it).
Original-PR: #95007
Dropped-scope: renderer routing half of d1c0fb093; all of 65d106e41
selectProfile(name) / newSessionInProfile(name) keep only $newChatProfile and
clear $newChatRoute. #94147 taught the send path to capture the (registry
source, profile) pair as the draft's exact owner, but three gaps still let a
session created on the composite gateway conn:local::omar degrade to the bare
string "omar" — and requestGatewayForProfile("omar") is a DIFFERENT socket
than the one that minted the runtime, so the next session-scoped RPC 4001'd
"session not found" while the runtime was ws-orphan-reaped:
1. Ownership persistence. The exact owner lived only in the bounded,
in-memory owner-hint map. The primary aggregate serves a `local` registry
source's rows WITHOUT connection_id (the unified-list splice tags non-local
sources only), and mergeSessionPage replaced the optimistic row with that
untagged row on the first sidebar refresh. After a hint eviction or a
relaunch nothing exact was left.
- One canonical SessionOwnerRoute type (store/session-request-router);
AgentProfileRoute / SessionProfileRoute / SessionRpcOwnerRoute alias it.
- Every owner ladder gains the connection-tagged ROW rung
(knownSessionOwner / sessionOwnerRouteFromRow): the session-RPC
dispatcher, knownOwnerForSession, the tile delegate, foregroundSessionScopes
and the async probe (resolveSessionOwner) all yield the exact route when
the row carries its connection.
- mergeSessionPage carries connection_id onto a row that comes back
untagged for the same profile (merge, don't clobber).
- Owner hints are persisted (bounded LRU, hermes.desktop.sessionOwnerHints.v1)
and rehydrated in LRU order; a removed registry connection drops its
hints (use-gateway-boot onChanged) so fail-closed can't pin sessions to a
dead source.
2. Fail closed. A request carrying session_id whose owner no rung could name
silently fell to the ambient presentation gateway, turning missing metadata
into a misleading backend "session not found". createSessionRpcDispatcher
and requestForOwnedSession now reject with an explicit
SessionOwnerResolutionError (store/session-owner-resolution). The ONE case
where ambient is the owner by construction stays ambient: no registry
source live AND at most one profile (legacy single-backend Desktop, whose
older backends omit `profile` on rows). Main-pane runtime ids (native
approval.respond, queued sends) now translate to their stored id through
the per-runtime state mirror (storedSessionIdForRuntimeId), so they resolve
an owner instead of tripping the gate.
3. Lifecycle. Between session.create returning and the foreground publication
($selectedStoredSessionId via navigate → route effect, or $sessionTiles),
the owner entry had no active request and was not yet foreground-pinned:
a prune recompute or a refcount-0 lease release could close the socket
holding the just-minted runtime before the first prompt.submit. Both create
paths now hold retainGatewayForAgent across the create RPC and hand off to
holdSessionOwnerUntilForeground (session-states), which names the owner in
foregroundSessionScopes — every registry dispose path honors it — until the
session becomes selected/tiled, the caller releases it (failed create,
mid-create drift close), or a 60s TTL expires. Nothing latches.
Regressions:
- profile-rail-fresh-chat-owner.test.tsx: new case evicts the hint AND
merges an untagged refresh row after turn one; turn two still rides the
same conn:local::omar socket, no probe, no session.close, no v1 socket;
the existing case also asserts the owner is foreground-pinned from create.
- session-rpc-dispatcher.test.ts (new): fail-closed error + ambient never
called; legacy single-backend stays ambient; tagged-row rung routes by
runtime id; hint outranks untagged row; probe result routes exactly.
- session.test.ts: persisted hints survive a simulated relaunch in LRU order,
malformed storage is ignored, per-connection forget; knownSessionOwner;
mergeSessionPage carries / drops / preserves identity on connection_id.
- session-states-foreground-scopes.test.ts: tagged-row scope for the selected
thread; hold pins from create, retires on selection / tile mount, explicit
release, TTL expiry.
- session-states-runtime-map.test.ts: state-mirror rung; knownOwnerForSession
through mirror + hint / tagged row; requestForOwnedSession fail-closed vs
legacy ambient.
- wiring-routing.test.ts: connection-tagged row rung; hint still outranks it.
Stacked on #94145, #94147 and #94178 (merged as the integration base);
review from this commit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Sessions / profile-rail path (selectProfile, newSessionInProfile, a
connection switch, `/profile`) sets $newChatProfile and deliberately clears
$newChatRoute, so a fresh chat had no explicit owner. The session was created
on the active registry gateway (conn:local::omar) but its durable owner
degraded to the bare string "omar": follow-up RPCs dialed
requestGatewayForProfile("omar") — a different socket than the one that
minted the WebSocket-scoped runtime — and 4001'd "session not found" while
the runtime was ws-orphan-reaped.
- store/profile: capture the active registry source together with the
new-chat profile intent ($newChatConnectionId / captureNewChatSource) in
selectProfile, newSessionInProfile, newSessionInAgent, connection switches
and `/profile`; resolveNewChatOwnerRoute() derives the exact
{ connectionId, profile } route whenever a registry source is live, even
with $newChatRoute null (legacy v1 primary still yields null).
- use-session-actions: session.create, the owner hint, the optimistic row's
profile + connection_id, and the failed-create cleanup all use that
effective owner (main chat and tile paths).
- use-prompt-actions/submit: re-pin targetStoredSessionId after a fresh
create. It was captured before the create (null) and seedOptimistic handed
it to updateSessionState, which the state cache read as a DETACH — the
fresh stored↔runtime binding was severed the moment the chat existed, so
every later session-scoped RPC failed to translate the runtime id, never
saw the tile route / owner hint / row, probed REST by runtime id and fell
to the ambient socket.
- contrib: the session-RPC dispatcher is factored out of wiring.tsx
(createSessionRpcDispatcher) so the exact production routing is what the
integration test drives.
Regression (profile-rail-fresh-chat-owner.test.tsx) drives the real path:
mocked sockets under the real registry store, primary = remote default,
active source = local, selectProfile("omar") ($newChatProfile = "omar",
$newChatRoute = null), real useSessionStateCache / useSessionActions /
usePromptActions and the production dispatcher; asserts session.create and
BOTH prompt.submit calls hit the same conn:local::omar gateway object, no
session-scoped RPC reached the primary or a v1 "omar" socket, no
session.close, the binding survives both turns, no REST probe.
Verified with the packaged Linux Desktop against the real ~/.hermes
(primary = remote OAuth gateway, "This device" as registry source, omar via
the profile rail, two prompts): both prompts persisted on one session in
profiles/omar/state.db, no ws_orphan_reap.
Refs #94071
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A fresh chat created through $newChatRoute lost its owner the moment
session.create returned. The create RPC rode the captured route
(requestGatewayForAgent), but the optimistic row was stamped from
$activeGatewayProfile — still `default` in All-profiles / Bot routing —
and no owner hint was recorded. The first turn ran on the routed
backend (e.g. local::omar); every later session-scoped RPC resolved the
row as `default` and 4001'd "session not found", leaving the routed
runtime to be ws-orphan-reaped.
Make the ownership transition atomic with the create:
- record capturedRoute as the stored session's exact owner hint the
moment a routed create returns a stored id (main chat and tile paths);
- upsertOptimisticSession accepts an explicit owner and stamps
profile = targetProfile || profile plus the owning connection_id,
falling back to the ambient profile only for an unrouted create;
- contrib/wiring resolves session RPC owners as: persisted tile owner
route → exact unique owner hint → session-row profile → cross-profile
probe (resolveSessionRpcOwner, pure + unit-tested), so prompt.submit,
session.resume, attachments, interrupt, redirect and recovery all use
the exact owner; knownOwnerForSession follows the same ladder;
- the mid-create drift-abort session.close rides capturedRoute too.
Regressions: an integration test (ambient default, route local::omar,
two turns → both prompt.submit hit local::omar, no session-not-found,
no session.close) and a unit test asserting a routed fresh create never
yields { followupOwner: 'default', foregroundScope: 'conn:local::omar' }.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 95e2ef66becb641d3bdedf016dab6c104a777ae7)
The intro kickoff ('Hey, tell me about yourself!') now fires ONLY from
genuine New Agent creation. The bot-click canonical resolution path mints
silently: the eager session.title write already persists the lazy row on
modern gateways, so the kickoff's session-persistence job is obsolete
there. A resolution miss (retitled row, hidden-listing gap, post-update
skew) previously re-fired the kickoff on EVERY click — a burned model
turn plus a user-attributed prompt the user never typed (ScottFive
report). Older gateways that reject the eager title keep a narrow compat
kickoff, else the pruner reaps the empty lazy session.
teardownSshConnection closed the tunnel and SSH transport but never
killed the detached serve --isolated process. Spawn uses setsid/nohup,
so the backend reparents to pid 1, keeps state.db open, and accumulates
across Cmd+Q. Reuse cleanupStale via disconnect while SSH can still
exec, sequence remote kill before close, and seal the bootstrap
coordinator so reconnect during a prevented first quit cannot respawn.
The quit race is 6s to cover cleanupStale's 5s wait-for-exit loop.
Avoid opening a second SSH lifecycle when a migrated registry request targets the same primary/default backend already booted through the legacy route. Compare effective SSH configuration for representation-only drift, treat empty and default as the same root profile, and keep named profiles isolated.
The rebase re-landed pre-#74805 versions of the backend release gate,
venv-blocker rescan, mac entitlements/usage tests, and package.json from
the stale branch base — restored to main's versions (only the salvaged
enumeration/profileMetadata/profile:remember hunks kept in main.ts).
refreshActiveProfile's new bounded retry chain (#70679) is no longer
awaited inside the switch-completion barrier, so a slow/unhealthy backend
cannot hold $gatewaySwitching past the switch-ownership deadline; also
drop an unused $connection import from the earlier conflict compose.
Salvage of #94192's unique work (owner-hardening portions that overlap
the class-1 branch — #94824/#93451 seams — and out-of-cluster #94864 are
intentionally excluded):
- use-gateway-request: recognize the full transport-error family
(ECONNRESET & friends, including error.code and error.cause.code) so a
reset SSH/remote socket triggers the connection-owned reconnect instead
of surfacing as a request failure; background profiles keep the
registry reconnect path for composite remote/SSH sources.
- session-tile-actions: tile attachment uploads and session RPCs follow
the tile's composite owner (connectionId+profile) even when the active
gateway moved to a same-named profile on another source.
- knownSessionOwner: sessions expose their complete owner (registry
connection + profile) instead of a bare profile name that silently
collapsed the route back to the local path; delegate/wiring resolve
owners through it.
Fixes the SSH-reconnect share of #91365-adjacent routing gaps.
Salvaged (partial) from #94192.
Never-interacted remote bots painted as bare handles because roster rows
carried only profile names: display_name/title/ui_meta/has_avatar were
fetched lazily on first interaction (#91365). Thread credential-free
profile metadata from the enumeration-time /api/profiles body through
enumerateRegistryAgentSources (main.ts) and buildAgentRoster
(connection-registry.ts), keeping it attached to the connection-qualified
row across the same-install collapse. The plugin.js botRosterMeta half of
the original PR is dropped — superseded by landed #92731.
Fixes#91365
Salvaged (partial) from #92708.
The profile rail's live workspace switch never persisted the selection,
so the Desktop always booted back into the previous startup profile
(#79886). Route the successful primary-backend activation through a new
persistence-only hermes:profile:remember IPC (validated
writeActiveDesktopProfile) that records the choice WITHOUT tearing down
the backend or reloading the window like hermes:profile:set does.
Registry-source picks name another source's profiles and do not touch
the startup preference. Reapplied semantically over three weeks of
main.ts/preload.ts drift (selectProfile now routes through
activateOnCurrentSource, #91349/#91365 seams).
Fixes#79886
Salvaged from #79888.
Global remote mode fires refreshProfiles while the remote HTTP proxy is
still routing: the one-shot fetch failed silently and the rail stayed
empty until a manual refresh. Retry with 500ms/1000ms backoff, surface
terminal failures on the console, and dedupe concurrent callers into a
single retry chain (gateway open fires useBackgroundSync and the
activeGatewayProfile effect at once). Reapplied semantically on top of
the #85731 epoch guard: a stranded epoch stops the retry chain and
invalidation detaches the single-flight slot.
Fixes#70679
Salvaged from #74500.
Addresses AI-review feedback on #94653: note where the 'connect-on-demand'
sentinel is produced, and cover the interaction between
isLocalEnumerationFailure and localRouteFallbackProfiles directly (not just
the helper in isolation).
'connect-on-demand' means local roster enumeration was intentionally
skipped to avoid spawning a local backend on a remote-only workspace,
not that it failed. The plugin-profile-routes IPC handler passed
Boolean(error) straight through, so that deferral was treated as a
genuine failure and Bot Mode re-synthesized cached local profile rows
even though local was never dialed.
Fixes#94648
Old per-profile pin caches caused the stale unpin resurrection. Prove they are ignored and that sessions.pinned repopulates the gateway-wide key without a migration PATCH.
Pin localStorage was keyed per connection and profile, so an unpin
reloaded a stale copy on switch and re-asserted pinned=true.
Scope pins by connection only so they survive rescope and stay isolated per gateway.
A pooled remote backend (Bot Mode, group chat) keeps its descriptor and SSH
forward cached in the backend pool. When the remote Desktop relaunches, the
remote process dies but the local forward stays LISTENing, so
ensureRegistryBackend() keeps returning the dead descriptor and every dispatch
to that machine fails until the app is restarted.
The background sweep cannot cover this: revalidatePooledRemoteBackends() only
runs from the renderer reconnect IPC, which never fires while the primary
connection stays healthy.
Validate the exact cached descriptor at dispatch time with a short /api/status
probe (2.5 s). On failure, retire the pool entry and its SSH forward, then
reconnect on demand. Concurrent dispatches share one retire/reconnect sequence
through a RemoteRevalidationCoordinator keyed on the cached promise, and
identity checks make a late failure from an old descriptor unable to tear
down a replacement another caller already installed.
Verified on a two-Mac setup (MacBook + Mac mini over SSH): after relaunching
the Mac mini's Desktop, a group-chat turn from the MacBook now reaches the
mini's backend and its reply lands, where it previously failed forever.
reconcileBusyStatesOnReconnect downgraded stale busy/awaiting claims by
writing the $sessionStates mirror directly. The claim has four holders —
the wiring cache, that mirror, the focused view's draft $busy /
$awaitingResponse, and busyRef — and only the write path (the delegate's
updateSessionState) keeps them in lockstep. After a reconnect that orphans
a mid-turn runtime (a respawned backend re-mints runtime ids, so the
terminal busy:false never arrives) the mirror cleared but the composer
stayed latched: Send failed isTargetSessionBusy and silently no-oped until
restart, and warm resume could OR the stale cache copy back over the
backend's running:false.
- SessionTileDelegate.retireBusyClaim?: optional twin of
invalidateRuntimeBindings; writes through updateSessionState, returns
false (and writes nothing) for a runtime the cache never held.
- reconcileBusyStatesOnReconnect routes each in-scope downgrade through it,
keeps the mirror publish as the fallback, and on a primary reconcile also
clears the focused draft latches. Scoped reconciles leave the composer
alone.
Tests: hook (real useGatewayBoot + fake socket), store (write-path route,
miss fallback, primary vs scoped), cache (real updateSessionState) and
delegate (hit/miss) — RED on main, GREEN here. Full desktop UI suite,
typecheck and lint pass.
Written with LLMs under human direction: initial report and diagnosis by
GPT-5.6 (OpenAI Codex); root-cause refinement, design and review by
Claude Fable 5; implementation and tests by Claude Opus 5.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Follow-up to the #93892 keep-set salvage (#93916): the new
"remote tile keep-set must not pin a local same-named secondary" test
exposed a real scoping defect — route identity in the prune keep-set must
be full composite scope (connectionId + profile), never a bare profile
name inherited by accident.
setPrimaryGateway() moved g.primaryProfile without moving g.activeKey
when the active route WAS the primary. The stale bare-name activeKey
(e.g. 'default') then matched a later, unrelated LOCAL 'default'
secondary in pruneSecondaryGateways' `key === g.activeKey` spare, so a
keep-set of composite scopes like 'conn:homelab::default' appeared to
pin the local socket forever. Now the active key follows the primary
re-home, keeping the exact-scope identity contract intact.
Follow-up to the mode-switch teardown split: classify legacy secondaries
via an explicit isLegacySecondary() helper, keep open-pane gateway owners
in the boot keep-set, and cover the explicit `local` registry source not
being classified as legacy.
Salvaged from PR #94370. The PR's off-topic edit-composer changes
(user-edit-composer.tsx, user-message-edit.test.tsx — an unrelated edit
submit-cooldown tweak) were dropped from this cherry-pick.
Dropped-files: apps/desktop/src/components/assistant-ui/thread/user-edit-composer.tsx, apps/desktop/src/components/assistant-ui/thread/user-message-edit.test.tsx