`openSessionInNewWindow` → IPC `hermes:window:openSession` →
`buildSessionWindowUrl` emitted no `profile`, so a secondary window (⇧⌘-click
pop-out, subagent watch) was a full renderer that adopted the PRIMARY
backend's profile and resolved the session id against the wrong store —
blank/wrong session for any non-primary profile (#82768, #61286).
The owning profile now rides the URL as `&profile=`, exactly the carry the
HUD already does (buildHudWindowUrl / windowProfileOverride in
use-gateway-boot); the renderer picks it with the same ladder openHud uses:
the session's stamped owner wins, an unstamped/uncached id (a brand-new
subagent child) inherits the profile the user is looking at.
Diagnosis credit: @DomGrieco (#82794).
Co-authored-by: DomGrieco <6556434+DomGrieco@users.noreply.github.com>
When the bundle was swapped under a running process, the About banner sent the
user to the installer — a download and a reinstall for a state that a plain
restart repairs, and the reason reinstalling never helped these reports.
Report bundleSwapPending on hermes:version and give that case its own copy and
a "Restart Hermes" button. It gets its own headline too: reusing "App build out
of date" over a body that says the app is already installed repeats the
contradiction with the Updates card that the banner is supposed to resolve. The
installer link stays for the genuinely-stale-bundle case.
Packaged builds only — a dev `--build-only` rewrites the stamp under a running
`npm start`, and that is a rebuild the developer asked for, not a torn install.
Co-authored-by: tk-pkm111 <133480534+tk-pkm111@users.noreply.github.com>
A user who reopens Hermes while an update is running lands on the boot gate,
which is what it is for. But the updater swaps the packaged bundle on disk
after `hermes update` exits, and its `open` leg only focuses this already-
running process, so nothing ever loads the new build. The parked instance then
passes the gate and boots the new runtime under the old renderer — the "App
build out of date" banner immediately after a fully successful update, over an
Updates card that says "You're on the latest version" and so offers no remedy.
Compare the install stamp this process loaded at boot with the one on disk when
the gate clears. On positive proof of a swap — different commit, or a different
builtAt at the same commit — relaunch instead of starting a backend. Detection
fails quiet like bundle-skew, so a swap that never happened (the Windows
locked-binary case) is unchanged. A one-shot argv flag makes a relaunch loop
impossible and a 15s failsafe falls back to the old behavior.
Co-authored-by: tk-pkm111 <133480534+tk-pkm111@users.noreply.github.com>
Co-authored-by: aeonsong <aeonsong@users.noreply.github.com>
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
First-boot migrateActiveProfileIfMissing only listed ~/.hermes/profiles/*
and scored profiles/<name>/state.db. Default's real DB is ~/.hermes/state.db,
so a tiny named profile could be pinned after an update.
Always candidate default, score/pid-check it at HERMES_HOME, and do not
write active-profile.json when the winner is default.
Fixes#100576
- curly: brace all single-line if statements in profile-migration.ts and
profile-migration.test.ts (17 errors in CI check:lint)
- perfectionist/sort-imports: node:fs builtin import before vitest external
- padding-line-between-statements: blank lines after block statements
- prettier: normalize formatting (fmt script style) in the three touched files
All 967 electron project tests still pass.
Polish from antigravity review of the rebase resolution (GPT-OSS):
the previous comment said "BEFORE the first primaryProfileKey() /
primaryBackendIsRemote() read" but those two calls live at different
points — primaryBackendIsRemote() is the very next line, primaryProfileKey()
is inside the connection IIFE. Be explicit about which is where so a
future reader who moves one of them knows what to preserve.
Addresses teknium1's review (#64195) finding #1: the previous PR placed
the migration inside the connection IIFE, AFTER
`resolveRemoteBackend(primaryProfileKey())`. When the preference file
was missing, `primaryProfileKey()` resolved to 'default' and the remote
branch returned immediately without ever reaching the migration. Remote-
mode users got no migration at all.
Move the call site to the top of `startHermes()`, before the connection
IIFE that reads `primaryProfileKey()`. Both remote and local branches now
flow through this path before any profile-dependent resolution, so the
migration runs on first boot regardless of mode.
The inlined implementation is replaced with a thin wrapper that builds a
`MigrationDeps` bag and delegates to `migrateActiveProfileIfMissing` from
`profile-migration.ts`. No production behavior change beyond the call-
site move.
Tests added in a separate commit.
When active-profile.json does not exist (fresh install or first boot after
update), seed it from the best available signal so the Desktop launches
into the user's primary profile instead of always defaulting to "default".
Priority ladder:
1. Legacy ~/.hermes/active_profile (explicit CLI choice via hermes profile use)
2. Running gateway (gateway.pid with verified liveness + hermes identity check
via /proc/cmdline or ps -o args= to avoid PID recycling false positives)
3. state.db heuristics — hybrid recency×size score picks the primary workspace
(e.g. a 409MB coder DB beats a 28MB default DB even if touched at similar times)
The stored JSON includes _migrated:true for priority 3 (heuristic guess) so
the renderer can optionally surface a one-time notification. Priority 1 and 2
are higher-confidence signals and skip the flag.
The migration is a no-op once active-profile.json exists, and only writes
when a non-default profile is confidently identified — preserving the legacy
fallback-to-default behavior for single-profile users.
Fixes#64160 (active-profile half).
The spawn-time output tail (#93608) puts child stdout into flowing mode at
spawn. Both backend spawn paths in main.ts then await claimBackendChild
(whose Windows Get-Process probe cold-starts in 2-8s) and boot-progress IPC
BEFORE waitForDashboardPortAnnouncement attaches its stdout listener. Node
streams never replay consumed chunks to late listeners, so a READY sentinel
printed during that window was lost forever — the wait hit its 90s timeout
and a healthy backend was killed (deterministic on Windows, racy on
macOS/Linux; still firing on v0.21.0 incl. concurrent multi-profile boots).
Fix (belt and suspenders, both spawn paths — primary and profile pool):
- create the port-announcement promise immediately after spawn, before any
await
- new bufferedOutput option on waitForDashboardPortAnnouncement: after
attaching its own listener, waitForDashboardPort scans the output tail's
already-buffered text for the sentinel, making listener-attach ordering
irrelevant regardless of call-site shape
The readyFile path was already ordering-safe (it polls a file, not the
stream). Approach follows stale PR #60986 by @ParaWheeler, rebased onto the
output-tail/readyFile plumbing added since.
Fixes#60323
releaseBackendLock tree-kills only the Desktop's own backends
(backendConnectionState + backendPool). A messaging gateway launched by
the gateway-launcher desktop plugin via /api/gateway/start lives outside
those structures; on Windows its launcher (venv\Scripts\python.exe)
keeps the venv mandatory-locked, so the 15s release gate aborts the
hand-off BEFORE the venv-blocker scan — the scanner's pausable-gateway
exemption and the CLI updater's pause machinery never get their chance.
Delegate to `hermes gateway stop --all` (launcher + worker discovery
across every profile; gateway.pid records only the uv WORKER and
taskkill /T from the worker never reaches its parent). Per review, adds
the drain-semantics counterpart: every applyUpdates abort path
(lock-held, venv-blocked, probe-failed, updater-spawn-failed) restores
gateways via `gateway start --all`, so a failed update no longer
strands every profile's gateway stopped. Pure DI'd module + tests.
Salvaged from PR #76057 (issue #70337; overlap credit: #70477 by
@JonthanaHanh targeted the same symptom earlier via ZIP-dir preservation).
The pre-handoff teardown tree-kills only the backends the Desktop owns
(backendConnectionState + backendPool). The memory plugin's hindsight
daemon is spawned DETACHED off venv\Scripts\pythonw.exe, so it survives
the teardown, keeps venv files mapped, and either dead-ends the
venv-blocker scan with no in-app remedy or (pre-#74805 shim-only gate)
raced the updater into a half-updated venv.
Add a narrowly-scoped reap: kill only processes whose exe lives under
venv\Scripts (ordinal case-insensitive prefix — no PowerShell -like
wildcard hazards) AND whose cmdline references hindsight_api.main.
External holders (user terminals, unrelated scripts) are never killed —
scanVenvBlockers still reports them and the hand-off aborts, per existing
design. Selection logic is a pure DI'd module with unit tests.
Salvaged from PR #75477 (scoped per review: the narrow daemon kill; the
PR's generic every-exe kill was rejected as over-broad, its updateInFlight
half was superseded by #75778/#73822, and its generic lock-probe half by
the #74805 release gate + #99724 scanner classification).
Auto-fix perfectionist import/export ordering across the new
preview-annotate files, and extract the conversation-switch annotate
reset into a useCallback so the effect body carries no .current writes
(no-restricted-syntax ref-mirror rule).
Click or drag on the preview page to pin a numbered comment. Saving a pin only adds it to the stack. Add comments drops the crops and a short prompt into the composer and never sends the turn.
Capture takes the visible page then crops in bitmap space, because Electron's rect crop on guest webviews is empty on Windows and shifted on high-DPI.
Show a Restart backend action that kills the SSH serve before the local child
so reconnect cannot reuse a stale lockfile, instead of dumping raw IPC JSON.
Window X calls app.quit(); backendShutdown.finally() calls it again.
teardownSshConnection deletes the map entry before SSH exec kill, so
the second before-quit saw an empty map and exited while disconnect
was still running. Latch the in-flight teardown so the re-entrant quit
still waits.
ensureRegistryBackend() reuses the ambient primary descriptor for a
non-local, non-ssh registry primary, but the reused descriptor was
missing sharedRemote: true. The request router only appends
?profile=<profile> for sharedRemote backends, so Capabilities/Skills
requests went out unscoped and the gateway served the default profile
while a named profile was selected.
Add the flag and a regression test asserting the scoped URL.
The gateway's session-backed MCP OAuth flow (mcp.servers.oauth.start) binds
its browser-callback listener on the BACKEND machine's 127.0.0.1. When the
Desktop app connects to a remote backend (SSH/Tailscale), the user's browser
resolves that loopback to the user's machine, the redirect dies, and every
OAuth catalog server (ClickUp, Hospitable, ...) fails in-app with no working
path — the exact topology from the 'MCP Recurring erros' support thread.
Fix mirrors the Desktop's native gateway login (native-oauth-login.ts):
- gateway: mcp.servers.oauth.start accepts client_redirect_uri (loopback-only,
RFC 8252-style validation); when supplied no gateway listener is bound and
the OAuth redirect_uri pins to the client's listener.
- gateway: new mcp.servers.oauth.callback RPC relays the client-captured
code/state into the flow; state verification stays in
DashboardOAuthFlow.deliver_callback (constant-time compare, replay-safe).
- desktop: mcp-oauth-callback-ipc.ts hosts a one-shot 127.0.0.1 listener in
the main process (hermes:mcp-oauth:listen/wait/cancel via preload bridge).
- desktop: hermes-bots mcp-setup.tsx prefers the client listener for local
AND remote backends, falling back to the legacy gateway-listener flow on
older gateways (feature-detect via start rejection).
- docs: remote-host MCP OAuth section documents the automatic Desktop path.
Validation: 19 new gateway tests (validator allowlist, listener skip, relay
accept/reject/replay) — sabotage-verified; 5 new desktop tests against a real
ephemeral listener; E2E through the real session registry + flow bridge with
a stubbed provider probe; tsc electron+renderer builds clean.
`pumpStreamToFile` opened the user-chosen destination with
`fs.createWriteStream`, which truncates the target the instant it opens,
and its error path then unlinked that same path. When a user picked an
existing file in the Save dialog (and confirmed the overwrite) and the
gateway dropped mid-stream, the original was gone: truncated first,
deleted second, with nothing written in its place. The data-URL
compatibility fallback (`saveGatewayFileViaDataUrl`) had the same class
of bug via `fs.promises.writeFile`, which truncates before the write
completes.
Both paths now go through one failure-atomic primitive. Bytes land in a
short, randomly named sibling temp file (`.hermes-download-<hex>.part`,
same directory so the final step is a same-volume rename), created with
`flags: 'wx'`, and are renamed onto the destination only after the whole
body has been written and the descriptor released. The destination is
never opened before that point, so a failed download leaves whatever was
there untouched.
- Ownership-gated cleanup: the temp file is unlinked only after the
stream's 'open' event proved THIS operation created it. An exclusive
create that fails before open (EEXIST collision, EACCES, missing
parent) never removes a file that belongs to someone else.
- `WriteStream.close(cb)` rather than `end(cb)` before renaming: `end`'s
callback fires on 'finish' while the fd may still be open, and Windows
refuses to rename a file with an open handle. Falls back to `end` for
stream shapes without `close`.
- The failure path waits for 'close' (bounded by a 2s grace period)
before unlinking, for the same reason: `destroy()` releases the fd
asynchronously and an unlink racing the open handle would leak the
`.part` file on Windows.
- A rename failure (destination locked, permissions) removes the owned
temp file and rejects; nothing is left behind.
- Fixed-length temp name so a long user-chosen filename cannot push it
past the filesystem's name limit.
- `fsPumpDeps()` is the single production deps factory (`'wx'` create,
`fs.promises.rename`, `fs.promises.unlink`); `writeBufferToFile()`
routes the data-URL fallback through the same pump. `PumpDeps` gains
`rename` and a `tempPathFor` test seam.
Tests. Fakes: temp-then-rename on success, close-before-rename ordering,
the regression itself (destination neither opened nor unlinked when the
response fails mid-stream), write-error cleanup, close-before-unlink
ordering, rename-failure cleanup, pre-open EEXIST leaves the colliding
file alone, `writeBufferToFile` success and post-open write failure, the
temp-name length bound, and the `main.ts` wiring. Real filesystem
(`gateway-file-download.fs.test.ts`, exact production deps in a scratch
dir): completed download replaces the destination with no temp left;
mid-stream failure leaves the pre-existing destination byte-for-byte
with no `.part`; failure into a fresh name leaves nothing; seeded temp
path survives a pre-open EEXIST with no rename; rename failure (directory
at the destination) cleans the owned temp; data-URL fallback success and
missing-directory failure.
Adds the contributor email mapping required by the attribution check.
Fixes#96597
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015u8q2pHVPZmxpSrgkt94jC
Composer drag added renderer CSS-pixel deltas onto a window AppKit
clamps to the current display, so the bar could not follow the cursor
onto a second monitor (and drifted on mixed-DPI Windows). Track the OS
cursor in main and lift that clamp.
Co-authored-by: Biotrioo <biotrioo@protonmail.com>
The renderer pings each pool backend every 60s (`hermes:backend:touch` →
`touchPoolBackend` → updates `lastActiveAt`). The LRU eviction cap used a
keepalive-fresh window of 90s — only 1.5× the ping cadence — to decide
whether a backend was "plausibly still alive". One missed or delayed ping
pushed a live backend past the threshold and the cap-driven eviction killed
the active profile's backend mid-session, restarting the gateway and
re-minting runtime ids. On WSL2, where the renderer→Electron IPC roundtrips
through 9p, brief 9p hiccups commonly stretch a ping to seconds of observed
silence, producing the ~80–90s exit / ~2 min cycle reported in #95189
(122 gateway starts on 2026-08-26 alone, driving renderer OOM via reconnect
churn at ~5GB/day).
Widen POOL_KEEPALIVE_FRESH_MS to 4 minutes (3× ping cadence + IPC stall
headroom, still bounded well below POOL_IDLE_MS=10min). Backends with one or
even two missed pings are now spared; truly idle backends (multiple lapses,
minutes idle) remain eligible for eviction by the cap and the idle reaper.
The constant is also overridable via HERMES_DESKTOP_POOL_KEEPALIVE_FRESH_MS
to make this tunable without a rebuild.
(connectionId, profile) scope, but it was only wired into hermes:connection,
hermes:connection:for, and the power-resume pool rebuild. Five other real
call sites still invoke ensureRegistryBackend()/ensureBackend() directly:
the media-protocol connection resolver, the terminal-pane backend resolver,
the ~5s roster-enumeration probe, the connections update-all dispatch, and
dispatchRegistryApiRequest (every registry-scoped hermes:api REST call).
ensureRegistryBackend() has a genuine await-then-check race
(reuseMatchingPrimarySshBackend before the pool entry check-and-set), so any
of these five racing a guarded dial for the same scope can each bootstrap
their own SSH tunnel / remote dashboard for the same connection — exactly
the symptom class #90812 was written to prevent, just reached through an
unguarded path instead of two renderer windows.
Left the ensureRegistryBackend() self-call inside its own dispatch-time
health-probe reconnect branch (registryDispatchRevalidation) unguarded —
that recursive path has its own coordination semantics and touching it
without modeling reentrancy against the newly-guarded entry points here is
a separate, riskier change better done on its own.
Re-implements the #93042 main.ts wiring against current main (post-#94724
drift), making the extracted managed-ssh-update engine reachable:
- ManagedConnectionUpdateGate instance + owner-only recovery journal at
DESKTOP_MANAGED_SSH_RECOVERY_PATH (read/write/persist/mark/clear with
strict record validation).
- IPC: hermes:connections:update-managed (requestManagedSshUpdate with
correlation-id claim + in-flight dedupe); update-all's ssh rows now route
through the transactional drain/update/restore lifecycle instead of
POSTing the remote backend updater.
- Gate enforcement at every dial/mutate seam: bootstrapSshConnectionInner
(pre-dial + publication fence with exact-serve rollback via
rollbackSshBootstrapResult), resolveRemoteBackend, ensureRegistryBackend,
saveRegistryConnection dial-field edits, connections:remove, and
primary-routing mutations (set-primary, set-launch-mode,
connection-config save/apply, profile:set) via
assertCanMutateManagedPrimaryRouting.
- Scope capture/drain/restore drivers: captureManagedSshScopes (pool +
primary discovery, bootstrap fence join), drainManagedSshScope (exact
identity-re-proof termination, no-kill forward recovery),
ensureManagedSshBackend(AtKey)/restoreManagedPrimarySshBackend restores,
openManagedSshUpdateTransport for serve-less connections.
- Startup recovery (resumeManagedSshRecoveries before createWindow) and
before-quit join of in-flight update/recovery operations BEFORE the SSH
coordinator is sealed, so restore dials are not refused during quit.
- Extended sshConnections state (spawnNonce/creationTime(Ns)/startedAt/
hermesPath/hermesHome/pythonPath/remoteProfile/registryConnectionId/
primaryRegistryScope) so drain can prove the exact serve it owns;
bootstrap coordinator entries carry metadata for the update fence;
persistSshConnectionToken mirrors tokens per managedSshTokenPersistencePlan.
- preload/global.d.ts: connections.updateManaged +
DesktopManagedConnectionUpdateResult/Receipt types.
Renderer UI (fleet-updates store, about-settings, system.ts ProfileScope
plumbing, i18n) intentionally NOT wired — it belongs to the deferred fleet
rollout UI and follows separately.
Wiring re-implemented against current main; design from #93042 by @andrexibiza
tsc -p apps/desktop clean; electron project 1924/1924 passed (133 files,
incl. 131/131 across the three engine suites); eslint clean on touched files.
A torn renderer bundle (update replaced the app while its files were
locked, e.g. antivirus or a still-running instance) loads index.html
fine and then dies on the first lazy import — a white screen with only
a desktop.log line. A main-frame load failure (missing index.html,
blocked file) was likewise log-only.
- resolveRendererIndex() already detects torn bundles; the primary
window now refuses to load one and shows a visible repair page
(error code, missing assets, 'hermes desktop --force-build', Reload)
instead of a blank window.
- did-fail-load on the main frame now gets bounded auto-reload through
the shared rolling reload budget (transient failures self-heal) and,
once the budget is exhausted, surfaces the visible error page.
ERR_ABORTED and sub-frame failures stay log-only, and helper windows
(OAuth/portal) keep their log-only policy (opt-in via
reloadOnFailedLoad).
Regression tests cover the policy decisions (reload / abort /
budget-exhausted surface), budget sharing with render-process-gone,
and the error page content + data: URL loading.
new BrowserWindow({ icon }) and app.dock.setIcon() decode the icon file
synchronously on the main process and throw on undecodable bytes. The
icon ladder was resolved with statSync().isFile(), which only proves a
file exists — a truncated or zero-byte PNG inside a packaged app.asar
(interrupted electron-builder run, partial copy) killed the main process
inside createWindow(): the window never appeared, running turns lost
their renderer, and the desktop log showed 'Uncaught exception: Error:
Failed to load image from path .../app.asar/public/apple-touch-icon.png
at createWindow'.
Resolution now runs through a decoding probe (nativeImage.createFromPath
must yield a non-empty image); a candidate that exists but does not
decode is skipped like a missing one, so the app falls through to the
next rung or starts with the platform default icon instead of dying.
The ladder and probe live in a pure module (electron/app-icon.ts) so
precedence is unit-testable without a running Electron app; window
factories re-resolve per call exactly as before.
Regression tests cover: skip-first-undecodable, all-fail -> undefined,
first-pass wins, missing/empty/directory rejection, and the unchanged
mac/Windows precedence ladder.
Co-authored-by: brooklyn! <brooklyn.bb.nicholson@gmail.com>
'Make primary' on a registered remote/cloud/ssh gateway only rewrites
connections.json — the v1 config.mode stays 'local', so startHermes()
resolved no remote route and spawned a loopback 'hermes serve' the
desktop never uses (full MCP set duplicated, port squat, respawn on
poll). resolveDesktopRemoteRoute gains a lowest-precedence registry-
primary rung (source: 'registry', existing v1/env/profile precedence
untouched), and globalRemoteActive() now recognizes a remote registry
primary so local-entry routes force pooled local children instead of
delegating into a primary that dials remote. A 'local' registry
primary still resolves null — genuinely-local desktops unchanged, and
local-profile secondaries keep their forced-local pooled backends.
- normalizeRegistry now preserves every malformed entry (unknown kind,
url-less remote/cloud, host-less ssh, mangled non-object items, and
any entry whose normalization throws) under a capped 'quarantined'
key that survives write cycles — healthy entries keep loading and
user data is never silently deleted.
- A whole-file parse failure preserves the original bytes in a
connections.json.corrupt-<ts> sidecar BEFORE the drift reconciler or
a save can overwrite the file with the degraded local-only registry.
- Loads log a quarantine notice and sanitizeConnectionsRegistry
surfaces reason+label summaries (never raw entries/token envelopes).
Two registered basic-auth gateways shared the single
persist:hermes-remote-oauth cookie jar, so signing in to gateway B
evicted gateway A's session cookies (Chromium jars ignore the port) and
A's cookie was silently presented to B on every request. Non-primary v2
registry remotes with cookie auth now ride a per-connection partition
(persist:hermes-remote-oauth:conn:<id>) resolved at the jar boundary;
the registry primary, v1 remote, cloud cascade, and portal flows keep
the legacy shared jar so upgrades do not sign anyone out. Fail closed:
a connection's requests can never see another connection's cookies.
reconnectGateway()'s in-flight lock lives at renderer module scope, so it
only dedupes reconnects inside ONE window. Two windows racing the same
wake both invoke the main-process backend ensure IPC, and for a pooled
SSH connection the loser of the pool-entry race could bootstrap a
duplicate remote backend (two tunnels, two remote serve processes).
Electron main is the single owner of backend lifecycles, so the claim
now lives there: BackendDialClaims keys in-flight dials by the pool
scope key from backendScopeKey(connectionId, profile) — the composite
identity seam wave-1 #93189 established for effective-identity reuse.
'hermes:connection' and 'hermes:connection:for' route through
backendDialClaims.run(), so concurrent renderer dials for one scope
coalesce onto one spawn and the second caller receives the first's
result. A claim exists only while its dial promise is unsettled: both
outcomes release it, a failed dial is never cached (fail closed, not
latched), and a synchronously-throwing dial rejects the claim instead
of escaping the seam.
The #93910 resume rebuild re-dials retired pool keys through the same
claim (redialPoolBackendAfterResume + new parseBackendScopeKey), so a
resume-driven rebuild and a concurrent renderer reconnect also coalesce
instead of racing.
Live-confirmed on the Phase B build: hermesDesktop.connections.save()
succeeds and the registry on disk gains the row, but the switcher menu
(fed by the renderer $connectionsRegistry snapshot) keeps painting the
stale list until reload. remove() already broadcasts
hermes:connections:changed; save() only did so on the dial-material-edit
branch, so a brand-new connection or a label rename never reached the
switcher's onChanged re-pull (or any other window).
Fix at the publish seam only: saveRegistryConnection now broadcasts a new
'saved' reason for every successful save that isn't a dial-material edit.
'saved' is a pure registry-refresh signal — the use-gateway-boot listener
explicitly ignores it (nothing moved, so no dispose/redial/forget), while
the switcher's existing onChanged listener re-pulls the snapshot.
Tests:
- electron/hardening.test.ts pins both broadcast branches in
saveRegistryConnection (source-assertion pattern; main.ts has no exports).
- connection-switcher.test.tsx mirrors the live repro scenario
(/tmp/mg-ab/w2_95393.py): menu before save lacks the row, Electron's
'saved' push arrives, menu after — without reload — shows it.
Partial cherry-pick of PR #95007 (weismanfamily). Surviving scope:
- electron connection-registry: registrySourceOwnsPrimaryBackend() —
descriptor-level proof that a registry-scoped request names the
already-running primary backend, wired into ensureRegistryBackend as
the generic (non-SSH-fingerprint) primary-owns short-circuit so a
cloud/url registry primary cannot spawn a second isolated server.
- store/connections: waitForInitialConnection() before the boot-time
source restore, so the sidebar registry cannot dial the preferred
source a second time while the identical primary backend is still
publishing its connection identity.
Dropped scope (superseded on main): the renderer routing half —
primaryConnectionId plumbing, primaryOwnsAgent short-circuits in
requestGatewayForAgent/openGatewayForAgent/ensureGatewayForAgent
(main has isPrimaryRegistryRoute via 1ec32e738), the 3-arg
setPrimaryGateway boot wiring (main has setPrimaryGatewayConnection),
the use-session-list-actions stampConnectionOwner (row stamping lands
via #94656/#94901), and the wiring.tsx bare-profile promotion commit
65d106e41 (main's knownSessionOwner covers it).
Original-PR: #95007
Dropped-scope: renderer routing half of d1c0fb093; all of 65d106e41
teardownSshConnection closed the tunnel and SSH transport but never
killed the detached serve --isolated process. Spawn uses setsid/nohup,
so the backend reparents to pid 1, keeps state.db open, and accumulates
across Cmd+Q. Reuse cleanupStale via disconnect while SSH can still
exec, sequence remote kill before close, and seal the bootstrap
coordinator so reconnect during a prevented first quit cannot respawn.
The quit race is 6s to cover cleanupStale's 5s wait-for-exit loop.
Avoid opening a second SSH lifecycle when a migrated registry request targets the same primary/default backend already booted through the legacy route. Compare effective SSH configuration for representation-only drift, treat empty and default as the same root profile, and keep named profiles isolated.