Commit Graph

4143 Commits

Author SHA1 Message Date
686f6c61 0ecbf1ce91 fix(desktop): snapshot Close All pane ids before persist-close
closeSessionTile can rewrite the layout tree. Iterating the live
group array then skips tiles. Copy the list first.
2026-08-26 09:47:23 -07:00
686f6c61 f1755cc1c5 fix(desktop): persist Bot Mode Close All session tiles
Close All only dismissed layout-tree panes. Bot tiles live in the
shared __bots_workspace__ bucket, so clicking a bot or swapping
profiles rehydrated the closed tabs. Persist-close those tiles
before dismissing the rest of the strip.
2026-08-26 09:47:23 -07:00
funky-xamarin 9ea7a37cc1 fix(desktop): restore live group after bot open failure 2026-08-26 09:47:23 -07:00
funky-xamarin fbd149035b test(desktop): verify group-to-local-bot handoff in Electron 2026-08-26 09:47:23 -07:00
funky-xamarin 35ee27b1c6 test(desktop): add RED group-to-local-bot E2E 2026-08-26 09:47:23 -07:00
funky-xamarin 7c255dbbbf test(desktop): focus group-to-local-bot handoff on close-callback RED
Rewrite the handoff proof to ≤150 lines on the hermes-bots VM seam so
base fails because the registered group closer is not invoked, not a
missing helper marker. Cover BotRow/Active Now local close-before-open,
remote no-dismiss, and no-group/old-host safety.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 09:47:23 -07:00
funky-xamarin 2fa39fac8c test(desktop): make group-to-local-bot handoff proof behavioral
Exercise real BotRow and Active Now open handlers so close-before-open
and remote stay-put are load-bearing, not source-regex false greens.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 09:47:23 -07:00
funky-xamarin f72786f764 fix(desktop): close group main tab when opening a local bot chat
Local BotRow and Active Now opens now share a handoff that retires the
selected group workspace before canonical chat open, so the bot chat can
own the main surface after leaving a group tab.
2026-08-26 09:47:23 -07:00
Teknium 9aa7530f7b style(desktop): sort oauth-partition import (lint) 2026-08-26 08:40:07 -07:00
Teknium 697087f2eb fix(desktop): SIGKILL-escalate the owned SSH backend when it survives the graceful quit wait (#91668 remainder)
The #95085 quit teardown kills the owned serve --isolated before the
SSH tunnel closes, but a backend mid-turn (in-flight LLM call, live MCP
children) can ride out SIGTERM past cleanupStale's 5s graceful wait.
The old code then gave up (threw, kept the lockfile) and before-quit's
6s race closed SSH anyway — reparenting the still-running serve to
pid 1: the reported leak, now specific to quit-during-active-turn.
Escalate to kill -9 with a confirmed-exit wait; only an unkillable pid
(D-state, permissions) still throws and preserves the lock record so
the next connect's reap pass retries.
2026-08-26 08:40:07 -07:00
Teknium 8497edb2ac test(desktop): pin the activation-epoch guard against mid-handshake profile switches (#92434 close-candidate)
Reproduces the reported Bot↔Default switch shape at the gateway.ts
activation seam: a switch-back that lands while the outgoing switch's
WS handshake is still pending keeps the route, the late-completing dial
neither steals the foreground nor breaks its socket, and re-activating
the bot works without an app restart. The guard (activation epochs +
open-socket-publish, landed via #89622/#92265/#81094) already prevents
the reported permanent break; this pins it so it cannot regress.
2026-08-26 08:40:07 -07:00
Teknium 1e9a12a71f fix(desktop): stop spawning loopback serve children when the registry primary is remote (#91564, #90316)
'Make primary' on a registered remote/cloud/ssh gateway only rewrites
connections.json — the v1 config.mode stays 'local', so startHermes()
resolved no remote route and spawned a loopback 'hermes serve' the
desktop never uses (full MCP set duplicated, port squat, respawn on
poll). resolveDesktopRemoteRoute gains a lowest-precedence registry-
primary rung (source: 'registry', existing v1/env/profile precedence
untouched), and globalRemoteActive() now recognizes a remote registry
primary so local-entry routes force pooled local children instead of
delegating into a primary that dials remote. A 'local' registry
primary still resolves null — genuinely-local desktops unchanged, and
local-profile secondaries keep their forced-local pooled backends.
2026-08-26 08:40:07 -07:00
Teknium 62e2d6e1e4 fix(desktop): quarantine malformed connections.json entries per-entry instead of dropping or nuking the registry (#94246)
- normalizeRegistry now preserves every malformed entry (unknown kind,
  url-less remote/cloud, host-less ssh, mangled non-object items, and
  any entry whose normalization throws) under a capped 'quarantined'
  key that survives write cycles — healthy entries keep loading and
  user data is never silently deleted.
- A whole-file parse failure preserves the original bytes in a
  connections.json.corrupt-<ts> sidecar BEFORE the drift reconciler or
  a save can overwrite the file with the degraded local-only registry.
- Loads log a quarantine notice and sanitizeConnectionsRegistry
  surfaces reason+label summaries (never raw entries/token envelopes).
2026-08-26 08:40:07 -07:00
Teknium 31250da505 fix(desktop): key cookie-auth session partitions on connection identity, not auth mode (#92183)
Two registered basic-auth gateways shared the single
persist:hermes-remote-oauth cookie jar, so signing in to gateway B
evicted gateway A's session cookies (Chromium jars ignore the port) and
A's cookie was silently presented to B on every request. Non-primary v2
registry remotes with cookie auth now ride a per-connection partition
(persist:hermes-remote-oauth:conn:<id>) resolved at the jar boundary;
the registry primary, v1 remote, cloud cascade, and portal flows keep
the legacy shared jar so upgrades do not sign anyone out. Fail closed:
a connection's requests can never see another connection's cookies.
2026-08-26 08:40:07 -07:00
Zeus-Deus fd565c80e9 feat(desktop): fleet profile rail — every registered gateway's agents on one strip
With several gateways registered, the Sessions profile rail only ever showed
the active gateway's profiles; reaching a bot on another machine meant a
gateway switch first, then a click on the rail that appeared afterwards. Bot
Mode (#91134) and Capabilities already read the union agent roster; the rail
is now its third consumer.

- Every registered gateway's profiles sit on the one strip, in registry order
  (This device first, then by label), each group headed by that gateway's
  kind glyph. The active gateway's squares are unchanged; the others are
  "at rest" (dimmed) with tooltips/accessible names qualified by machine
  (`inbox · Homelab`), so same-named profiles never read alike.
- Clicking an at-rest square performs the same dial → commit → re-home as
  the statusbar switcher, landing on that exact (gateway, profile):
  `selectConnection(id, { profile })`. The spinner sits on the clicked
  square; the previous source stays painted until the target answers.
  Groups keep their slots whichever gateway is active, so a square never
  moves under the pointer that clicked it.
- Right-click on an at-rest square: Switch to / Color / Rename / Edit
  SOUL.md / Delete, executed on the owning gateway (renameProfile,
  getProfileSoul and updateProfileSoul accept the same scope deleteProfile
  already had); the delete confirmation names the machine. The legacy
  per-profile "Connect to a remote host…" item is hidden on multi-gateway
  setups, where the rail shows machines directly.
- Unreachable gateways keep their squares with an amber dot on the glyph;
  two registrations of one backend collapse to one group; past thirteen
  squares across the fleet the strip condenses into a menu sectioned by
  gateway. Roster is fetched on mount / focus / registry change only — no
  periodic fleet polling.
- Single-gateway Desktops render exactly as before: no roster fetch, same DOM.

Also fixes a boot race the e2e surfaced: initializeConnectionsRegistry()
"restored" the launch-mode source over a switch the user had already made
while boot was settling (same class as #91047). The restore now yields when
a switch is pending or already landed.

Tests: pure grouping (fleet-rail.test.ts), rail component fleet mode
(profile-rail-fleet.test.tsx), store (explicit profile pick; restore yields),
and a Playwright e2e (fleet-profile-rail.spec.ts) that boots Desktop with two
REAL backends — the local one plus a second `hermes serve` registered as a
remote URL connection — and verifies layout, a real re-home, gateway-scoped
actions, and order stability.

Docs: multi-connection-desktop.md describes the fleet rail.

Refs #89304, #92384, #91047, #94724

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 08:39:45 -07:00
Teknium dd0aae4173 fix(desktop): paint stored Bot Chat history immediately instead of stranding the wake on an unsatisfiable profile gate
Fixes #89843. On a shared-remote connection every profile is served
through the primary socket, so waitForFocusedSessionHydration's
profileMatches gate could never become true — a bot chat whose stored
transcript painted within seconds still burned the whole 20s hydration
budget and then stranded the pane with 'Timed out loading <bot>'s
session history'.

The wake now resolves paint-first: once the stored transcript is painted
on exactly the target session, the content is its own proof — the pane
opens immediately and a subtle 'Syncing…' badge (new $hydrationSyncProfile
atom + ChatSyncBadge) shows until the profile gate catches up in the
background. Fail-closed everywhere content is not its own proof: a
superseded/conflicting concurrent wake still rejects, and an
expected-empty chat still waits for the full runtime gate.
2026-08-26 08:39:45 -07:00
etzelvon dbbed456c1 style(desktop): satisfy curly + padding-line lint for guarded switch
- braces for single-statement if(confirmNotificationId) guards
- blank line before return in staleness branch (padding-line-between-statements)
2026-08-26 08:39:45 -07:00
etzelvon a72d6d6338 fix(desktop): address review — staleness guard, neutral confirm fallback, clarify pending return
- staleness guard in applyConfirmedSwitch: bail and dismiss if
  current model/session no longer matches snapshot this warning was
  created for, preventing stale Confirm from clobbering newer pick
  (Enough1122 review #92492)
- neutral fallback for missing confirm_message: 'Confirm this model
  switch?' instead of modelSwitchFailed
- document selectModel boolean: false means not-applied (pending
  confirmation or failed) — pending already shows warning, not an error
- test mock: add dismissNotification mock for guarded flow

Addresses https://github.com/NousResearch/hermes-agent/pull/92492#issuecomment-5387151656
2026-08-26 08:39:45 -07:00
etzelvon 7450266a3a fix(desktop): confirm guarded model switches instead of snapping back
Desktop composer ignored gateway's confirm_required response for
contributor / expensive models. It painted the target optimistically
then invalidated model-options and refetched the still-active session,
so muse-spark-1.2-contributor appeared to instantly snap back to
gpt-5.6-sol with no explanation.

Now handle the confirmation protocol: on confirm_required rollback the
optimistic state, surface the backend's confirm_message as a warning
notification with a Confirm action, and on Confirm retry config.set with
confirm_expensive_model:true. Preserves data-training consent, keeps
the fix session-scoped and test-covered.
2026-08-26 08:39:45 -07:00
konsisumer b3a2065ff3 fix(installer): target Windows updater shim kills 2026-08-26 07:59:01 -07:00
hermes-seaeye[bot] 277899f10c fmt(js): npm run fix on merge (#95614)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 14:35:33 +00:00
beplee ec8ca8f2cb fix(desktop): re-bind open pane to rebuilt runtime after model switch
A mid-conversation model/provider switch rebuilds the agent runtime. The
rebuilt runtime emits session.info (and all later events) under a NEW
explicit session_id while the pane still holds the dead one as its
active id — isActiveEvent is false for the same conversation from that
moment on, so view-scoped updates stop and the chat freezes until a
full resume (#93942 scenario B; backend even logs 'client should resume
the stored session', but the client never does).

Fix: when a session.info event lineage-matches the selected conversation
(sessionMatchesStoredId over stored_session_id) but carries a different
runtime id, adopt the new runtime id as the active session id — keeping
the durable selection untouched — so every subsequent isActiveEvent gate
keeps matching without a resume. Guarded: the old runtime must show no
live turn (not busy/awaiting/streaming) or the adoption is refused, so
an overlapping manual switch can never split one conversation across
two panes.

The existing compression-rotation path does not cover this case: it
fires when the SAME runtime's stored id rotates, while a rebuild
produces a NEW runtime with a NEW stored id.

Together with #94255 (tile reconcile on sessions.changed), closes
#93942.

Regression tests verified failing pre-fix on 41447a6d70.
2026-08-26 07:28:10 -07:00
beplee 7cfed25370 test(desktop): behavior tests for tile reconcile + signature pruning (#94255 review)
Enough1122 review points on #94255, all addressed:

1. Source-grep Python tests replaced with real vitest behavior tests:
   - a tile whose stored transcript gained a background delivery IS
     reconciled (updater invoked, correct stored id fetched)
   - an unchanged transcript is skipped entirely (no updater call —
     the signature gate proven, not asserted by regex)
   The structural smoke test in test_bots_chat_live_append.py stays as
   a cheap drift alarm; the contract now lives here.

2. Shared-sequence latest-wins semantics documented in the reconcile
   docstring (review point 2).

3. Signature map pruning: a tile closed/superseded mid-read now deletes
   its signature entry instead of leaking one map slot per ever-opened
   tile for the app lifetime (review point 3).

4. Test harness: typed updateSessionState mock via Parameters<> instead
   of the {}-as-state cast; no more silent spread corruption.

No production behavior change beyond pruning: 18/18 hook suite green,
typecheck clean.
2026-08-26 07:28:10 -07:00
beplee 3669fa3095 fix(desktop): satisfy lint — read busy atom directly in tile reconcile
CI caught two lint issues in the tile-reconcile path:
- no-restricted-syntax: don't mirror the $busy atom into a ref via
  useEffect (stale-read hazard); reconcileTileTranscripts now receives a
  live getter view so the loop reads the current value at tick time.
- react-hooks/exhaustive-deps: add updateSessionState to the
  sessions.changed effect deps (stable useCallback from
  useSessionStateCache, so no extra re-subscription).

No behavioral change: regression tests 2/2, adjacent hook suite 58/58,
typecheck clean.
2026-08-26 07:28:10 -07:00
beplee db8ff4eb75 fix(desktop): reconcile workspace-tile transcripts on sessions.changed
Bot canonical chats open as workspace tiles (workspaceMode: 'bots') and
are deliberately hidden from $sessions/$messagingSessions, so the
sessions.changed transcript refresh skipped them twice over: it covers
only the main pane's selection, and its resolveSession() bails on hidden
sessions. A background delivery (bot-to-bot DM via bot_relay.deliver, a
cron run's output, another machine) therefore never reached an open bot
chat — the roster updated but the pane stayed stale until remount
(#93942 scenario A).

Fix: the sessions.changed tick now also reconciles every visible
workspace tile through a dedicated signature-gated path. Each tile
carries its own stored↔runtime id pair so no resolution step is needed;
per-tile signatures make no-change ticks free; busy tiles are skipped
(their own stream owns the view); closed/superseded tiles discard their
in-flight read.

Slice 1 of 2 for #93942 (scenario A only). Scenario B (stream re-key
after mid-conversation model switch) follows separately.

Fixes part of #93942
2026-08-26 07:28:10 -07:00
funky-xamarin ef2710d1f4 fix(desktop): refresh hidden Bot Chat transcripts 2026-08-26 07:28:10 -07:00
Hermes (Lift-Off agent) 1396a30cd7 test(desktop): cover canonical activity stale-runtime split 2026-08-26 07:28:10 -07:00
Teknium 254af55728 fix(desktop): force session.resume on explicit bot-switch open so Bot Chat never paints a stale cached transcript (#93604)
The post-open surface-health check in host.openSession trusts any
non-empty cached transcript ($messages.length > 0), so re-opening a
bot's canonical chat after switching bots could pass the check while
painting a stale snapshot kept by the session-states cache — skipping
requestSessionResume and leaving old messages on screen until an app
restart.

Add an opt-in forceResume flag to PluginOpenSessionOptions, honored only
alongside awaitHydration, and set it on the one explicit bot-switch open
path (openStoredBotChat, which serves openBotCanonicalChat and stored
opens). Resume is cheap and idempotent per the route-resume effect's own
contract, so the extra request is a no-op when the transcript is already
fresh. All other navigation paths keep the existing heuristic.
2026-08-26 07:28:10 -07:00
Teknium bc21808e6a fix(desktop): legacy group members stored under display names seat their real bot once, not as ghosts (#92794)
Older builds persisted group members with a FRIENDLY name as the
descriptor's `name` (e.g. '大司命' for slug 'taiyi'), some predating
connection scoping entirely (no connectionId). Key matching alone seated
those descriptors as ghosts NEXT TO their own live rows ('4 bots' in a
2-bot room, reproduced live), and any path passing ghost identity onward
targeted a profile that does not exist on disk.

groupChatMemberBots now normalizes stored descriptors before seating:
an unmatched descriptor re-tries by case-drifted slug or friendly name
(botFriendlyNames precedence) against rows on its own connection —
connectionless pre-scoping descriptors match local rows only. The next
persistence pass rewrites storage to slugs, so the repair self-heals.
Unresolvable descriptors still seat as degraded ghosts and are never
used as profile targets.
2026-08-26 06:26:09 -07:00
toprakeker 8624c1e8f7 fix(desktop): reject SSH backends with replaced runtimes 2026-08-26 06:24:30 -07:00
hermes-seaeye[bot] 86ae906e88 fmt(js): npm run fix on merge (#95511)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 11:57:06 +00:00
miha 1bd5da3ac6 fix(desktop): skip macOS TCC-protected media dirs in git repo scan
The sidebar's home-dir repo crawl descends into ~/Pictures, ~/Music,
~/Movies and ~/Public — and into Photos/Music library packages — which
triggers Photos, Media Library and Files & Folders permission prompts
attributed to Hermes.app. Because the app is ad-hoc signed and re-signed
on every self-update (#49110), macOS drops all TCC grants after each
update, so these prompts re-fire every time.

Skip the media folders as direct children of a search root (nested dirs
like ~/dev/Music are ordinary and still scanned; an explicitly passed
root is still walked), and skip Apple library packages
(*.photoslibrary/*.musiclibrary/*.tvlibrary/*.aplibrary) at any depth.

Partial mitigation for #49110 / #52010: removes the Photos and Media
Library prompts entirely; the identity reset itself needs Developer ID
signed release artifacts (tracked in #49110).
2026-08-26 04:51:48 -07:00
Teknium bb3421bf25 fix(desktop): single-owner backend dial claim in Electron main (#90812)
reconnectGateway()'s in-flight lock lives at renderer module scope, so it
only dedupes reconnects inside ONE window. Two windows racing the same
wake both invoke the main-process backend ensure IPC, and for a pooled
SSH connection the loser of the pool-entry race could bootstrap a
duplicate remote backend (two tunnels, two remote serve processes).

Electron main is the single owner of backend lifecycles, so the claim
now lives there: BackendDialClaims keys in-flight dials by the pool
scope key from backendScopeKey(connectionId, profile) — the composite
identity seam wave-1 #93189 established for effective-identity reuse.
'hermes:connection' and 'hermes:connection:for' route through
backendDialClaims.run(), so concurrent renderer dials for one scope
coalesce onto one spawn and the second caller receives the first's
result. A claim exists only while its dial promise is unsettled: both
outcomes release it, a failed dial is never cached (fail closed, not
latched), and a synchronously-throwing dial rejects the claim instead
of escaping the seam.

The #93910 resume rebuild re-dials retired pool keys through the same
claim (redialPoolBackendAfterResume + new parseBackendScopeKey), so a
resume-driven rebuild and a concurrent renderer reconnect also coalesce
instead of racing.
2026-08-26 04:49:56 -07:00
Teknium 3123624c07 fix(desktop): revalidate pooled remote/SSH backends on power resume (#93910)
After macOS sleep/resume, pooled remote SSH descriptors kept serving dead
tunnels: a remote entry has no child 'exit' to clear it, the renderer
keepalive spares it from the idle reaper, and the wake-path nudges from
b90289b04/febed060a only re-drive the PRIMARY renderer socket. The
background failure-streak policy needs several probe rounds before it
drops a descriptor, so the Bots pane showed 'Gateway offline' long after
the network was back.

New revalidateSuspectPooledRemoteBackends(): on resume every pooled
remote is suspect — probe each once (bounded by
REMOTE_LIVENESS_TIMEOUT_MS), retire the dead ones immediately (pool
entry + SSH bootstrap + tunnel/master teardown) and rebuild them through
the caller's dial path, while healthy descriptors are left untouched. A
failed retire skips the rebuild (never dial on top of an installed
descriptor); a failed rebuild is logged and left to the renderer's
normal reconnect — the sweep never throws.

attachPowerResumeRemoteRevalidation() wires the sweep to the Electron
powerMonitor 'resume'/'unlock-screen' seam with a 15s holdoff so the
near-simultaneous macOS wake signals coalesce into one sweep and can
never form a hot loop; overlapping kicks additionally join the one
in-flight sweep via the existing RemoteRevalidationCoordinator.
2026-08-26 04:49:56 -07:00
Teknium b455abe0b3 fix(desktop): poll-guard reset is fire-and-forget off the redial path
composer-status imports $gateway from this module (cycle forces the
dynamic import), and awaiting the module load inside openSecondary sat on
the timed redial path — under CI load that pushed cold-start redials past
waitFor budgets in the lifecycle suite. The reset needs no ordering
guarantee relative to the dial; detach it.
2026-08-26 04:49:22 -07:00
Teknium 62534e2b5a fix(desktop): isolate the poll-guard reset import + sort-imports lint
The combined dynamic import meant a failed composer-status import (mocked
test graphs) silently skipped resetTileRuntimeBindings too — the exact
lifecycle regression CI caught. Separate best-effort trys per module.
2026-08-26 04:49:22 -07:00
Teknium fe615a0099 fix(desktop): republish the connections registry to renderers after every successful save (#95393)
Live-confirmed on the Phase B build: hermesDesktop.connections.save()
succeeds and the registry on disk gains the row, but the switcher menu
(fed by the renderer $connectionsRegistry snapshot) keeps painting the
stale list until reload. remove() already broadcasts
hermes:connections:changed; save() only did so on the dial-material-edit
branch, so a brand-new connection or a label rename never reached the
switcher's onChanged re-pull (or any other window).

Fix at the publish seam only: saveRegistryConnection now broadcasts a new
'saved' reason for every successful save that isn't a dial-material edit.
'saved' is a pure registry-refresh signal — the use-gateway-boot listener
explicitly ignores it (nothing moved, so no dispose/redial/forget), while
the switcher's existing onChanged listener re-pulls the snapshot.

Tests:
- electron/hardening.test.ts pins both broadcast branches in
  saveRegistryConnection (source-assertion pattern; main.ts has no exports).
- connection-switcher.test.tsx mirrors the live repro scenario
  (/tmp/mg-ab/w2_95393.py): menu before save lacks the row, Electron's
  'saved' push arrives, menu after — without reload — shows it.
2026-08-26 04:49:22 -07:00
Bruno Bza 06be6cffbc fix(desktop): release reconnect-orphaned warm transcripts once their authoritative state settles
A gateway connection that dies mid-turn leaves cached session snapshots
whose busy/awaitingResponse flags can never settle: the respawned
backend re-mints runtime ids, so no terminal publish ever reaches the
orphaned snapshot again. #isWarmSettled treated those frozen flags as
live work, so every orphan pinned its full warm transcript until app
restart — roughly 5MB per reconnect cycle, which turned the restart
loop in #95189 into renderer OOM.

SessionStateCache now accepts an optional isAuthoritativelyActive
probe. When wired, in-flight flags only block eviction while the
authoritative $sessionStates record still claims work for the same
runtime id; without the probe the legacy always-block behavior is
preserved byte-for-byte. Eviction remains gated on needsInput, pending
drafts, and active references, so a genuinely running turn (which
re-asserts busy on every publish) is never a casualty.

The useSessionStateCache hook wires the probe to the store it already
imports. Reconnect reconciliation (reconcileBusyStatesOnReconnect)
settles the authoritative record, and the next prune drains the
orphaned cache entry through the normal LRU path, ownership included.
2026-08-26 04:49:22 -07:00
Teknium a7ea156470 fix(desktop): harden the dead-session poll guard per #94950 review
Two review-thread deltas on the salvaged #94950 latch:

- Match the gateway's structured 4001 code, not a message substring, when
  the rejection carries one (JsonRpcGatewayError). A coded error that
  merely mentions 'session not found' in wrapped text (e.g. a 5007 tool
  failure) must not latch the guard and freeze the status stack on a
  healthy session. The substring fallback survives only for codeless
  legacy errors.

- Reset the latch on runtime re-mint, not only on status-stack rebind:
  wire resetBackgroundPollingGuard() at both reconnect seams that already
  drop stale runtime bindings (use-gateway-boot's post-reconnect
  resetTileRuntimeBindings and gateway.ts's reopening path), so ids the
  dead runtime 4001'd resume polling once a respawned backend re-mints
  them.

Tests: code-specific match both directions; full-reset resumes every
latched session.
2026-08-26 04:49:22 -07:00
Justin Johnson c19849cd02 fix(desktop): stop the status-stack poll storming a dead session with 4001s
The composer status stack polls `process.list` every 5s while a background
process row is on screen. `process.list` is session-scoped, so against a
runtime id the gateway no longer holds it returns 4001 "session not found".

`refreshBackgroundProcesses` swallowed *every* failure with a bare `catch {}`
commented "transient socket loss". A gone session is not transient: the poll
re-sent the same dead runtime id every 5 seconds for the lifetime of the
window. On one machine this produced 31,518 gateway rejections in a day
(vs 663 the day before), 18,614 of them against a single runtime id, and it
is what users see reported as "sessions stopped with a session not found
error" after an update.

The trigger is a reconnect, not the poll itself: anything that mints a fresh
runtime (gateway restart, the #94219 reconnect/replay work, an idle-reaped
pooled backend) strands the id the status stack is still holding, and nothing
in this path ever re-checked it.

Distinguish the two failure classes:

- 4001 / "session not found" is TERMINAL for that runtime id — latch the id
  and stop polling it.
- A timeout or transport error is transient — keep retrying, since the
  session may well still be alive. Misclassifying that direction would
  silently freeze the status stack on a healthy session.

The latch is cleared when the status stack (re)binds a session id, so a
session that comes back under a fresh runtime resumes polling normally
rather than staying dark for the life of the app.

Also name the method in the gateway's 4001 warning. That line was added in
c305839442 "for diagnosability", but without the RPC name it cannot say WHICH
client call is looping — the reason this storm could not be attributed from
the logs alone. A ContextVar set in `handle_request` carries it; it is
diagnostic only and never used for authorization.

Tests:
- composer-status: 4001 stops the poll, a timeout does not, one gone session
  never suppresses a healthy sibling, and a rebind resumes polling.
- tui_gateway: the rejection warning names the method.
2026-08-26 04:49:22 -07:00
Teknium 574bd7175c fix(desktop): unify boot-class getConnection() budgets on one shared 45s constant
Follow-up to the #95039 salvage: the cherry-picked bound used the 20s
RECONNECT_ATTEMPT_TIMEOUT_MS on boot()/softSwitch() getConnection(), but a
reviewer note (and the Phase A registry-restore work) established that
boot-class awaits must ride out a full backend cold spawn — main's spawn
budget is 45s (DEFAULT_BACKEND_READY_TIMEOUT_MS). A 20s renderer bound would
latch boot errors on healthy-but-slow cold boots.

Introduce BACKEND_BOOT_WAIT_TIMEOUT_MS (45s) in lib/with-timeout.ts as the
single shared boot-class budget, point boot()/softSwitch() getConnection()
and connections.ts BOOT_DESCRIPTOR_WAIT_TIMEOUT_MS at it, and keep the 20s
reconnect budget only for reconnect-class awaits against an already-spawned
backend. No magic-number drift: 45_000 now appears once in renderer code.
2026-08-26 04:49:22 -07:00
nftpoetrist 31f3de1f06 fix(desktop): bound getConnection() on the boot and soft-switch paths (#93454)
resolveGatewayWsUrl() in attemptReconnect() because a wedged IPC
round-trip into the main process (e.g. a stuck revalidation after a
liveness-probe trip) can hang these awaits forever. A later fix
(e8d5660bae) extended the bound to resolveGatewayWsUrl() in boot() and
softSwitch() too, but left the getConnection() call immediately above
it in both functions unbounded.

If that call wedges: during initial boot the 'Starting Hermes...'
screen never resolves (bootCompleted never flips, nothing hits catch),
and during a soft gateway/profile switch  latches
true forever since the try block's finally never runs.

Wrap both with the same withTimeout()/RECONNECT_ATTEMPT_TIMEOUT_MS
pattern already used for the sibling calls.
2026-08-26 04:49:22 -07:00
Teknium d22e2b9f6e fix(desktop): a canonical-title race adopts the winner instead of forking the forever chat (#92473, part 2)
Between the registry miss and the eager session.title write, another
writer can take the canonical title (peer dm minting server-side, a
second machine, cross-connection sync). UNIQUE(title) rejects our write
with 'already in use' — which the compat path previously read as 'old
gateway' and prompted into OUR stray lazy session, forking the forever
chat. A uniqueness rejection now re-consults the registry and adopts the
winner; the zero-message stray is abandoned to the gateway pruner.
Genuine old-gateway failures (unknown method) keep the compat kickoff.
2026-08-26 04:03:35 -07:00
hermes-seaeye[bot] c427367938 fmt(js): npm run fix on merge (#95457)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 10:27:25 +00:00
Teknium 8e9459c97f style(desktop): sort registrySourceOwnsPrimaryBackend imports (lint) 2026-08-26 03:22:37 -07:00
Teknium 2947272233 refactor(desktop): one canonical write shape for connection_id row stamping
Reconcile #94901's API-layer row stamping with #94656's durable-owner
persistence: extract lib/session-owner-stamp.ts as THE canonical
stamp-untagged-rows write path (never clobbers an explicit owner,
never stamps `local`) and re-express api/sessions'
stampActiveConnectionOwner through it. #94656's writers (optimistic
row from the captured owner route, mergeSessionPage carry, cache
patch) are exact-owner writers and stay as-is; the helper's contract
documents why it must not overwrite them.

Credit: row-stamping concept from PR #94901 (joe-rodgers) and
PR #95007 (weismanfamily); persistence shape from PR #94656
(Zeus-Deus).

Co-authored-by: joe-rodgers <25499388+joe-rodgers@users.noreply.github.com>
2026-08-26 03:22:37 -07:00
joe-rodgers fb393ee08b fix(desktop): stamp remote list rows with their owning connection; retry one transient projects.tree loss
Partial cherry-pick of PR #94901 (joe-rodgers). Surviving scope:
- api/sessions: stampActiveConnectionOwner — rows returned by the
  active non-local gateway are stamped with its registry connection_id
  (explicit owners from multi-source responses preserved), so a later
  resume cannot fall back to a same-named local profile.
- store/projects: one-shot projects.tree retry when a remote source
  switch leaves the first read RPC on a newly-opened socket without a
  response (request timed out / gateway connection closed), only while
  the same gateway/profile is still foreground. Component fix for the
  live-confirmed #92352 sidebar-never-paints gap.

Dropped scope (superseded on main / by the #94656 anchor landed just
below): knownSessionOwner+SessionOwnerScope rewiring in session.ts,
session-states.ts, wiring.tsx (main and #94656 carry richer variants),
and the $connection-derived optimistic-row stamp in
use-session-actions/utils.ts (#94656 stamps the optimistic row from
the captured exact owner route instead of ambient state).

Original-PR: #94901
Dropped-scope: routing half of 2cb5bdbf1 (session.ts, session-states.ts, wiring.tsx, use-session-actions/utils.ts hunks)
2026-08-26 03:22:37 -07:00
Teknium b2a58dbb39 fix(desktop): bound the boot descriptor wait so a dead primary cannot strand the registry restore
Follow-up to the #95007 partial cherry: waitForInitialConnection() was
an unbounded listen on $connection — a primary that never publishes its
descriptor (spawn failure, dead SSH target) would strand
initializeConnectionsRegistry() forever and the last-used source would
never be restored. Bound it with the codebase's withTimeout helper
(same pattern as the sibling SWITCH_* call sites in this file): after
45s (the primary spawn budget) the restore proceeds exactly as it did
before the wait existed, and the listener is torn down either way.

Regression test: boot restore proceeds after the deadline with the
descriptor never arriving.

Original-PR: #95007
2026-08-26 03:22:37 -07:00
Michael Weisman fd031488aa fix(desktop): prove registry-primary backend ownership electron-side; wait for the primary descriptor before boot restore
Partial cherry-pick of PR #95007 (weismanfamily). Surviving scope:
- electron connection-registry: registrySourceOwnsPrimaryBackend() —
  descriptor-level proof that a registry-scoped request names the
  already-running primary backend, wired into ensureRegistryBackend as
  the generic (non-SSH-fingerprint) primary-owns short-circuit so a
  cloud/url registry primary cannot spawn a second isolated server.
- store/connections: waitForInitialConnection() before the boot-time
  source restore, so the sidebar registry cannot dial the preferred
  source a second time while the identical primary backend is still
  publishing its connection identity.

Dropped scope (superseded on main): the renderer routing half —
primaryConnectionId plumbing, primaryOwnsAgent short-circuits in
requestGatewayForAgent/openGatewayForAgent/ensureGatewayForAgent
(main has isPrimaryRegistryRoute via 1ec32e738), the 3-arg
setPrimaryGateway boot wiring (main has setPrimaryGatewayConnection),
the use-session-list-actions stampConnectionOwner (row stamping lands
via #94656/#94901), and the wiring.tsx bare-profile promotion commit
65d106e41 (main's knownSessionOwner covers it).

Original-PR: #95007
Dropped-scope: renderer routing half of d1c0fb093; all of 65d106e41
2026-08-26 03:22:37 -07:00
Deus d6e323bd63 fix(desktop): fail closed through registry boot and drain owner holds 2026-08-26 03:22:37 -07:00