Three fixes from one remote-gateway (VPS) debug bundle, all live-reproduced
and re-verified on a headed Electron seat via CDP:
- Bots roster no longer shrinks during a gateway outage: source enumeration
is bounded (10s/source instead of wedging the roster IPC >30s behind a
dead dial) and a bounced remote source keeps painting its last-known
profile list (was SSH-only), so 4 bots never show as 2 mid-outage.
- Pool backend spawns that die before the child exists (forced-local spawn
of a profile that only exists on the remote) now log the failure to
desktop.log, and the profile-exists guard runs BEFORE the Starting line —
no more orphaned no-READY/no-exit spawn bursts in bundles.
- An SSH host-key change (VPS reinstall) is classified terminal like a
reauth rejection: it latches, the boot-failure overlay shows the
ssh-keygen -R guidance, and the renderer stops the infinite boot-retry
loop (one bundle had 157 consecutive failures over 2.5h). Reset/repair/
apply-config clear the latch; live-verified Retry-after-fix boots clean.
A dropped registered remote connection (SSH or HTTP) never recovered on
its own: the next boot attempt failed with a transient transport error
("Could not verify the existing SSH backend", ERR_CONNECTION_RESET,
mint timeout), the failure was correctly NOT latched, but nothing ever
re-attempted the boot — the renderer's reconnect machinery only arms
after a completed boot. The app parked on "Desktop boot failed" until
the user manually deleted and re-entered the same connection details,
which merely forced the fresh bootstrap an automatic retry would have
performed (issue 82679, feature ask 80430).
Root causes and fixes:
- electron/backend-start-failure.ts: new isRetryableRemoteBootFailure()
predicate — a remote, non-reauth boot failure is transient and may be
retried; local failures and confirmed 401/403 rejections are not
(a missing capability differs from a transient failure).
- electron/main.ts: the boot-failure progress broadcast now carries
`retryable` (rides with `error` through updateBootProgress), and a
failed reuse probe against a cached SSH master tears the stale
master/tunnel down so the next attempt bootstraps fresh — exactly
what manual re-entry did.
- use-gateway-boot.ts: bounded self-heal loop for a failed boot whose
progress is marked retryable — up to 5 re-attempts with the same
full-jitter backoff as the socket reconnect loop (2s base, 15s cap).
Exhausted retries end in the real boot-failure recovery overlay,
never an infinite spinner. Reset on success and on soft switch;
timer cleared on unmount.
- store/boot.ts: resumeDesktopBootForRetry() re-arms the overlay with a
retry status while an automatic retry is in flight.
Secondaries already had full-jitter backoff (store/gateway.ts); this
closes the same class for the PRIMARY/registered-connection path.
Tests: predicate matrix (retryable vs reauth-latch mutually exclusive),
plus renderer hook tests proving a transient SSH failure self-heals on
the next attempt, retries are bounded (6 total dials then the recovery
overlay, no further attempts), and non-retryable failures never enter
the loop. Sabotage-verified (disabling either half fails 4 tests).
Fixes#82679Fixes#80430
shouldLatchBackendStartFailure deliberately never latches a remote failure:
remote faults are usually transient and must stay retryable without an app
restart. A confirmed reauth rejection is the exception — it cannot self-heal,
because nothing changes until the user signs in again.
Worse, not latching actively prevents the recovery it was protecting. A
non-latching remote boot failure re-runs startHermes on every
getConnection/api call, re-emits running: true, and the boot-failure overlay
(visible = Boolean(boot.error) && !boot.running) hides itself — so the
"Sign in" button flickers out from under the user before it can be clicked.
Add shouldLatchRemoteReauthFailure as a separate predicate rather than
changing the existing one, so transient remote failures keep self-healing.
The two latches are complementary and never fire for the same failure.