Commit Graph

7 Commits

Author SHA1 Message Date
HexLab98 19e56b446b test(desktop): cover unsigned OAuth latch vs needsOauthLogin-only retry 2026-08-23 19:26:42 -05:00
hermes-seaeye[bot] 6f31cfad78 fmt(js): npm run fix on merge (#90690)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-20 09:20:15 +00:00
Teknium 6170021714 fix(desktop): remote-gateway desktop stops lying after disconnects — roster survives outages, spawn failures log, host-key change stops the retry wall
Three fixes from one remote-gateway (VPS) debug bundle, all live-reproduced
and re-verified on a headed Electron seat via CDP:

- Bots roster no longer shrinks during a gateway outage: source enumeration
  is bounded (10s/source instead of wedging the roster IPC >30s behind a
  dead dial) and a bounced remote source keeps painting its last-known
  profile list (was SSH-only), so 4 bots never show as 2 mid-outage.
- Pool backend spawns that die before the child exists (forced-local spawn
  of a profile that only exists on the remote) now log the failure to
  desktop.log, and the profile-exists guard runs BEFORE the Starting line —
  no more orphaned no-READY/no-exit spawn bursts in bundles.
- An SSH host-key change (VPS reinstall) is classified terminal like a
  reauth rejection: it latches, the boot-failure overlay shows the
  ssh-keygen -R guidance, and the renderer stops the infinite boot-retry
  loop (one bundle had 157 consecutive failures over 2.5h). Reset/repair/
  apply-config clear the latch; live-verified Retry-after-fix boots clean.
2026-08-20 02:07:10 -07:00
hermes-seaeye[bot] 1826310f49 fmt(js): npm run fix on merge (#88128)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-17 04:19:24 +00:00
Teknium 2b7f49673f fix(desktop): self-heal dropped SSH/HTTP registered remote connections
A dropped registered remote connection (SSH or HTTP) never recovered on
its own: the next boot attempt failed with a transient transport error
("Could not verify the existing SSH backend", ERR_CONNECTION_RESET,
mint timeout), the failure was correctly NOT latched, but nothing ever
re-attempted the boot — the renderer's reconnect machinery only arms
after a completed boot. The app parked on "Desktop boot failed" until
the user manually deleted and re-entered the same connection details,
which merely forced the fresh bootstrap an automatic retry would have
performed (issue 82679, feature ask 80430).

Root causes and fixes:

- electron/backend-start-failure.ts: new isRetryableRemoteBootFailure()
  predicate — a remote, non-reauth boot failure is transient and may be
  retried; local failures and confirmed 401/403 rejections are not
  (a missing capability differs from a transient failure).
- electron/main.ts: the boot-failure progress broadcast now carries
  `retryable` (rides with `error` through updateBootProgress), and a
  failed reuse probe against a cached SSH master tears the stale
  master/tunnel down so the next attempt bootstraps fresh — exactly
  what manual re-entry did.
- use-gateway-boot.ts: bounded self-heal loop for a failed boot whose
  progress is marked retryable — up to 5 re-attempts with the same
  full-jitter backoff as the socket reconnect loop (2s base, 15s cap).
  Exhausted retries end in the real boot-failure recovery overlay,
  never an infinite spinner. Reset on success and on soft switch;
  timer cleared on unmount.
- store/boot.ts: resumeDesktopBootForRetry() re-arms the overlay with a
  retry status while an automatic retry is in flight.

Secondaries already had full-jitter backoff (store/gateway.ts); this
closes the same class for the PRIMARY/registered-connection path.

Tests: predicate matrix (retryable vs reauth-latch mutually exclusive),
plus renderer hook tests proving a transient SSH failure self-heals on
the next attempt, retries are bounded (6 total dials then the recovery
overlay, no further attempts), and non-retryable failures never enter
the loop. Sabotage-verified (disabling either half fails 4 tests).

Fixes #82679
Fixes #80430
2026-08-16 19:58:37 -07:00
Brooklyn Nicholson 36e2228a4c fix(desktop): latch a confirmed remote reauth failure so the overlay stays clickable
shouldLatchBackendStartFailure deliberately never latches a remote failure:
remote faults are usually transient and must stay retryable without an app
restart. A confirmed reauth rejection is the exception — it cannot self-heal,
because nothing changes until the user signs in again.

Worse, not latching actively prevents the recovery it was protecting. A
non-latching remote boot failure re-runs startHermes on every
getConnection/api call, re-emits running: true, and the boot-failure overlay
(visible = Boolean(boot.error) && !boot.running) hides itself — so the
"Sign in" button flickers out from under the user before it can be clicked.

Add shouldLatchRemoteReauthFailure as a separate predicate rather than
changing the existing one, so transient remote failures keep self-healing.
The two latches are complementary and never fire for the same failure.
2026-07-25 21:52:45 -05:00
HexLab 3f199f5c51 fix(desktop): don't latch remote backend boot failures so remote gateway reconnect recovers (#65756) 2026-07-16 20:45:43 -04:00