A pooled remote backend (Bot Mode, group chat) keeps its descriptor and SSH
forward cached in the backend pool. When the remote Desktop relaunches, the
remote process dies but the local forward stays LISTENing, so
ensureRegistryBackend() keeps returning the dead descriptor and every dispatch
to that machine fails until the app is restarted.
The background sweep cannot cover this: revalidatePooledRemoteBackends() only
runs from the renderer reconnect IPC, which never fires while the primary
connection stays healthy.
Validate the exact cached descriptor at dispatch time with a short /api/status
probe (2.5 s). On failure, retire the pool entry and its SSH forward, then
reconnect on demand. Concurrent dispatches share one retire/reconnect sequence
through a RemoteRevalidationCoordinator keyed on the cached promise, and
identity checks make a late failure from an old descriptor unable to tear
down a replacement another caller already installed.
Verified on a two-Mac setup (MacBook + Mac mini over SSH): after relaunching
the Mac mini's Desktop, a group-chat turn from the MacBook now reaches the
mini's backend and its reply lands, where it previously failed forever.
Electron pre-installs its own uncaughtException listener and only warns on
unhandled rejections, so a main-process fault usually leaves the app running
with the reason on stderr — which nothing captures when the app is launched
from Finder or the Start menu. The fault never reaches desktop.log, so it is
absent from `hermes debug share` and the user can only describe symptoms.
Record both to desktop.log and flush synchronously, since a fault that does
prove fatal leaves no chance for the batched async flush. Five loadURL calls
were also unhandled, each able to leave a blank window with no explanation
anywhere the user can send us; they now name the surface that failed.
Co-authored-by: Rodrigo Fernandez <rod@nxtlevel.dev>
A pooled backend entry pointing at a remote host has no child process, so
the 'exit' handler that clears a dead local backend never fires. The
renderer's 60s keepalive touch also spares it from the idle reaper. Nothing
was left to retire the descriptor, so once the host went away the pool kept
serving it and every profile bound to that host stayed broken until restart.
Pooled remote descriptors now share the primary's liveness policy: probed on
the same revalidate tick, keyed per base URL, and dropped only after the
same consecutive-failure limit, so the next ensureBackend() rebuilds.
Co-authored-by: Rodrigo Fernandez <rod@nxtlevel.dev>