fix(desktop): widen pool keepalive-fresh window to absorb WSL2 IPC stalls (#95189)

The renderer pings each pool backend every 60s (`hermes:backend:touch` →
`touchPoolBackend` → updates `lastActiveAt`). The LRU eviction cap used a
keepalive-fresh window of 90s — only 1.5× the ping cadence — to decide
whether a backend was "plausibly still alive". One missed or delayed ping
pushed a live backend past the threshold and the cap-driven eviction killed
the active profile's backend mid-session, restarting the gateway and
re-minting runtime ids. On WSL2, where the renderer→Electron IPC roundtrips
through 9p, brief 9p hiccups commonly stretch a ping to seconds of observed
silence, producing the ~80–90s exit / ~2 min cycle reported in #95189
(122 gateway starts on 2026-08-26 alone, driving renderer OOM via reconnect
churn at ~5GB/day).

Widen POOL_KEEPALIVE_FRESH_MS to 4 minutes (3× ping cadence + IPC stall
headroom, still bounded well below POOL_IDLE_MS=10min). Backends with one or
even two missed pings are now spared; truly idle backends (multiple lapses,
minutes idle) remain eligible for eviction by the cap and the idle reaper.
The constant is also overridable via HERMES_DESKTOP_POOL_KEEPALIVE_FRESH_MS
to make this tunable without a rebuild.
This commit is contained in:
Finn763
2026-08-26 12:31:20 +08:00
committed by Teknium
parent 8246c4f92a
commit 00988f3b89
2 changed files with 90 additions and 2 deletions
+21 -1
View File
@@ -1405,7 +1405,27 @@ const POOL_IDLE_MS = Math.max(60_000, Number(process.env.HERMES_DESKTOP_POOL_IDL
// pings every 60s for every open profile). LRU eviction must spare these — a
// concurrent multi-profile session keeps several backends "fresh" at once, and
// killing one to honor the soft cap would abort a running agent.
const POOL_KEEPALIVE_FRESH_MS = 90_000
//
// The window is intentionally MUCH wider than the 60s ping cadence:
// * 1 missed ping = +60s of apparent silence
// * WSL2 IPC stall = the renderer's `hermes:backend:touch` roundtrips
// through 9p; a single brief 9p hiccup can stretch a
// ping to ~30s of observed silence (#95189: gateways
// exited every ~2 min on WSL2 because the previous
// 90s window left no headroom — one delayed ping
// pushed a live backend past the threshold and the
// cap-driven eviction killed the active profile's
// backend mid-session, re-minting runtime ids and
// re-allocating pooled gateway secondaries ~700×/day).
// * 3× ping + 60s headroom = ~4 min, comfortable margin for two missed
// pings + WSL2 IPC stall. The hard ceiling for the cap-eligible set is
// POOL_IDLE_MS above (default 10 min) — this constant only governs the
// "is this backend plausibly still alive" question for LRU eviction,
// not when the idle reaper definitively tears a backend down.
const POOL_KEEPALIVE_FRESH_MS = Math.max(
120_000,
Number(process.env.HERMES_DESKTOP_POOL_KEEPALIVE_FRESH_MS) || 4 * 60_000
)
let poolIdleReaper = null
let backendOrphanReapPromise = null
// Auto-reload budget for renderer crashes, shared by EVERY window (primary,
+69 -1
View File
@@ -14,7 +14,8 @@ import { test } from 'vitest'
import { selectPoolEvictions } from './pool-eviction'
const NOW = 1_000_000
const FRESH_MS = 90_000
// Mirrors main.ts POOL_KEEPALIVE_FRESH_MS (4 minutes — see #95189).
const FRESH_MS = 4 * 60_000
/** A spawned local backend entry (has a child process). */
const spawned = (idleMs: number) => ({ process: { pid: 123 }, lastActiveAt: NOW - idleMs })
@@ -83,3 +84,70 @@ test('descriptor-only pools never evict', () => {
assert.deepEqual(selectPoolEvictions(entries, 2, NOW, FRESH_MS), [])
})
// ── #95189 — Keepalive-fresh window must tolerate transient missed pings ──
// Symptom: gateway restarts every ~2 minutes on WSL2. Root cause: the renderer
// pings every 60s; the LRU cap declared a backend "stale" if `lastActiveAt`
// was > 90s ago — only 1.5× the ping interval. WSL2 IPC roundtrips (renderer
// → 9p → Electron main → ipcMain.handle) commonly stall several seconds; one
// delayed or missed ping pushed a live backend past the threshold and the
// cap-driven eviction killed the active profile's backend mid-session,
// forcing a restart loop that re-minted runtime ids and re-allocated pooled
// gateway secondaries ~700×/day (#95189, related #87906/#84716/#88054).
//
// These tests pin the new tolerance: the keepalive-fresh window is wide
// enough to absorb ≥1 missed ping (and the IPC stall headroom around it)
// without evicting an active backend. Truly stale backends (multiple lapses,
// minutes idle) are still evicted as before.
test('#95189: one missed keepalive ping must NOT make the most-recently-touched backend evictable', () => {
// Renderer's keepalive cadence is 60s. With the old freshMs=90s window,
// a backend last touched 95s ago — i.e. exactly ONE missed/delayed ping —
// was eligible for LRU eviction even though it had been actively
// touched the moment before and would be touched again imminently.
//
// Build a pool where the only entry over the cap is the active one (95s
// idle). With the old 90s window, it was evicted; with the widened window
// the cap must instead be honored by leaving the pool over-cap for one
// extra cycle — killing an active backend is far worse than briefly
// exceeding the soft cap.
const entries: [string, ReturnType<typeof spawned>][] = [
['active', spawned(95_000)], // 1 missed ping on a 60s cadence
['fresher', spawned(2_000)] // touched recently — must NOT be evicted
]
// keep=1 → pool over cap by one. Active backend must be spared; the cap
// may be exceeded rather than kill a live backend (the long-standing
// "spare fresh backends" rule from #94381 / earlier pool-eviction tests).
assert.deepEqual(selectPoolEvictions(entries, 1, NOW, FRESH_MS), [])
})
test('#95189: two missed keepalive pings (2-min IPC stall) must NOT evict an active backend', () => {
// The reported symptom: gateways exited ~80–90s after start with a clean
// disconnect (no stderr), recurring every ~2 min. Reproduce the boundary:
// 125s of silence = just over two missed pings at 60s. A backend in this
// state is still actively serving — the renderer is mid-reconnect, not
// gone — so eviction here was the trigger for the restart loop.
//
// The other entries are all FRESH (well within the keepalive window), so
// no eviction is correct even before any cap considerations.
const entries: [string, ReturnType<typeof spawned>][] = [
['active', spawned(125_000)],
['fresher', spawned(2_000)],
['fresher-2', spawned(5_000)]
]
assert.deepEqual(selectPoolEvictions(entries, 1, NOW, FRESH_MS), [])
})
test('#95189: a backend genuinely idle for minutes IS evicted (#95189 long-window scenario)', () => {
// Sanity: the widening does NOT make the pool unbounded. 10 minutes of
// silence (the documented POOL_IDLE_MS) is still fair game.
const entries: [string, ReturnType<typeof spawned>][] = [
['idle', spawned(10 * 60_000)],
['fresh', spawned(5_000)]
]
// keep=1, idle is over the cap AND past the fresh window → evicted.
assert.deepEqual(selectPoolEvictions(entries, 1, NOW, FRESH_MS), ['idle'])
})