botConnectionRoute() stays the strict, throwing dispatch path for real
routing (requestForBot, session creation). botRosterMeta() is passive
display code and previously reached that throw through a bare catch,
which would have swallowed any unrelated failure the same way. It now
calls a new non-throwing resolveBotConnectionRoute() and branches on a
typed resolved | owner_removed | not_scoped status instead.
Adds witnesses for the split: the typed statuses themselves, that
strict dispatch still fails closed on an orphaned row, and that an
unrelated failure while resolving meta for a live route still
propagates instead of being swallowed.
botRosterMeta() calls botConnectionRoute() for every sourceScoped/remoteSource
row to look up its metadata. That's a passive display lookup, but
botConnectionRoute() throws whenever connectionId can't be resolved -- which
is exactly what a stale group-chat roster row looks like once its connection
is deleted (its persisted descriptor keeps remoteSource: true but loses
connectionId). Since botRosterMeta() is called for every member on every
group-chat render, opening a group that still references a deleted
connection threw on render and crashed the pane's error boundary in a loop
that survived app restarts (the poisoned row is in Local Storage).
botConnectionRoute()'s fail-closed throw is correct and stays for its actual
callers -- routing a real request to a bot (requestForBot, session
creation, etc., covered by remote-routing-races.test.mjs). botRosterMeta()
now catches that throw and treats the row as having no resolvable route,
same as a bot with no meta at all, instead of letting it blow up rendering.
Fixes#93492
Follow-up to the reconnect-loop fix: the same unbounded ticket-mint await
exists on the soft gateway-switch path and the initial boot() path. Bound
both with the same withTimeout/RECONNECT_ATTEMPT_TIMEOUT_MS so a wedged
IPC round-trip fails into the existing retry paths instead of hanging the
switch or the 'Starting Hermes…' screen forever.
attemptReconnect() awaited desktop.revalidateConnection?.() unbounded,
immediately before the two IPC calls the previous commit wrapped in
withTimeout(). A wedged revalidation after a liveness-probe trip -
the exact trigger #93454 and this file's own comment describe - hung
that await forever, so the reconnecting guard never cleared and the
prior fix never got reached.
Wrap it in the same 20s withTimeout() (still swallowing the result via
.catch, matching its existing best-effort semantics) and extend the
regression test to hang revalidateConnection() specifically, proving
getConnection() and the socket still proceed once the stall times out.
After a liveness-probe-triggered reconnect on a remote gateway,
attemptReconnect() awaits desktop.getConnection() and resolveGatewayWsUrl()
with no timeout. If either stalls (e.g. main process wedged mid-revalidation
even though the backend itself is reachable), the `reconnecting` guard never
clears, so every later scheduleReconnect()/attemptReconnect() early-returns
forever and the UI stays stuck in "reconnecting" until the app is restarted.
Bound both awaits with a 20s timeout so a stall rejects instead of hanging;
the existing catch/finally already clears the guard and resumes backoff on
rejection. gateway.connect() keeps its own separate connect timeout.
Fixes#93454
After a remote backend update restarts the gateway, the window WebSocket often dies without a close event (SSH/tailscale tunnels) and users force-quit to recover. finishBackendApply now nudges the registered reconnect handler, which rides the new ping liveness probe: healthy sockets are left untouched, dead ones are force-closed and re-dialed. The old blind gateway.close() on every wake signal is removed in favor of the probe.
macOS sleep/wake (or a silent network drop) can leave the renderer's
WebSocket half-open: no close event fires, so connectionState stays
'open' while every RPC hangs until its per-call timeout. prompt.submit's
timeout is 30 minutes, so the user's next message reads as "enter does
nothing until I restart the app".
- Add a minimal ping RPC (tui_gateway/server.py) answered synchronously
on the WS reader thread.
- On wake signals, reconnectNow now probes the open-looking socket with a
5s-bounded ping and force-closes it on failure, letting the existing
reconnect machinery (backoff, tile rebinding, session refresh) take
over. A pre-ping backend answering -32601 is treated as healthy.
- Tests: half-open socket force-reconnects; healthy socket untouched;
method-not-found backend untouched; backend ping envelope contract.
The whole test file legitimately runs ~14s on CI (heavy dynamic import paid
by the first test), brushing the global 15s per-test budget. Slow runners
tip the first test over and cascade-fail all 11 — hit twice in a row on PR
#93612 and on a main run in the same hour. Raise the file's describe-level
timeout to 60s; individual tests still run in milliseconds locally.
Sabotage-verified: timeout:1 fails all 11, 60s passes all 11.
macOS Chinese pinyin IME: pressing Enter to confirm a candidate word in
the group-chat composer submitted the draft as a message mid-composition.
The GroupMentionInput onKeyDown checked only `event.key === 'Enter' &&
!event.shiftKey` with no IME guard, unlike the core composer which guards
isComposing + keyCode 229 (#44135).
Add the same guard to the three Enter handlers in the bots plugin:
- GroupMentionInput (group composer + reply box) — the reported bug
- GroupClarifyCard free-text answer input — same premature-submit
- skill-hub search input — same premature-trigger
Closes#93528
Follow-up hardening on #92977 (issue #92976). The cherry-picked retry
wrapped every verb, so an ECONNRESET arriving after the backend had
already processed a POST (prompt submitted, session created) would
silently double-submit on retry.
- Extract the transport policy into electron/api-transport.ts so it is
unit-testable without Electron: keep-alive agent pools, transient
error classification, and a verb-gated withRetry.
- Retry rule: GET/HEAD/OPTIONS retry on any transient transport error;
POST/PUT/PATCH/DELETE retry only when the request provably never
reached the server (connect-phase failures like ECONNREFUSED /
ENOTFOUND, or an error thrown before the body was flushed —
requestState.bodySent === false). Ambiguous resets after the body
went out surface to the caller; when in doubt, don't retry.
- Separate keep-alive pools for JSON calls vs streaming downloads so
long downloads can't starve latency-sensitive JSON calls.
- Destroy pooled agents on app will-quit.
- Tests: shouldRetryRequest truth table, withRetry behavior, plus LIVE
transport tests against real misbehaving node HTTP servers: a GET
burst where the server resets keep-alive sockets (bare attempt fails,
retried succeeds) and a POST whose socket is RST after server-side
processing (hit counter stays 1 — no double submit).
While a fullscreen app owns the screen, the HUD becomes an in-game chat frame:
the idle bar steps back to a glanceable opacity, and the transcript is held
open for as long as the game is there rather than fading on a timer — you look
back at a chat log during a lull, not while the text happens to be fresh.
Detection is a pure pass over the same front-to-back window enumeration
read_window_below uses (electron/hud-game-overlay.ts); main polls it while the
HUD is open and pushes changes to the renderer, which owns the treatment. Two
details the enumeration forced:
- Hysteresis. Entering needs the game to be what the user is actually looking
at, so a windowed app on top vetoes it. Staying only needs the game to still
exist: clicking the HUD to type de-foregrounds the game and floats every
other open window above it, which otherwise dropped overlay mode at the
moment the user engaged with it.
- The last state is replayed on did-finish-load. The watch pushes only on
change and its first tick fires at window creation, before the renderer has
mounted its listener, so a HUD opened over an already-fullscreen game
consumed its only message and sat at 'no game' forever.
The band itself is reworked for living over someone else's window:
- Light-on-dark unconditionally. The theme's near-black body ink is unreadable
over a dark game, and every attempt to gate the light ink on some
condition — focus, then the game flag — produced a state where it evaluated
false and the words went black on black. The sheet is a dark scrim in every
theme so white is always right; anything that paints its own light surface
(a clarify question, an approval card, a code block, a form control) opts
back into theme ink by re-pointing the ink variable, matched on the fill it
paints rather than the feature it belongs to.
- Your own lines are gold rather than bubbled. With no card the log otherwise
reads as one voice; blue and purple are what most game UIs use for their own
text, so they disappear into the background.
- The scrollback ramps out at the top instead of being cut off, masked on the
scroller (the band is a static box — its rows overflow the thread viewport
nested inside it, so a mask on the band ramps over empty space).
- The sheet is inset under the bar, so its square top corners no longer poke
out past the bar's rounded ones.
setBackgroundMaterial on a transparent window permanently kills per-pixel
alpha on Win11 — every transparent pixel composites as opaque white, so the
HUD showed a white slab instead of the desktop behind it. Verified against a
minimal repro on Electron 40.10.2: the break happens with ANY material value
including 'none', which is exactly what the idle HUD asks for, and neither
'auto' nor a follow-up setBackgroundColor('#00000000') restores it.
The DWM backdrop and window transparency are mutually exclusive, so the
Windows HUD keeps the CSS tint its sheet already paints and skips the native
frost. macOS is untouched: setVibrancy composites correctly.
window-below asked node_modules for get-windows, whose lib/windows.js locates
its native binding through preGyp.find() — by HOST platform. When the tree was
installed on one OS and Electron is running on another (a WSL-hosted dev run
driving a win32 Electron), pre-gyp picks the host's slot, ignores the correct
binding sitting beside it, and upstream's fail-soft path returns no-op stubs.
Enumeration then reports 'unavailable' on a machine that answers perfectly
well, which silently disables read_window_below.
scripts/stage-native-deps.mjs already writes a staged lib/windows.js that
requires its binding directly, so prefer it and keep the bare import as the
fallback.
message_agent callers previously got provider prose (a raw 401
paragraph, a missing-provider essay) and could not branch on the
failure class. Now the #93091 item-1 reason enum rides the whole relay
roundtrip:
- Desktop relay drain forwards bot_relay.deliver's error.data.reason
into bot_relay.reply (and prefers it for the attention badge over
free-text re-parsing);
- write_reply already persisted reason / classified fallbacks;
- the sender-side waiter prints "[reason: <code>]" ahead of the free
text, so the completion notification the sending agent receives is
machine-branchable.
Additive everywhere: healthy replies unchanged, reasonless errors
classify to a code, old consumers keep working.
ensureGroupChatSession's resume loop caught ANY session.resume error
(stored sid, then title lookup) identically and fell through to
session.create — the same bug findExistingCanonicalChat was fixed for
hours earlier (87b645f52c) in the same file: a transient failure (the
backend still warming up after a restart, a network blip on a
cross-connection lookup, an oversized-resume refusal) read as "no
session, mint a new one". That forks the member's real session AND
silently overwrites room.sessions[key], making the original
unreachable from the room. ensureGroupChatSession is actually more
exposed than the 1:1 case: it runs every group turn
(runGroupChatMemberTurn), with two independent swallow points.
Distinguish "genuinely doesn't exist" from "transient failure" the
same way the gateway itself does: session.resume's own handler
(tui_gateway/methods_session.py) returns JSON-RPC code 4007 only when
the target truly isn't found; every other failure (including 4130,
"session too large to resume" — a session that DOES exist) now
surfaces instead of being silently swallowed. The existing outer
try/catch at the call site already treats a thrown error as "this
member passes the round" (recordGroupActivity kind: 'failed'), so
nothing new needs to catch it — a transient hiccup now costs one
skipped round instead of a permanent fork.
The toggle was gated on having 2+ registered sources, which hid it in exactly
the local-only state the drift produces — the state where a user most needs to
change what launch restores.
Reconciliation repairs the drift at its source, but it can still fail to
persist (read-only or full userData), which leaves a window live on a source
the registry cannot name. $activeConnectionId is null there, the preferred-id
guard misses, and the restore re-homes a working connection.
Return early when a connection is live but unnameable. The registry has no
claim on a source it does not know about.
migrateV1ToRegistry runs exactly once, only when connections.json is absent.
A user who was local at that moment and pointed Settings -> Gateway at a
remote afterwards gets a live remote the registry cannot name: the descriptor
resolves to no connectionId, primary still says 'local', and the boot-time
launch pick force-switches the window onto a fresh local backend seconds after
the sessions list paints. That backend has no provider, so onboarding pops.
Reconcile on read: when the v1 global route names a remote with no matching
registry entry, register it and adopt it as primary/last-used, then persist so
the repair happens once. Narrow on purpose — an already-registered route is
left alone even when primary names something else, because that is the user's
pick in the Connections panel, not drift.
Replaces the hand-edit-connections.json workaround users have been trading.
The showAllProfiles browse-mode flag is persisted to localStorage, but
every restart it was force-collapsed anyway: initializeConnectionsRegistry
restores the last-used source via selectConnection, and selectConnection's
post-activation path unconditionally ran $showAllProfiles.set(false).
That collapse is correct for a user click on the connection picker (a
concrete-source action), but the silent boot restore is not a user action.
Gate both reset sites on pendingTarget === null && activeConnectionId ===
null (the fresh-boot state) so the persisted preference survives restart,
while any user-initiated switch still collapses browse mode.
Regression tests cover both directions: boot restore preserves true, a
user switch collapses it.
Fixes#93197
The salvaged hardening tests matched main.ts source text to assert that
fetchConnectionStatus reaches for a bearer and that Apply preflights before
persisting. A rename breaks them while a real auth regression that keeps the
substrings passes.
Make the preflight a first-class option on applyConnectionConfigAtomically so
its ordering is observable, and assert it through the seam: preflight runs
before either write, and a rejected preflight leaves both stores and the
activation untouched.
Follow-up to #93339: the auxiliary.review slot existed in config but was
missing from every model-picker surface, so users could only set the
review model by hand-editing config.yaml.
- hermes_cli/web_server.py: review in _AUX_TASK_SLOTS (REST allowlist,
stale-aux warning sweep)
- hermes_cli/main.py: review in _AUX_TASKS (hermes model aux picker)
- apps/desktop model-settings.tsx + all 5 i18n locales (en/ja/zh/
zh-hant/ar): review slot with label/hint
- web/src/pages/ModelsPage.tsx: review row in dashboard Models page
- tests: registry-sync test pinning review across DEFAULT_CONFIG,
_AUX_TASKS, and _AUX_TASK_SLOTS (curator pattern)
- docs: aux-task table in fallback-providers.md (en) + zh-Hans mirrors
of fallback-providers and the delegation /review section missed in
#93339
CI sibling-test blast radius from the cluster branch:
- singleFlightSessionResume crashed on run() doubles that return
non-promises (Cannot read 'finally'); wrap via Promise.resolve().then(run).
- Three use-session-tile-delegate tests pinned the pre-#92961 ambient
dispatch for default-profile sessions; the routing-authority change
intentionally routes every known owner through the profile router, so
the tests now assert requestGatewayForProfile('default', ...) instead.
Fixes#85834 (Electron REST intercept fall-through). The
/api/sessions/{id}[/messages] intercept in electron/main.ts required an
explicit ?profile= (or request.profile) to route a read to its remote owner;
callers without a hint fell straight through to the LOCAL backend and 404'd
on its state.db even though the session lives on a configured remote — while
the list endpoints happily showed the row (remoteSessionList tags s.profile).
When no explicit profile resolves, consult the same remote session lists the
list endpoints use to find the owning profile (matching id or lineage root
id), memoized for 30s so a transcript+messages burst costs one sweep. Only
when the id is genuinely unknown remotely does the request fall through to
local, exactly as before. Pure lookup lives in profile-session-routing.ts
with unit tests (owner hit, lineage-root match, null on miss/dead
remotes/no remotes).
Maintainer commit (cluster salvage).
Client half of #91684. The approval bar (approval.tsx) and the native
notification action path (native-notifications.ts) sent approval.respond on
the AMBIENT gateway socket. Ambient follows foreground focus; for an approval
raised by a cross-profile or tile-owned session it points at a backend that
never held the approval, so Run/Reject silently failed after a profile swap
or reconnect.
- New knownOwnerForSession/requestForOwnedSession in store/session-states.ts:
resolve the owner sync (tile owner route -> known session profile via row or
open-time hint; runtime ids translated to stored ids first) and dispatch via
requestForSessionProfile. Ambient only when no owner is known — never a
fall-back to "active".
- approval.tsx and native-notifications.ts respond through it, binding the
ambient dispatcher so the no-owner path keeps the exact 2-arg call shape.
- Tests: owner resolution (tile route first, row-profile fallback,
undefined for unknown/null) and ambient arity preservation in
session-states.test.ts; existing approval + native-notification suites
still pass unchanged on the ambient path.
Maintainer commit (cluster salvage).
After sleep/wake or a reconnect, many surfaces discover the same dead runtime
at once (submit recovery, slash/rewind recovery, tile resumes, the target
resolver, session switch) and each fired its own session.resume — the gateway
minted a runtime per call and the losers fed the orphan reaper (#91276 storm).
- New use-prompt-actions/single-flight-resume.ts: module-level in-flight map
keyed by storedSessionId; all resume call sites (utils.ts recovery, submit.ts
direct rung, resolve-target-session.ts, use-session-actions switch resume,
use-session-tile-delegate resumeTile) share one in-flight promise per stored
id. Failed flights are not cached.
- Drift-abort paths no longer abandon a freshly-minted runtime: utils.ts
SessionRecoveryAborted and submit.ts post-routed-resume / post-resume aborts
register it in a stored->runtime recovery cache; the next action for that
stored session adopts it (via onRecovered) or reuses it instead of minting
another. Cache entries are take-once and skip a known-dead id.
- Unit tests: one RPC for two concurrent callers of the same stored id,
drift-abort registers (not strands) the recovered runtime, independent
stored ids resume independently, cached-runtime adoption.
Maintainer commit (cluster salvage, part of the session-not-found-after-
reconnect consolidation).
Follow-up to PR #91357 (salvaged, author enwaiax): the committed #90428
explicit-target regression fixture started foreground B with a valid active
runtime and a positive B->runtime cache entry, so routedSessionNeedsResume was
false and the formerly broken foreground-recovery branch was never exercised —
the test passed even on the broken head (b9df1f9c2).
Strengthen it per the review: B now starts with activeSessionIdRef null and an
empty ownership cache, resumeStoredSession(B) fully publishes B's runtime and
cache binding, and the assertions still require no high-level resume of B, an
authoritative session.resume(C), exactly one queued prompt.submit to C's
recovered runtime, and no mutation of foreground refs/cache.
Salvaged-from: PR #91357 (author enwaiax); fixture hardening by maintainer.
Step 2 of removing 'active gateway' as a routing input. A session's backend
is a property of the SESSION (its profile), never of whatever the window is
currently showing. The active-profile fallback was the root cause of Bot Mode
'session not found' / hangs: a hidden/unlisted session with an unknown owner
was silently dispatched to the active profile's backend, which never owned it.
- sessionRpcNeedsProfileRoute: drop the active-profile comparison entirely. A
KNOWN owner (route or profile name) ALWAYS routes to its own profile's
socket; only a null/empty owner (fresh draft, global chrome) routes ambient.
A primary-profile owner collapses back to the primary socket inside
gatewayForProfile, so the reauth-aware reconnect path is unchanged.
- session.ts: split knownSessionProfile (row -> hint, undefined when unknown)
out of rememberedSessionProfile. rememberedSessionProfile keeps its active
fallback but is now documented as PRESENTATION-only (navigation keying),
never routing.
- wiring requestGateway: resolve the owner from the tile route -> known
profile -> a cross-profile REST probe (resolveSessionProfile, stamps
ownership) before dispatch; only a request with no session at all falls to
ambient. Never the silent active fallback.
Tests updated to the new contract + knownSessionProfile coverage asserting it
returns undefined (not active) for an unknown session. tsc 0 errors.