A start-marker probe failure (Get-Process timing out on a PowerShell 5.1
cold start, #87169) in claimBackendChild used to stop the freshly spawned
backend and rethrow — killing a healthy backend, triggering the renderer's
repair respawn, and looping. And because stderr piping only attached after
the claim, every before-ready failure surfaced as a bare exit code.
- extract probe + claim policy into electron/backend-claim.ts:
processStartMarker/execText (moved verbatim from main.ts), probeStartMarker,
and a pure claimDecision(childAlive, probe) a Windows CI lane can drive
with real PowerShell
- probe failure + LIVE child now degrades to PID-only identity
(pid-only:<pid> marker, WARNING logged), matching the existing
createParentStartMarkerResolver degrade pattern; processIdentityMatches
verifies degraded identities by PID liveness (command check still layers
on top in backendIdentityMatches)
- probe failure + DEAD child keeps the fail-closed throw, now carrying the
child's buffered stderr/stdout tail
- ring-buffered ~8KB output tail attached at spawn time in BOTH spawn paths
(pool + primary); tail appended to claim errors, before-ready exit
messages, and backend-ready's exited-before-port-announcement errors so
the real exit reason reaches desktop.log and the boot UI
- tests: claimDecision matrix (degrade test fails against the old
stop+throw behavior), real processStartMarker probe, ring-buffer caps,
and output-tail suffixes on backend-ready exit errors
A persisted tall/narrow size has no way out on Linux. Put a reset next
to Exit HUD so the default size (and position, where the compositor
allows it) is one click away.
Co-authored-by: Shawn Wang <32839114+enwaiax@users.noreply.github.com>
Wayland clients cannot place themselves, so the JS setBounds drag is a
no-op there. Make the composer bar a -webkit-app-region drag handle on
Linux (input carved out with no-drag) and let the compositor move it.
Co-authored-by: Tony Simons <214744153+asimons81@users.noreply.github.com>
The shared-client half of the gateway.ping heartbeat contract (#89958);
tracks lastInboundAt, sends pings, invalidates a silently-dead socket.
Part of #83166.
A group member turn is a session-scoped RPC sequence (resume → attach →
prompt.submit → poll) issued with the runtime id its first RPC minted, but
requestForBot routes every RPC through its own request-scoped socket lease
(retained:false secondaries in store/gateway). Between two RPCs the refcount
hits 0, the leased socket closes, the gateway detaches the runtime session on
WS disconnect, the orphan reaper frees it after grace, and the next RPC —
prompt.submit, unwrapped — dies 4001 'not in memory'. The member turn aborts
and the sub-profile bot goes silent in the room.
- store/gateway: retainGatewayForAgent(connectionId, profile) — refcounted
hold on the pooled socket with an idempotent release, mirroring the
existing request-lease machinery.
- sdk: host.retainProfile(route) exposes the retain to plugins
(feature-detected by consumers; older hosts keep working).
- hermes-bots plugin: runGroupChatMemberTurn acquires the lease before
ensureGroupChatSession's first RPC and releases in finally, so the socket
that minted the runtime id stays open across attach+submit+poll; and
prompt.submit gets a one-shot catch-and-retry that re-resumes via the
STORED session id on 4001-class failures (belt-and-braces for routes the
lease can't cover). 4007 'never existed' keeps flowing to session.create.
Tests: simulated 4001 on first submit recovers via re-resume and delivers;
lease held across attach+submit (mock refcount never hits 0 mid-turn); lease
released after success AND failure; no-retainProfile host feature detection;
store-level retain/release + idempotent double-release + the unretained
disposal race.
#93515 reports auto-speak reading each reply twice when the Edge TTS
streaming attempt falls back to the POST endpoint and the reply's
renderer id gets rewritten to its durable id mid-flight. That was true
before 63565fa26b, but resolveSpokenReply()'s ordinal-anchored dedupe
(landed 2026-08-19, five days before this issue was filed) already
follows the rewrite. No source change — this pins the behavior with a
regression test at the hook/store integration level, one layer above
the existing spoken-reply.ts unit tests.
forkBranch was unconditionally routing every branched session into the
main pane via resumeSession, including sidebar/background branches of
a session the user isn't currently viewing. That reintroduces the
#69750 focus-stealing bug for that path: branching a different session
from the sidebar yanked the active view away from whatever was open.
Only take over the main pane when the branch's parent is the session
already selected; otherwise keep opening it as its own tile.
forkBranch ended by opening the branch as a session-tile and leaving the
primary selection on the parent (#69750). In the default layout there is
no visible tile pane, so branching only added a sidebar row with no
feedback in the main area — and openSessionTile no-ops when the target
is already the selected session, the common case of branching the chat
you're viewing.
Load the branch as the primary session via resumeSession instead, which
reuses the runtime already warm-cached by forkBranch's
ensureSessionState/updateSessionState calls, so it doesn't cost an extra
resume RPC.
Fixes#93444
The unlayered *:focus-visible reset in styles.css intentionally zeroes
--tw-ring-shadow ('No focus rings, anywhere'), so any control that relied
solely on focus-visible:ring-* had no visible keyboard focus state at all.
Mirror each control's hover treatment as a focus-visible background/text
affordance instead, keeping the global reset intact:
- ui/sidebar.tsx: group label, group action, menu button, menu action,
menu sub-button get focus-visible:bg-sidebar-accent + accent foreground
- ui/tabs.tsx: TabsTrigger gets focus-visible:bg-background + text-foreground
- ui/text-tab.tsx: focus-visible:text-foreground (matches its hover)
- chat/composer/micro-actions.tsx: pill gets focus-visible chrome-action-hover
- right-sidebar/index.tsx HEADER_ACTION_CLASS: focus-visible sidebar-accent
- right-sidebar/terminal/rail.tsx RAIL_ACTION: focus-visible chrome-action-hover
- chat/sidebar/cron-jobs-section.tsx (row body + run rows): focus-visible
chrome-action-hover
- chat/sidebar/session-row.tsx <time>: focus-visible:text-foreground
Sweep verified: remaining focus-visible:ring-* usages under apps/desktop/src
already pair with a border/bg/text companion (button/checkbox/switch/input,
starmap share-controls) or are covered by PR #93460's row-hover work
(cron/index.tsx run rows).
Fixes#93462. Reported by @fred0m.
Every opened file registered its preview pane with dock dir 'right', so
each open split a new zone off the right edge — three file opens made
three ever-narrower columns (#93610). The first preview still opens its
own zone docked beside main; every subsequent preview now anchors to an
existing preview-tile pane with dir 'center', so it stacks as a tab in
the same preview zone. Covers files, artifacts, and the Browser tab
alike (all flow through openPreview/$previewTabs); session tiles are
untouched.
Fixes#93610
Cron runs finish unwatched by design, so counting them in
$unreadSessionCount turned the titlebar badge into a permanently-lit
cron run counter (#93552). The badge now counts regular + messaging
sessions only; cron unread state stays visible on the sidebar cron
section rows, and 'Mark all as read' (markAllSessionsRead +
ackAllSessionsRead, which iterates cron rows) still clears them.
Fixes#93552
The bot relay's drain loop RPCs every registered connection through
requestGatewayForAgent's per-request lease. With no other consumer
holding the route, the refcount hit 0 after every tick and the pooled
secondary was disposed — a fresh WebSocket dial + teardown per
connection every 4s, flooding the gateway logs with connect/disconnect
pairs (#93594).
Two changes, both directions from the issue:
- Retained relay-route secondaries: retainGatewayForRelay pins a
route's pooled socket with a counted retention (never clobbering the
foreground 'retained' flag) for the relay's active lifetime, reusing
the existing scheduleReconnect/full-jitter machinery on drops. The
plugin pins each registered connection once via the new feature-
detected host.retainProfileSocket door, reconciles pins with the
current connection set on every drain, and releases everything in
stopBotRelay/dispose. Local routes (null/'local') are exempt so the
idle reaper can still reclaim spawned local backends. The live-work
pruner also respects the pin.
- RELAY_DRAIN_INTERVAL_MS 4s -> 30s: the push path (#93091,
bot_relay.outbox.pending) carries envelope latency, so the poll is
purely a backstop — 30s matches LIVE_SESSION_STATUS_BACKSTOP_INTERVAL_MS.
Tests: relay-push-drain updated to the new backstop semantics; new
gateway-relay-retention.test.ts proves one socket construction across
5 drain ticks (vs 3 constructions for 3 unretained ticks) and that
release/prune/local-exemption behave; new relay-socket-retention
plugin test pins the pin-once / release-on-departure / stop-releases
contracts.
missingRendererAssets only checked the module refs index.html itself names
(<script type=module> + modulepreload), so a torn install whose boot-critical
files were intact but whose lazy chunks were gone passed the generation check
and died minutes later on the first React.lazy() route with 'Failed to fetch
dynamically imported module' (#93479: syntax-diff-*, shiki-*, mermaid-embed-*).
Walk the generation's module graph: for every present JS chunk, parse its
inline __vite__mapDeps filename table (the lazy-import manifest Vite bakes
into each chunk) and check those files too, transitively and cycle-safe.
resolveRendererIndex now skips a lazy-chunk-torn candidate in favor of the
intact copy instead of shipping a delayed crash.
Tests cover the mapDeps parser (definition table vs index-only call sites,
CDN refs), the exact #93479 tear shape, transitive/cyclic walks, and the
torn-vs-intact preference end to end.
The renderer index resolver tried APP_ROOT/dist/index.html — inside app.asar
when packaged — before the app.asar.unpacked copy that asarUnpack (dist/**)
ships and that resolveWebDist() already prefers for the embedded dashboard.
Loading the asar-internal index is how lazily imported chunks (syntax-diff-*,
shiki-*, mermaid-embed-*) end up fetched from a path that cannot serve them,
killing the workspace pane (#93479).
Reorder the candidate ladder to prefer the unpacked web dist when packaged,
following the unpackedPathFor/resolveWebDist precedent. All window loaders
(main, overlay, quick) share resolveRendererIndex, so one reorder covers
every surface. Dev behavior is unchanged: outside an asar both candidates
collapse to APP_ROOT/dist and keep the original order.
React.lazy(() => import('./syntax-diff')) only has its pending state
covered by Suspense. When the dynamic import rejects (e.g. a packaged
app whose renderer window resolves to the app.asar copy of dist/ while
the chunk exists only in app.asar.unpacked, #93479), the rejection
throws past Suspense to the nearest error boundary, which is the whole
workspace ContribBoundary. One missing highlighter chunk then blanks
the entire chat transcript instead of just the diff falling back to
the plain colored DiffBody, the way markdown-text.tsx already isolates
this failure class for markdown.
Wraps LazySyntaxDiff in a local ErrorBoundary that renders DiffBody on
catch, so a failed highlight chunk degrades in place.
Audit of the remaining unguarded botConnectionRoute() callers a pane
render can reach (#93492 follow-up to the botRosterMeta split). Each now
uses the non-throwing resolveBotConnectionRoute() and degrades on an
owner_removed row instead of throwing into the pane's error boundary:
- botWorkspaceOwnerKey / setBotsWorkspaceOwner: sidebar visibility
listener, Bots home open, and roster context menus recompute these on
passive UI edges; an orphaned selection now yields the name-keyed owner
and the blocked workspace target.
- durableGroupChatMembers: rebuilt on every group send over the whole
seated roster; one orphaned member no longer aborts the room update,
and a swept member's degraded mark now survives the rebuild.
- useModelOptions: hook body runs during render; the query is disabled
for an orphaned row and the picker paints its error/disabled state.
- AdvancedProfileConfig: dialog falls back to the bot's own name scope.
Strict dispatch callers (requestForBot, session creation, deleteBot,
duplicateBot, ensureBotMetadata, routines) intentionally keep the
fail-closed throw — remote-routing-races.test.mjs still asserts it.
Adds orphaned-connection-members.test.mjs covering the removed-connection
sweep, the hydrate annotate (with/without a readable registry), the
degraded 'Gateway removed' rendering of swept rows, and every guarded
caller.
Rows poisoned before the removed-connection sweep existed (their
connection was deleted while an older Desktop ran, so no lifecycle push
ever swept them) are what made #93492 survive app restarts. After the
persisted 'group-chats' hydrate, run a pure annotate pass over the rooms:
- a descriptor that lost its connectionId (route unresolvable — the exact
shape that threw on render) is always marked;
- a descriptor whose connectionId is absent from the live connection
registry is marked only when the registry could actually be read —
an unavailable registry must not read as 'everything is orphaned'.
Marked rows keep their identity and degrade to the existing 'Gateway
removed' state; nothing is deleted.
Root cause of #93492: deleting a cloud/remote connection disposed its
gateways (store/gateway.ts) but never touched the persisted 'group-chats'
storage, so every member descriptor referencing the deleted connection
stayed behind as a poisoned row (remoteSource: true, connection gone) that
render-path route lookups tripped over forever.
Subscribe to the connection registry's 'removed' lifecycle push
(window.hermesDesktop.connections.onChanged, feature-detected — older
Electron mains don't emit it) and annotate every persisted group-chat
member owned by the deleted connection. Rows are marked
(sourceMissing/sourceReachable), never silently deleted: the member keeps
its identity and panes render the existing degraded 'Gateway removed'
botSourceStatus state. Writes ride updateGroupChat so the durable record
keeps its full shape, and the listener unbinds on plugin dispose.
botConnectionRoute() stays the strict, throwing dispatch path for real
routing (requestForBot, session creation). botRosterMeta() is passive
display code and previously reached that throw through a bare catch,
which would have swallowed any unrelated failure the same way. It now
calls a new non-throwing resolveBotConnectionRoute() and branches on a
typed resolved | owner_removed | not_scoped status instead.
Adds witnesses for the split: the typed statuses themselves, that
strict dispatch still fails closed on an orphaned row, and that an
unrelated failure while resolving meta for a live route still
propagates instead of being swallowed.
botRosterMeta() calls botConnectionRoute() for every sourceScoped/remoteSource
row to look up its metadata. That's a passive display lookup, but
botConnectionRoute() throws whenever connectionId can't be resolved -- which
is exactly what a stale group-chat roster row looks like once its connection
is deleted (its persisted descriptor keeps remoteSource: true but loses
connectionId). Since botRosterMeta() is called for every member on every
group-chat render, opening a group that still references a deleted
connection threw on render and crashed the pane's error boundary in a loop
that survived app restarts (the poisoned row is in Local Storage).
botConnectionRoute()'s fail-closed throw is correct and stays for its actual
callers -- routing a real request to a bot (requestForBot, session
creation, etc., covered by remote-routing-races.test.mjs). botRosterMeta()
now catches that throw and treats the row as having no resolvable route,
same as a bot with no meta at all, instead of letting it blow up rendering.
Fixes#93492
Follow-up to the reconnect-loop fix: the same unbounded ticket-mint await
exists on the soft gateway-switch path and the initial boot() path. Bound
both with the same withTimeout/RECONNECT_ATTEMPT_TIMEOUT_MS so a wedged
IPC round-trip fails into the existing retry paths instead of hanging the
switch or the 'Starting Hermes…' screen forever.
attemptReconnect() awaited desktop.revalidateConnection?.() unbounded,
immediately before the two IPC calls the previous commit wrapped in
withTimeout(). A wedged revalidation after a liveness-probe trip -
the exact trigger #93454 and this file's own comment describe - hung
that await forever, so the reconnecting guard never cleared and the
prior fix never got reached.
Wrap it in the same 20s withTimeout() (still swallowing the result via
.catch, matching its existing best-effort semantics) and extend the
regression test to hang revalidateConnection() specifically, proving
getConnection() and the socket still proceed once the stall times out.
After a liveness-probe-triggered reconnect on a remote gateway,
attemptReconnect() awaits desktop.getConnection() and resolveGatewayWsUrl()
with no timeout. If either stalls (e.g. main process wedged mid-revalidation
even though the backend itself is reachable), the `reconnecting` guard never
clears, so every later scheduleReconnect()/attemptReconnect() early-returns
forever and the UI stays stuck in "reconnecting" until the app is restarted.
Bound both awaits with a 20s timeout so a stall rejects instead of hanging;
the existing catch/finally already clears the guard and resumes backoff on
rejection. gateway.connect() keeps its own separate connect timeout.
Fixes#93454
After a remote backend update restarts the gateway, the window WebSocket often dies without a close event (SSH/tailscale tunnels) and users force-quit to recover. finishBackendApply now nudges the registered reconnect handler, which rides the new ping liveness probe: healthy sockets are left untouched, dead ones are force-closed and re-dialed. The old blind gateway.close() on every wake signal is removed in favor of the probe.
macOS sleep/wake (or a silent network drop) can leave the renderer's
WebSocket half-open: no close event fires, so connectionState stays
'open' while every RPC hangs until its per-call timeout. prompt.submit's
timeout is 30 minutes, so the user's next message reads as "enter does
nothing until I restart the app".
- Add a minimal ping RPC (tui_gateway/server.py) answered synchronously
on the WS reader thread.
- On wake signals, reconnectNow now probes the open-looking socket with a
5s-bounded ping and force-closes it on failure, letting the existing
reconnect machinery (backoff, tile rebinding, session refresh) take
over. A pre-ping backend answering -32601 is treated as healthy.
- Tests: half-open socket force-reconnects; healthy socket untouched;
method-not-found backend untouched; backend ping envelope contract.
The whole test file legitimately runs ~14s on CI (heavy dynamic import paid
by the first test), brushing the global 15s per-test budget. Slow runners
tip the first test over and cascade-fail all 11 — hit twice in a row on PR
#93612 and on a main run in the same hour. Raise the file's describe-level
timeout to 60s; individual tests still run in milliseconds locally.
Sabotage-verified: timeout:1 fails all 11, 60s passes all 11.
macOS Chinese pinyin IME: pressing Enter to confirm a candidate word in
the group-chat composer submitted the draft as a message mid-composition.
The GroupMentionInput onKeyDown checked only `event.key === 'Enter' &&
!event.shiftKey` with no IME guard, unlike the core composer which guards
isComposing + keyCode 229 (#44135).
Add the same guard to the three Enter handlers in the bots plugin:
- GroupMentionInput (group composer + reply box) — the reported bug
- GroupClarifyCard free-text answer input — same premature-submit
- skill-hub search input — same premature-trigger
Closes#93528
Follow-up hardening on #92977 (issue #92976). The cherry-picked retry
wrapped every verb, so an ECONNRESET arriving after the backend had
already processed a POST (prompt submitted, session created) would
silently double-submit on retry.
- Extract the transport policy into electron/api-transport.ts so it is
unit-testable without Electron: keep-alive agent pools, transient
error classification, and a verb-gated withRetry.
- Retry rule: GET/HEAD/OPTIONS retry on any transient transport error;
POST/PUT/PATCH/DELETE retry only when the request provably never
reached the server (connect-phase failures like ECONNREFUSED /
ENOTFOUND, or an error thrown before the body was flushed —
requestState.bodySent === false). Ambiguous resets after the body
went out surface to the caller; when in doubt, don't retry.
- Separate keep-alive pools for JSON calls vs streaming downloads so
long downloads can't starve latency-sensitive JSON calls.
- Destroy pooled agents on app will-quit.
- Tests: shouldRetryRequest truth table, withRetry behavior, plus LIVE
transport tests against real misbehaving node HTTP servers: a GET
burst where the server resets keep-alive sockets (bare attempt fails,
retried succeeds) and a POST whose socket is RST after server-side
processing (hit counter stays 1 — no double submit).
While a fullscreen app owns the screen, the HUD becomes an in-game chat frame:
the idle bar steps back to a glanceable opacity, and the transcript is held
open for as long as the game is there rather than fading on a timer — you look
back at a chat log during a lull, not while the text happens to be fresh.
Detection is a pure pass over the same front-to-back window enumeration
read_window_below uses (electron/hud-game-overlay.ts); main polls it while the
HUD is open and pushes changes to the renderer, which owns the treatment. Two
details the enumeration forced:
- Hysteresis. Entering needs the game to be what the user is actually looking
at, so a windowed app on top vetoes it. Staying only needs the game to still
exist: clicking the HUD to type de-foregrounds the game and floats every
other open window above it, which otherwise dropped overlay mode at the
moment the user engaged with it.
- The last state is replayed on did-finish-load. The watch pushes only on
change and its first tick fires at window creation, before the renderer has
mounted its listener, so a HUD opened over an already-fullscreen game
consumed its only message and sat at 'no game' forever.
The band itself is reworked for living over someone else's window:
- Light-on-dark unconditionally. The theme's near-black body ink is unreadable
over a dark game, and every attempt to gate the light ink on some
condition — focus, then the game flag — produced a state where it evaluated
false and the words went black on black. The sheet is a dark scrim in every
theme so white is always right; anything that paints its own light surface
(a clarify question, an approval card, a code block, a form control) opts
back into theme ink by re-pointing the ink variable, matched on the fill it
paints rather than the feature it belongs to.
- Your own lines are gold rather than bubbled. With no card the log otherwise
reads as one voice; blue and purple are what most game UIs use for their own
text, so they disappear into the background.
- The scrollback ramps out at the top instead of being cut off, masked on the
scroller (the band is a static box — its rows overflow the thread viewport
nested inside it, so a mask on the band ramps over empty space).
- The sheet is inset under the bar, so its square top corners no longer poke
out past the bar's rounded ones.