DPR-sized canvas backing store separated from CSS footprint, tracking
zoom/display changes live. Rendering-fix subset of PR #75307; the
overlay-placement feature portion is out of scope here.
Covers the devicePixelRatio half of #83216.
All three artifact timestamp sources (message.timestamp,
session.last_active, session.started_at) are epoch SECONDS — the
transcript reader and session-date-groups both multiply by 1000 — but
the collector passed them straight to new Date() (ms), so every
artifact rendered as 1970-01-21. Normalize seconds to ms once at
collection; the Date.now() fallback stays ms.
Local file artifacts (e.g. D:\ComfyUI\output\*.png) fell through to
mediaExternalUrl() which yields a file:// URL the renderer cannot
load. Route through the desktop fs bridge whenever it exists —
readDesktopFileDataUrl already dispatches remote REST vs local
Electron internally (#83380).
Require explicit provenance for tool-result artifacts while preserving assistant links, MEDIA deliveries, generated outputs, file mutations, and browser screenshots. Normalize persisted Unix-second timestamps at collection time and retain millisecond fallbacks.
Consolidates current-main-compatible work from #41156 and #48577.
Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>
Co-authored-by: tt-a1i <53142663+tt-a1i@users.noreply.github.com>
The Artifacts view reads message.timestamp, session.last_active, and
session.started_at from the SQLite database, which stores all timestamps
as Unix epoch seconds (REAL). These values were passed directly to
JavaScript's Date() constructor, which expects milliseconds — causing
every artifact timestamp to display as January 1970 dates.
Fix by multiplying the database value by 1000 to convert seconds to
milliseconds at the storage point, using nullish coalescing (??) instead
of logical OR (||) so that valid zero timestamps are not skipped.
Review fixes from #86679 comments (trevorgordon981, helix4u, kshitijk4poor):
- Edit inheritance: mergeConnectionInput preserves fields the editor does
not carry (cloud org, ssh remoteHermesPath/remoteProfile) so a rename no
longer wipes them. When the payload carries an ssh host string, stored
user/port are NOT inherited — the composite host field is authoritative,
fixing the stale user/port resurrection on edit.
- Token hygiene: tokens only persist on token-auth remotes; switching an
entry to oauth (or cloud) clears the stale envelope.
- Plain-text opt-in: the panel now surfaces the same consent dialog as
Settings -> Gateway on keyring-less machines (registry list exposes
secureTokenStorage; save retries with allowPlainTextToken after consent).
- Registry test isolation: hermes:connections:test builds the probe directly
from the registry entry instead of coercing against v1 connection.json —
no more inheriting the v1 global token for a different host, and the local
entry now probes the app-managed backend (never v1 remote/ssh state, so
the test button can no longer trigger a v1 file write).
- 'local' id reserved at the validation boundary: a crafted IPC payload can
no longer replace the local entry via upsert.
- Cloud creation hidden in the editor (a dialable cloud entry comes from the
Cloud sign-in/discovery flow); migrated cloud entries stay editable.
- First-run migration write is guarded: a failed write keeps the migrated
registry in memory instead of hard-failing every connections IPC call.
- uniqueLabel(): single label-dedup helper — counts up instead of "X 2 2",
clamps 253-char migrated URL-host labels under LABEL_MAX; used by
normalizeRegistry and both migration paths.
- UI copy: staged-rollout note replaces the "side by side" claim; test
failure toast leads with the failure wording; dropped unused i18n keys.
Tests: +9 pure-module cases (reserved id, token-drop rules, merge
inheritance, ssh host precedence, uniqueLabel); electron+settings suites
1355 passed.
First slice of multi-source agent support: the desktop can now persist ANY
number of named backends (local runtime, remote gateways, Hermes Cloud
instances, SSH hosts) side by side instead of one global connection plus
per-profile overrides.
- electron/connection-registry.ts: pure v2 registry module — required
case-insensitively-unique labels (device names), @name-device handle rule
for duplicate profile names across sources (agentHandle), defensive
normalizeRegistry for corrupt files, one-time v1→v2 migration that imports
the global block + per-profile overrides (deduped by URL/host) and leaves
connection.json untouched for older builds.
- main.ts: connections.json storage beside connection.json (same secret
posture: safeStorage-encrypted tokens, 0600, tighten-before-parse, mtime
cache) + hermes:connections:* IPC (list/save/remove/set-primary/test).
Test maps registry entries onto the existing testDesktopConnectionConfig
probe stack — no new probe code.
- Settings → Connections: manage the registry (add/edit/remove/test/make
primary) with forced naming; local entry is non-removable; removing the
primary retargets to local. en + zh locales.
Storage-level only by design: routing/pool generalization to composite
(connection, profile) keys, the multi-source roster, plugin SDK surface, and
fan-out updates land as follow-up PRs.
The composer's "Use skill: X" pill only checked the draft text and the
workspace-name collision - it happily re-offered a skill the session had
already loaded (skill_view), edited (skill_manage), or that the user had
invoked via its /name command. Clicking it would re-inject the full
SKILL.md into a context that already carries it.
The draft provider now scans the session transcript (per-runtime
$sessionStates mirror, $messages for the active session) for
skill_view/skill_manage tool calls naming the skill - exact or qualified
(category/name, plugin:name), from parsed args or hydrated argsText -
and for user turns starting with the skill's slash command, and stands
down on a hit. The scan only runs when the draft actually matched a
skill, so ordinary typing pays nothing.
Three fixes for generated/displayed images in the desktop chat:
- Shell fallback context menu no longer swallows right-clicks on images:
the guard now yields to Electron's native image menu (Copy Image, Copy
Image Address, Save Image As...) for img/picture/video/canvas targets,
matching the existing editable/selection carve-outs.
- Save Image As / download button: generated-image URLs (fal.media etc.)
end in an extensionless content hash, so saves produced an unopenable
"All Files" blob. The main-process save dialog and the renderer anchor
fallback both now append a MIME-derived extension, add image type
filters, and default to the user's Downloads directory instead of the
process cwd (win-unpacked on packaged Windows installs).
- New will-download handler routes any Chromium-initiated download
through the same Downloads-dir + guaranteed-extension policy.
Validation: new unit tests for the filename derivation (6 passing);
npm run check:lint green (tsc x3 + eslint, 0 errors).
The update hand-off spawned the detached updater, called unref(), and
quit unconditionally after the 2.5s dwell. Node reports exec failures
(ENOENT/EACCES) asynchronously via the child 'error' event, and a
short-lived updater can die inside that window — in both cases the app
vanished with no updater, no relaunch, and no evidence (the reported
macOS incident, and the posix.sh early-death reports on the same
thread).
Add observeUpdaterHandoff(): watch the just-spawned child for 'error'
and early 'exit' during the existing dwell (no added latency — the
dwell doubles as the settle window). Clean exit 0 inside the window
stays a success (the Windows `cmd start` wrapper exits immediately by
design); a spawn error, non-zero exit, or signal death is a failed
hand-off. On failure:
- applyUpdates (Windows hand-off): don't quit — restart the backend and
surface a structured error to the UI.
- applyUpdatesPosixHandoff (mac/linux): don't quit — surface the error.
- handOffWindowsBootstrapRecovery: return false so the caller falls
through to its next recovery path instead of quitting into nothing.
The pre-written update marker names the dead child pid, so
readLiveUpdateMarker self-heals it; no marker cleanup needed. Children
without an event interface settle ok after the window, keeping the
observation a best-effort hardening rather than a new way to wedge an
update.
Covered by 7 new unit tests (spawn-error, non-zero exit, signal death,
clean exit 0 wrapper, survival, double-settle, event-less child).
Closes#66753
The desktop window opened blank white on a fresh install: React threw
"Minified React error #527" before the first paint, from the
`vendor-react-<hash>.js` chunk.
`apps/desktop` pins react and react-dom to the same exact version, but
`vite.config.ts` aliased both to a hardcoded `../../node_modules/<pkg>` —
straight into the monorepo root, where npm is free to hoist a different
react. `@streamdown/math` is a root dependency whose react peer is
`^18.0.0 || ^19.0.0` and which declares no react-dom peer, so npm hoists
the newest react (19.2.8) to the root while react-dom stays at the
version hoisted from the workspaces (19.2.7). react-dom's own peer is
`react: ^19.2.7`, which 19.2.8 satisfies, so the install reports success
and nothing warns. The bundle then shipped react 19.2.8 with react-dom
19.2.7 and React refused to run.
`npm ci` masks this because the lockfile pins the root react to 19.2.7,
which is why CI is green. The recurrence engine is
`_run_npm_install_deterministic()`: when `npm ci` fails it falls back to
`npm install --no-save`, which re-resolves the whole tree and never
records the result — so the split comes back on the next update and
leaves no trace.
Fix the resolution rather than the hoist. The aliases now resolve both
packages from the desktop workspace itself, where npm guarantees the
declared versions are reachable (it nests a copy under the workspace
exactly when the hoisted one differs), so the pair can only ever match.
Pinning react at the npm layer instead was rejected: every manifest-level
pin tried (root dependency, root `overrides`, a scoped override on
`@streamdown/math`) breaks a fresh install with
`ERESOLVE ... peer ink-text-input@"6.0.0" from @hermes/ink@0.0.1`.
Two guards keep it from silently returning:
- `assert-root-install.mjs` (the existing preflight of `build`,
`dev:renderer` and `preview`) now fails the build when the resolved
react and react-dom versions differ, so a split surfaces as an
actionable error instead of a white window.
- A `tests-js` contract test asserts every workspace pins the two to the
same exact version, and that the desktop bundler no longer points at a
hardcoded `node_modules` path.
Verified on a synthetic split tree (root react 19.2.8 / react-dom 19.2.7,
workspace react 19.2.7): the old aliases resolve 19.2.8 + 19.2.7, the new
ones resolve 19.2.7 + 19.2.7, and the preflight exits 1 when the
workspace itself resolves the mismatch.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sibling site of the idle-resume rule from the stale-fold fix: the
assistant-tail append exit (user row persisted, no projection row) still
carried the journal's streamId onto a not-running resume, which kept the
journal entry alive (persistInFlightTurnState only clears when streamId is
null) and re-folded the same tail on every open. Apply the same
keepPending gate and pin it with a regression assertion.
Refs #85308
The inflight-turn journal can outlive the turn it recorded (reclaim,
reconnect or restart races skip the settle that clears it). On session
resume the fold then re-appends journaled assistant rows to a transcript
that already holds the committed replies, so the conversation ends with
duplicate answers in scrambled order. The fold also carried the stale
entry's streamId onto the resumed state on an idle resume, which kept the
journal entry alive (persistInFlightTurnState only clears when streamId is
null) and re-folded the same tail on every open.
Detect text-level staleness before the append path: when every recoverable
journaled assistant row already exists as committed text in the base
transcript, treat the entry as caught up and clear it. Only keep a stream
target when the resumed session is genuinely running (keepPending), so an
idle resume self-heals instead of re-folding.
The inflight journal regression test always spied on Storage.prototype,
but Node 26's jsdom setup can provide a plain in-memory localStorage fallback.
That left the test unable to observe the setItem call in CI even though the
per-session journal write was correct.
Select the native window.Storage prototype when available and otherwise spy
on the active localStorage object, preserving the assertion across both
storage implementations.
Refs #82832
The desktop journal synchronously read, parsed, cloned, and rewrote one
aggregate localStorage value while streamed turns were repainting. Large
tool results and multi-session state could therefore block the renderer and
leave the app unresponsive, while the existing macOS diagnostic path lacked
a real native hide/restore regression check.
Store bounded recovery projections under per-session keys, migrate legacy v1
data once, isolate quota and storage failures, and preserve the newest
recoverable tail without allowing oversized writes to replace valid state.
Add a real Electron/CDP macOS-arm64 A/B harness with native visibility control,
renderer heartbeat and Settings/composer/transcript checks, plus focused
regressions. Keep bulk tool payloads, diagnostics, and the existing recovery
merge behavior out of the hot path.
Fixes#63047
- Shared per-job trigger controller (apps/shared) coalesces duplicate
clicks for the same profile+job inside a mounted client while letting
unrelated jobs run independently; the backend durable claim remains
authoritative across windows/processes.
- Two-phase feedback everywhere: the action stays disabled/spinning
while the request is in flight and the terminal success/error is
reported once, after the HTTP response — no premature success toast
(Web), matching the Desktop info notification.
- Desktop keeps the 24h trigger timeout for the synchronous long
operation and fences stale profile/list responses and unmounted
surfaces; the sidebar trigger button shows a spinner while busy.
Fixes the #70449 symptom that remained after the two salvaged commits:
opening or viewing a chat whose turn is still running cleared its working
indicator. `session.activate` / `session.resume` report `running` as a
snapshot taken when the RPC was issued; a turn that started or kept
streaming while the RPC was in flight has already marked the runtime busy
in the live cache, and both resume paths overwrote that newer truth with
the stale `running: false`, dropping the session out of the working set
and painting it done mid-turn.
Add `resolveResumedBusy`: a snapshot saying running always wins (adopting
a live turn is never stale), but a snapshot saying idle can no longer
rewind a live busy — the turn's own terminal signal (running:false via
session.info / the settle path) stays the only authority that ends it,
and the background-sync reaper still clears truly lost turns. Wired into
both the warm `session.activate` path and the cold `session.resume` path,
reading the freshest cache entry rather than the pre-await state.
Includes an eslint --fix formatting pass over the touched files.
Salvaged from #51358 (razultull), rebuilt against the rewritten status
architecture on main. The original PR patched setSessionWorking /
noteSessionActivity / the statusbar counter, all of which have since been
replaced (busy now lives in session-states.ts, the watchdog only marks
stalled, and the statusbar Agents item already shows a pure subagent count
— the conflated counter the PR split no longer exists).
What still applied is the core bug: a parent that delegates via
delegate_task(background=true) ends its own turn the moment the handle
returns, so the sidebar row dropped to a plain idle dot while the spawned
subagents kept working for minutes — the session read as "done" mid-task.
Add $delegatingSessionIds — sessions whose subagents are still queued or
running — as an input to the session dot projection, claiming the same
'background' treatment as running background processes (and yielding to
'working' while the parent turn itself is live). Uses the same
runtime→stored bridge, lineage aliasing, and fresh-chat runtime-id
fallback as $backgroundRunningSessionIds, and clears by construction the
moment the last subagent reaches a terminal status.
Three safeguards keep finished chats from looking busy: tool rows seal on turn settle, vanished runtimes clear awaiting state and open tool parts, and late stream events no longer land in a freshly opened session.
The pane-resize sash's 9px grab band was centered on the split boundary,
so ~4.5px reached into the leading pane and sat exactly on top of the
session list's 4px scrollbar — the pointer always hit the sash (cursor
flipped to col-resize) and the scrollbar thumb was unclickable/undraggable.
Make the grab band asymmetric: 1px into the leading pane, 7px into the
trailing one. The scrollbar regains its full hit area while the sash stays
an easy 8px target; the hairline and hover strip are repositioned onto the
actual boundary.
Fixes#79157
Fixes#73495. Two cold-start defects made the configured Hermes Cloud
agent vanish after a Desktop restart even though the persisted Portal
session was still renewable:
1. hasLivePortalSession() trusted the FIRST cookies.get() on the lazy
`persist:` partition. It now reuses the warmOauthCookieStore()
warm-up + bounded reread that hasLiveOauthSession() gained in
PR #67769, so a single hydration false-negative no longer clears the
agent list and flips the panel to signed-out.
2. Discovery required the short-lived `privy-token` access cookie but
treated its absence as a full interactive re-login, even when the
30-day `privy-session` / `privy-refresh-token` renewal cookies
survived the process exit. New cookiesHavePrivyAccessToken() splits
"signed in (renewable)" from "discovery can succeed right now";
discoverCloudAgents() and cloudAgentSilentSignIn() now mint a fresh
access token via one bounded, hidden, deadline-capped portal load
(renewPortalAccessSilently) before or after a 401, and only surface
needsCloudLogin when renewal genuinely cannot complete.
PRIVY_SESSION_COOKIE_VARIANTS also learns `privy-refresh-token` so a
renewal-only jar still counts as signed in rather than demanding an
interactive login while usable refresh material sits in the partition.
Tests: connection-config.test.ts covers the access/session split,
including the exact renewal-only cold-start jar from the issue repro.
A message typed while a turn streamed rendered ABOVE assistant output the
user had already watched arrive (#73793), and the retired
insert-before-the-active-reply fallback could splice the bubble mid-thread
— halfway up the chat — when the stream id was missing or stale (#83151).
Fix the class at every path that assigns a transcript position to a
mid-turn user message:
- New shared appendMidTurnUserMessage (rewind.ts): seal the live stream
bubble in place (interim), append the correction at the live tail, and
clear streamId so post-redirect deltas seed a fresh bubble BELOW the
correction. Used by both the primary composer redirect path
(use-prompt-actions) and the session-tile steer path
(session-tile-actions), replacing the insert-before splice and its
last-assistant mid-thread fallback.
- appendLiveSessionProjection now projects the resume/reload turn in
arrival order (prompt → streamed output → correction → post-redirect
output) instead of prompt → corrections → reply, so the projection
agrees with the live transcript and messages no longer jump upward on
reconnect. With the gateway's new correction_offsets the flat dump is
split at each accepted-correction boundary; without offsets the
corrections follow the projected reply.
- tui_gateway/server.py records correction_offsets (assistant text length
at each accepted correction) on the inflight turn and carries them in
_inflight_snapshot, only when complete, so resume can rebuild true
arrival order. Older gateways/clients degrade cleanly.
- preserveLocalPendingTurnMessages and the projection's latest-user-run
matcher now treat a live-tail assistant row between the prompt and its
correction as part of the same turn's run, so arrival-ordered runs
survive refreshes without dropping the prompt.
Fixes#73793. Fixes#83151.
Completes the sidebar order/visibility class on top of the three salvaged
contributor commits:
- mergeSessionPage (#47203): interleave survivors against the
title-preserving merged rows using the backend's effective-recency key
(last_active with a started_at fallback), tie-preferring survivors so
keep-set rows with no timestamps retain the old prepend contract.
- sidebar order helpers (#73314): dedupe live ids as well as persisted ids
in reconcileFreshFirst/reconcileOrderIds/orderByIds so the shared-git-root
flatMap path can neither render one repo once per project nor write the
duplicates back into localStorage (the persisted feedback loop).
- Pinned section (#85969): resolvePinnedSessions falls back to the server
`pinned` flag when the localStorage pin set is cold or clobbered, so a
backend-pinned row is never simultaneously filtered out of every list and
absent from the Pinned section (the "session vanishes entirely" state).
session-pin-sync then adopts the pin locally on its next reconcile.
Regression tests cover survivor interleaving with optimistic bumps and
started_at fallback, duplicate live/persisted id dedup, and pin resolution
fallback (cold cache, lineage-root pins, undefined flag on old backends).
The desktop sidebar persists repo/lane order in localStorage
(hermes.desktop.workspaceParentOrder / workspaceOrder). If that saved
list ever contains the same id twice, orderByIds() pushes the matching
item once per occurrence, rendering the same repo header twice inside a
project. reconcileFreshFirst() then preserves the duplicates, so the
corruption self-perpetuates across restarts and storage clears.
Verified on a live install: the persisted workspaceParentOrder contained
the same repo path at two positions, the backend project tree was clean,
and the duplicated header matched the duplicated id.
Treat persisted UI order as untrusted input: orderByIds() now skips ids
it has already emitted, and reconcileFreshFirst() dedupes the retained
tail so the next persist writes a clean list (self-healing).
`refreshSessions` swaps the session page into `$sessions` only when
`sameCronSignature` reports a change, and that signature compared row
content — id, lineage root, title, source, profile, preview,
message_count, last_active, ended_at — but not row state. A page whose
only delta was `pinned` was judged identical and discarded, so the row
cached in the atom kept its old flag indefinitely. An idle conversation
never moves any of the compared fields again, which is exactly the kind
a user goes and unpins.
`session-pin-sync` treats that row as authoritative. Its write guard
(daeedf67c) is released by a page that CONFIRMS the value it wrote, and
falls back to letting the server win once WRITE_GUARD_MS elapses with no
confirmation. Because the confirming page was filtered out one layer up,
the fallback was the only branch that ever ran: ~10s after an unpin the
next reconcile read the frozen `pinned: true` row and called
pinSession() again. Adoption marks the id `mirrored`, so the push pass
never corrected the backend either — the local pin set and
sessions.pinned drifted apart permanently, which is why four of five
pins rendered in the sidebar read pinned=0 in state.db.
Compare both flags so a pin-only page reaches the atom. That restores
the guard's confirm path and makes WRITE_GUARD_MS a backstop again
rather than the load-bearing branch. `archived` is included for the same
reason: it is row state a consumer reads. Neither flag moves outside a
deliberate user action, so the churn the gate exists to prevent is
unaffected.
The existing `releases the guard once a page confirms the written value`
test passes on main because it hands `$sessions` the confirming page
directly — the gap was in the pipeline that decides whether such a page
is ever delivered.
Fixes#76919
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
When multiple sessions are active/settled simultaneously, survivors
(sessions the server omitted from the fresh page) were prepended as a
block in their old relative order from the previous $sessions array.
This caused recently-interacted sessions to appear below older ones.
Now survivors are sorted by last_active descending and merged into the
incoming array at the correct position using a two-pointer merge,
so the sidebar always reflects true recency.
Fixes#47203
Regression coverage for #68634. The indicator family mounts only on the
thread's any-role tail (ba756333): a running bubble that is not the tail
stays silent, even when only a user or system row trails it, and the
optimistic-placeholder flow renders exactly one status row, the
placeholder's own.
Mutation-checked: relaxing the mount gate to a last-assistant walk fails
the two silence cases, and removing the gate fails all five.
Clarify prompts share the same emitted-while-detached failure class the
pending-approval replay fixed: `clarify.request` rides `_block()`'s pending
registry, so a client whose transport was down when the event fired never
sees the question and the agent thread stays parked until timeout.
Widen the resume snapshot the same way:
- tui_gateway/server.py: `_live_session_payload` now carries
`pending_clarify` — a read-only snapshot of the clarify prompt still
blocking the session, scoped to the owning runtime sid. The registry stays
authoritative; the embedded request_id resolves via clarify.respond.
- Desktop resume paths (`use-session-actions`) restore the parked clarify
into the clarify store (multi_select preserved) and flag needsInput,
mirroring restorePendingApproval on both the activate and resume paths.
- pending_approval replay now also forwards the queue-injected request_id so
the restored prompt responds with exact-request correlation.
- Tests: server-side replay + scoping test; harmonized the #82087 replay
test with the request_id `_ApprovalEntry` now injects.
Reveal a blocking clarify card by re-arming the existing thread bottom-scroll bridge after the request row is hydrated. Keep background-session prompts isolated to their needs-input indicator.
Refs #53666.
Correlate approval requests, reject stale responses, replay pending approvals after reconnect or session resume, and preserve fail-closed timeout behavior.
Follow-ups on top of the two salvaged commits:
- widen the #83855 recently-interrupted cooldown to the session-tile
interrupt path (use-session-tile-delegate.interruptSession) — same race
class, sibling call site: a tile Stop also clears busy before the
gateway settles, so a quick tile edit/resend raced 4009 session busy.
The recovered runtime id is marked too.
- regression test for the tile cooldown.
- refresh three #65328 ownership-proof assertions to tolerate the
omit_messages flag main now sends on session.resume (toMatchObject).
- eslint import-order fix in utils.test.ts.
Fail closed on missing ownership cache entries and prove runtime ownership
forward+reverse against runtimeIdByStoredSessionIdRef before prompt.submit.
Thread the ownership cache through main wiring and session-tile submit.
Adds regression tests for forward mismatch, reverse-only proof, positive
map control, and cache-miss resume.
Closes#65328.
Stop clears frontend busy immediately while the gateway may still wind
down. Edit/restore then passed interruptFirst=false and raced 4009
session busy. Keep a short per-session cooldown after cancel so rewind
still interrupt-first, and expire the submit-in-flight lock so a hung
submit cannot block the session forever.
Fixes#83855
Co-authored-by: Olympusbuildz <Olympus.roots@outlook.com>
Signed-off-by: Olympusbuildz <Olympus.roots@outlook.com>
Two sibling sites still treated only 'completed'/'failed' as terminal:
- delegate-model.ts settled result rows as 'completed' for ANY status other
than 'failed', so a delegate result row with status 'timeout' or 'error'
(the statuses tools/delegate_tool.py actually emits on child timeout or
crash) rendered behind a green check. Settled rows now map ok/completed
to completed and everything else to failed.
- subagents.ts asStatus accepted a literal 'queued' payload status even on
a subagent.complete event, leaving the row active forever. The fail-closed
branch now runs before the queued fallback, so completion events always
settle.
Folded-in coverage from PR #80045 (gannotti, #80018): after a terminal
subagent.complete, a stray late 'running' progress event must not restart
the spinner — the upsert guard keeps the settled failed status.
Folded-in coverage from PR #85995 (smause): a subagent.complete event whose
payload still says 'running' or 'queued' must settle the row as failed —
the completion event itself is the source of truth that the child is done.
Review feedback: prev?.summary could shadow the 'Timed out after Xs'
reason when a live event had populated it. timeoutSummary() now wins for
raw timeout status; add coverage for the missing-duration placeholder.