Commit Graph

4729 Commits

Author SHA1 Message Date
teknium1 d735097376 fix(desktop): keep createdBy=learn through the starmap share code
The share-code codec encodes createdBy through a fixed table that only
knew none/agent/user, so idxOf returned 0 for 'learn' and a learned
skill round-tripped as createdBy: null. The 2-bit field has a free slot,
so 'learn' now occupies it, and metaBadges labels agent- and
learn-created skills alike as 'learned' — matching the Python side,
where both values count as learning signals.

Review finding: learn nodes lost createdBy in share-code round-trip.
2026-09-15 05:38:31 -07:00
teknium1 ee6fb0ed22 fix: clear the relay roster per sole connection, not once per below-two regime
rosterCleared was a single boolean latched the first time the peer set fell
below two connections. When the sole connection a was replaced by c between
ticks (still length 1) the flag stayed set, so c's gateway never received
the empty-roster push and kept the stale roster. Track the id of the
connection that got the clear instead and push whenever the sole
connection's id differs; reset once two or more connections relay again.

Review finding: routes [a] → tick → routes [c] → tick pushed no roster clear to c.
2026-09-15 05:19:12 -07:00
teknium1 c02a269351 fix(desktop): an empty relay route list does not spend the one roster clear
relayConnections() returns [] before the registry loads (and whenever the
host bridge is missing). The below-two clear must not latch on that empty
list, or the single connection that arrives on the next tick never gets its
roster cleared. Gate the clear on exactly one connection; formats the
salvaged test file.
2026-09-15 05:19:12 -07:00
John Paul Soliva 8b091e539c fix(desktop): the relay clears the remaining gateways' remote roster when the peer set drops below two
syncRelayRosters returned early with fewer than two connections, so a machine
removed from the registry stayed in every remaining gateway's
bot_relay/roster.json — still listed in each bot's system-prompt roster and
still a message_agent target — until a second connection appeared again.
Push the now-empty roster once when the set shrinks below two (and once at
start with a single connection, for a roster left behind by an earlier peer);
union pushes resume when a peer returns.
2026-09-15 05:19:12 -07:00
teknium1 33cd423d56 fix: remote probe watchdog kills the probe's process group, not just its child
The watchdog SIGKILLed only the direct child of the wrapper. A remote
`hermes` launcher that runs the CLI without exec leaves the hung grandchild
alive after its parent dies — exactly the broken-launcher class from
#110478 — so the probe still orphaned a process on the remote. Enabling
job control (`set -m`) around the probe start puts it in its own process
group; the watchdog now kills `-$__htp` first and the direct pid as the
fallback for shells that cannot turn job control on without a tty (dash),
where behaviour is unchanged.

Review finding: watchdog SIGKILLs only its direct child; a non-exec launcher leaves the hung grandchild orphaned.
2026-09-15 05:18:28 -07:00
teknium1 b06071c0c9 test(desktop): make the remote-watchdog probe test host-safe
The live shell leg skipped nothing on Windows (no POSIX sh) and its orphan
sweep matched ANY `sleep 30` on the host, so a busy dev box or a parallel
test failed it. Skip on win32 like the other shell-shape tests here, and
sleep a per-run unique duration so the sweep can only see our own child.
The lockfile-skew guard now rejects a kill of any literal pid instead of
only `kill -9`, so the probe watchdog's own `kill -9 $__htp` no longer
weakens it.
2026-09-15 05:18:28 -07:00
Kevin Rajan 5724cb3002 fix(desktop): kill hung remote SSH probes via a POSIX remote watchdog
runSsh SIGKILLs the LOCAL ssh child on timeout, but the remote command keeps
running as an orphan (ppid=1). A hung remote CLI (e.g. a wedged
`hermes --version`) therefore accumulates orphans on every timed-out probe.

Wrap the two remote Hermes CLI probes (--version and the serve --help
ownership probe) in withRemoteTimeout(), a pure-POSIX remote watchdog
(macOS remotes have no GNU `timeout`): the probe runs as the watchdog's
direct child and is kill -9'd remotely after 15s, before the local
20s exec timeout fires. The sleeper's stdio is detached so its orphaned
sleep cannot hold the ssh channel open on the healthy path.

Fixes #110478
2026-09-15 05:18:28 -07:00
teknium1 bc59c39cfd docs(desktop): point the copy-control inset comment at the real scroller
The salvaged comment cited tool/fallback.tsx, which does not exist; name ExpandableBlock/CodeCardBody (the .scrollbar-overlay scroller) and log-tail.tsx (the 12px sibling) instead.
2026-09-15 05:15:42 -07:00
David Metcalfe 6e04bf8e59 fix(desktop): keep the code-block copy control clear of the scrollbar
The fenced code card's scroller (CodeCardBody / ExpandableBlock) carries
`.scrollbar-overlay`, which hands the card's right edge back to the
platform's ~15-17px scrollbar lane. The hover copy control sat at a 6px
inset with a 10px icon, inside that lane and on top of the bar itself.
Inset it 16px and use the 12px icon the other corner copy controls use
(log-tail.tsx).

Salvaged from #110548 without its className change-detector test; the fix
is pure CSS (Tailwind inset/icon size) and is verified by a static trace
against the .scrollbar-overlay scroller.
2026-09-15 05:15:42 -07:00
KoNit-K fccdcdaea8 fix(desktop): expose Codex compression auto-raise 2026-09-15 05:14:59 -07:00
KoNit-K ec5c975902 fix(desktop): refetch vault sources after closed-to-open remount
Per-query staleTime: 0 so a fresh installed:false cache still refetches when the Passwords page remounts disabled and the gateway opens later.
2026-09-15 05:14:15 -07:00
KoNit-K c89a4404fd fix(desktop): refresh vault source detection on mount 2026-09-15 05:14:15 -07:00
DavidMetcalfe 6201a8236f fix(desktop): drop unsaved credential edits when the settings target changes
The shared Settings "Applies to" target re-fetches `vars`, but the in-flight
edit and revealed maps are keyed by var name alone and were not reset with it.
A value typed while targeting one profile therefore survived a switch to
another, where the still-live Save button wrote it into the profile then being
targeted — the credential landed in the wrong profile.
2026-09-15 05:13:32 -07:00
teknium1 7c47e5517b test(desktop): trim the live-draft helper coverage to two invariants
Keep the case the fix exists for (editor holds text the mirror has not
seen yet) and the pre-mount fallback; the whitespace, stale-non-empty and
cleared-editor cases all reduce to "the helper returns composerPlainText
of the editor", which the first test already pins.
2026-09-15 05:12:33 -07:00
DavidMetcalfe 7679556d90 test(desktop): tighten the live-draft race coverage and dedupe the sibling reads
Tighten the empty-editor case against a stale non-empty mirror, add JSDOM body cleanup for editorWith, and route the bare-Enter and Cmd/Ctrl+Enter live reads through liveComposerDraft.
2026-09-15 05:12:33 -07:00
DavidMetcalfe bc62b270ed fix(desktop): read the live composer draft for the recall guard
draftRef is a once-per-frame mirror, so an ArrowUp in the same frame as a keystroke or paste saw the pre-keystroke text and let the sent-message recall overwrite what the user just wrote.
2026-09-15 05:12:33 -07:00
teknium1 e5d57c7fc5 refactor(desktop): workspace lanes read the show-all preference at the leaf
The salvaged fix threaded `showAllSessions` through three components
(sessions-section → entered-content → RepoFlatSection → workspace-group) as
a prop. The repo's TypeScript rule is that a leaf subscribes to the shared
atom instead of state being threaded through intermediaries, and
overview-row.tsx already reads `$sidebarShowAllSessions` that way. Drop
the prop plumbing and let SidebarWorkspaceGroup `useStore` the atom, which
also covers the bare `groups.map(SidebarWorkspaceGroup)` branch in
sessions-section.tsx that the prop version left paged.

The regression test flips the atom instead of re-rendering with a prop.
Still red against origin/main source, green here.
2026-09-15 05:11:46 -07:00
KoNit-K 38a4e909c6 fix(desktop): honor show all sessions in projects 2026-09-15 05:11:46 -07:00
teknium1 9b5f5a0c84 fix(desktop): local-runtime job poller no longer throws after the test env tears down
The 700ms re-poll in local-runtime-jobs.ts reached for `window.setTimeout`;
when a test left a running job in the store, the timer fired after vitest had
torn down jsdom and threw `ReferenceError: window is not defined` as an
unhandled error, failing the whole check:test:ui lane on main even though all
7808 tests passed. Use the global timer functions (identical in the renderer)
and drain the running-job state after each local-models-settings test so no
poll is armed across test boundaries.
2026-09-15 04:50:07 -07:00
teknium1 9c689bee1a fix(bot-mode): an empty member seat surfaces an error instead of swallowing the group send
Ported from the pre-TSX PR onto the current module layout. main's
group-chat-view.tsx already restores the composer draft when
sendToGroupChat returns null (feat(desktop): retain Bot group drafts by
room, b42d8279ed), so the draft-loss half of the original fix is
FIXED_ON_MAIN and not re-applied here. What remained: sendToGroupChat in
group-rounds.ts still folded "no text" and "no members" into one silent
`return null`, so a fully typed message into a room whose roster had not
hydrated (or a legacy room record without member descriptors) was
rejected with no thread, no log entry and no error.

The guards are split: empty content stays silent, an empty member seat
raises host.notify with a Bot Mode i18n string (en/ja/zh/zh-hant). The
wording no longer promises a retry will help (review: a legacy room with
no member descriptors never recovers by retrying) and points at the two
real remedies.

The original source-regex .mjs tests are dropped per review; the
contract is pinned by a behavioural vitest in group-rounds.test.ts that
drives sendToGroupChat through the scripted room harness: empty members
→ null + one error toast + no log entry; blank text → null, no toast.
2026-09-15 04:14:09 -07:00
teknium1 81e89ad1c1 test(desktop): kanban block-loop title assertion follows the plain-language copy 2026-09-15 03:50:00 -07:00
teknium1 66878996dd fix(ux): plain-language, actionable user-facing messages (desktop-tui)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 03:46:45 -07:00
teknium1 ca5822b5a5 test(desktop): mark the loud scope note semantically instead of by class
Asserting the loud variant via the Tailwind `font-medium` class couples
the test to styling: any restyle of the accent breaks it without the
behaviour changing. Give the note role="status" (it announces which
profile edits land in) and a `data-scope-loud` flag for the non-default
variant, and assert those. Also fixes the padding-line lint warning in
settings-scope.test.ts.
2026-09-15 03:42:20 -07:00
teknium1 d85aaa3049 test(desktop): stub the new settings-scope atoms in the ConfigSettings test
config-settings.test.tsx mocks @/store/settings-scope with a hand-written
subset of exports; SettingsProfileScope now also reads
$settingsScopeProfile and $settingsScopeEditsNonDefault, so the partial
mock threw at render time in CI (JS & TS checks red).
2026-09-15 03:42:20 -07:00
teknium1 d7d9209795 fix(desktop): derive the loud scope note from the shared settings-scope store
Rewrite of the original component-local `editingNonDefault` computation
onto the existing scope architecture (owner's ask: "use the same
existing architecture" as the Capabilities/toolset scope handling).

- store/settings-scope.ts gains `$settingsScopeEditsNonDefault`, a
  computed over `$settingsScopeProfile` × `$profiles`, so "the settings
  pages are editing a non-default profile" is a store fact any surface
  can subscribe to, not a per-component recomputation.
- profile-scope.tsx reads `$settingsScopeProfile` for the selected chip
  (instead of re-deriving `override ?? active`) and the new selector for
  the loud/quiet note. User-visible outcome is unchanged: accented note
  for any non-default target, override or not; quiet note for an explicit
  override onto the default; nothing when following the active default.
- Review: an unloaded roster (no `is_default` entry yet) used to suppress
  the note exactly at the landing moment where a bot may already be the
  active profile. The selector now assumes the root profile's canonical
  key as the default, so an unknown default fails loud, not quiet. The
  four-cell truth table is documented at the JSX conditional.
- settings-scope.test.ts pins the selector across active-profile,
  override and unloaded-roster inputs; the component tests from the
  original PR are kept as-is and still pass.
2026-09-15 03:42:20 -07:00
Teknium b047b5db4c fix(desktop): settings pages state loudly when they edit a non-default profile's config
After any Bot Mode chat the active gateway profile is the bot's, so the
settings scope silently followed it — edits landed in
profiles/<bot>/config.yaml with only a faint chip tint as the tell
(#89190/#89162/#89597 report class, live-repro'd: Max Agent Steps written
to scout's config). The applies-to note now renders for ANY non-default
target, override or not, accented; default-profile editing stays quiet.
2026-09-15 03:42:20 -07:00
teknium1 8b9df066e2 test(gateway): trim block-loop wording coverage to two invariants; CLI keeps the needs_input distinction
The salvaged parametrized set collapsed to one neutral case (transient) plus
the needs_input positive control. hermes kanban block now mirrors the
notifier: 'needs a human decision' only when the block was typed
needs_input, 'orchestration attention needed' otherwise.
2026-09-15 03:42:00 -07:00
KoNit-K 5c970d9745 fix(kanban): neutral block-loop wording on the CLI, Desktop toast, wake text and docs
The sibling surfaces of the gateway ping rendered the same false claim:
`hermes kanban block` said "needs a human decision", the Desktop toast title
said "needs a decision", the wake status line (locales/*.yaml
gateway.kanban.wake.block_loop_detected) said "needs a decision" and the
docs described the triage route as "for a human decision". A repeated-block
circuit breaker only establishes that orchestration attention is needed.

Surface sweep from PR #111131 (notifier/test hunks dropped in favour of the
typed-kind formatter from PR #111132).
2026-09-15 03:42:00 -07:00
teknium1 0468acdce6 fix(desktop): key voice-fields autosave on the profile scope string
Parents (profile-config.tsx) build the `profile` scope as a fresh object
literal on every render, so `useMemo(..., [profile])` produced a new
cache writer each time and the autosave effect, which depends on both,
tore down and re-armed its 550ms timer on every unrelated re-render.
Key both on profileScopeKey(profile) — the same string the query cache
already uses as the scope's identity — so only a real scope change
re-arms them.
2026-09-15 03:41:04 -07:00
teknium1 ab1ef0c88e test(desktop): pin that a scoped Capabilities TTS panel saves into its scope
Review on the PR: the load-bearing direction (profile B's scope
forwarded into saveHermesConfigRecord) was only covered by the live
E2E; the unit test asserted just the unscoped default. Renders the panel
with profile={profile:'scout', connectionId:'gw-2'} and asserts the
autosave carries exactly that scope. The @/hermes full-replacement mock
gains profileScopeKey, which use-config-record reaches when a scope is
present. Sabotage (drop the profile prop from <VoiceProviderFields>)
fails this test with `expected undefined to deeply equal {profile:
'scout', …}`.
2026-09-15 03:41:04 -07:00
Teknium bdac5c8e2e fix(desktop): Capabilities TTS voice fields write the scoped profile's config, not the active one
ToolsetConfigPanel threads its profile scope into every fetch but rendered
VoiceProviderFields without it, and the fields were hard-wired unscoped —
configuring profile B's TTS from the Capabilities selector read and
autosaved profile A's whole config record. New capability-scoped
saveHermesConfigRecord (symmetric with getHermesConfigRecord), profile
prop threaded through, per-scope cache write-through. Unscoped callers
(Settings → Voice) unchanged.
2026-09-15 03:41:04 -07:00
teknium1 e860b8e4e4 fix(context): compute-host /context and session.context_breakdown carry the per-file manifest; report blocked files
Why: the tui_gateway live formatter (`_format_live_context_output`, used when
the session runs on a compute host) renders its own summary and never got the
"Context files" block, and `session.context_breakdown` had no structured rows,
so Desktop's popover could not show them. The formatter now appends
render_context_file_lines() with the session cwd bound (the RPC thread has no
session context, so the discovery walk would key on the backend's cwd), and
the RPC payload gains a `context_files` list (contract + generated TS/OpenRPC
+ Desktop type). The docs sentence is scoped to the surfaces that render it.

A file whose content _scan_context_content replaces with a BLOCKED marker was
reported "loaded"; the manifest now runs the same scan and reports `blocked`.
The module docstring names the frontmatter-strip / chain-cap approximations
and drops the product-name attribution (credit stays in the PR body).
2026-09-15 03:37:49 -07:00
kshitijk4poor e112da7578 docs(desktop): preflight comments describe the OAuth-only branch; scope assertion via objectContaining 2026-09-15 16:07:43 +05:30
kshitijk4poor 2821cb2d8c refactor(desktop): one profile-order sort for the active strip and the at-rest groups
sortByProfileOrder moves to a pure lib module with a key selector so
buildRestGroups sorts its named squares directly, replacing the collator
sort that the component then re-sorted with a different comparator. The
fleet rail test no longer needs importOriginal (and three store mocks) to
reach the helper.
2026-09-15 16:07:43 +05:30
kshitijk4poor 8751b3edd4 test(desktop): keep the unscoped getProfiles invariant separate from the scoped one 2026-09-15 16:07:43 +05:30
kshitijk4poor 9bb985c5d4 fix(desktop): OAuth REST preflight gets its own dial budget
Sharing the remainder of SWITCH_DIAL_TIMEOUT_MS meant a slow-but-successful
socket dial left the preflight 0 ms and the switch failed as "Timed out
connecting" although the socket had just opened.
2026-09-15 16:07:43 +05:30
kshitijk4poor b591df7c42 fix(desktop): at-rest rails read the roster only; active gateway keeps its single render path
The renderer's per-connection list no longer feeds buildRestGroups: with no
roster reconciler it could outlive a profile deleted elsewhere. The active
gateway is never rendered through FleetRestGroup — that path routed its own
squares through selectConnection (full dial + wipe) instead of selectProfile's
live swap. What survives from the original change is the order parity: at-rest
named squares follow $profileOrder like the active strip.
2026-09-15 16:07:43 +05:30
kshitijk4poor 256edfc1f5 fix(desktop): profile-list ownership follows the published source, without a roster reconciler
The per-connection list cache is kept only to repaint $profiles on re-home.
The $fleetRoster listener is dropped: a roster landing while the active
source's own /api/profiles read was in flight invalidated that read, so
$profiles stayed empty/stale after every switch or focus refresh that the
roster IPC won. Both are reads of the same backend; neither is "older".

A null descriptor is a reconnect blip (setConnection's contract) and keeps
the current owner instead of blanking the rail; the first published
descriptor adopts whatever list is already loaded. Legacy sources are keyed
by endpoint rather than a JSON tuple. Tests cover the roster race and the
null blip; the two use-session-actions tests now publish the descriptor
before seeding $profiles, matching the runtime order.
2026-09-15 16:07:43 +05:30
Zeus-Deus 5b49051276 fix(desktop): keep profile switches on the selected gateway 2026-09-15 16:07:43 +05:30
kshitijk4poor 8bdac1b17a fix(desktop): omit a null iss from the oauth.callback relay; hoist the loopback parse import
The Electron listener now always emits iss (null when the server sent none)
and McpOauthCallbackParams is extra="forbid", so a new Desktop against a
backend without this change would fail every remote MCP OAuth login with a
4000 - including providers that never send iss. Send the key only when set.
Also drop the deliver_callback_flow test the RPC test subsumes.
2026-09-15 13:00:12 +05:30
kshitijk4poor 1c243f86de fix(tui_gateway): relay RFC 9207 iss through the oauth.callback RPC
The oauth.callback handler parsed `iss` but never passed it to deliver_callback_flow, and McpOauthCallbackParams (extra="forbid") had no `iss` field, so the desktop renderer sending `iss: null` was rejected with 4000 "unknown key" — breaking every Desktop→remote-gateway MCP OAuth login. Add the field, forward it, and regenerate the OpenRPC/TS contract artifacts via scripts/gen_gateway_contracts.py.

Also update tests/hermes_cli/test_mcp_dashboard_oauth.py for the 3-tuple callback shape introduced by the cherry-picked commit (it was red on the stack).
2026-09-15 13:00:12 +05:30
OOOOOAO 1a6503a520 fix(mcp): thread RFC 9207 iss through every OAuth callback relay
mcp 2.x rejects an authorization response that omits the RFC 9207 `iss`
parameter when the authorization server advertised
`authorization_response_iss_parameter_supported`. Cloudflare advertises it
AND sends it; the CLI loopback handler has always forwarded it, but every
other callback producer parsed only code/state/error, so the SDK raised:

    OAuthFlowError: Authorization response missing iss parameter
    advertised by the authorization server

and the server parked. Same machine, same config, `hermes mcp login <name>`
from a terminal succeeded — the failure is specific to the non-CLI relays.

Forward `iss` on every producer, matching `_make_callback_handler()`:

- tools/mcp_dashboard_oauth.py: `deliver_callback()` accepts `iss`;
  `wait_for_callback()` returns `(code, state, iss)`. The bridge in
  tools/mcp_oauth.py already splats that tuple into
  `_authorization_code_result(code, state, iss)`, so it needs no change.
- tui_gateway/mcp_oauth_sessions.py: the gateway-hosted loopback listener
  parses `iss`, and `deliver_callback_flow()` forwards it.
- tui_gateway/methods_tools.py: the `oauth.callback` RPC passes `iss`.
- hermes_cli/web_routers/mcp.py: the dashboard callback route accepts it.
- apps/desktop/electron/mcp-oauth-callback-ipc.ts: the one-shot listener
  reads `iss` off the redirect (the renderer already spreads the whole
  callback object into the RPC, so it flows through unchanged).

Providers that omit `iss` round-trip as `None`/`null` rather than being
dropped, so servers that do not advertise RFC 9207 keep working.

Verified live on Windows against mcp.cloudflare.com, whose metadata sets
`authorization_response_iss_parameter_supported: true`: the server that
previously parked on the missing-iss error now reports
`Authenticated — 3452 tool(s) available` and `hermes mcp test cloudflare`
connects. State-mismatch and replay rejection are unchanged.

Tests (each fails on base, passes with the fix):
- test_dashboard_flow_preserves_rfc9207_iss
- test_deliver_callback_forwards_iss (client-redirect relay)
- test_loopback_listener_forwards_iss (real HTTP redirect)
- two vitest cases on the Electron listener, incl. the iss-absent case

Refs #92758, #99984. PR #92765 fixes the dashboard route and the loopback
listener but not the client-redirect relay
(`deliver_callback_flow` / `oauth.callback` / the Electron listener), which
is the path Desktop drives against a remote backend.
2026-09-15 13:00:12 +05:30
brooklyn! d128ce2e25 fix(desktop): keep sudo commands visible before password entry 2026-09-15 02:22:22 -05:00
brooklyn! b79107c565 fix(gateway): include command context in sudo password requests 2026-09-15 02:22:22 -05:00
Austin Pickett 5871d750bf test(bot-mode): pin thread-scoped session identity
The room-level test is the reported defect: two composer sends, two
threads, and neither backend transcript may contain the other thread's
prompt. Proven red on the unfixed code — both threads resolved to
sid-research-1.

Also pins same-thread continuity, the pre-thread adoption (first thread
continues the old conversation, later threads do not), and that the
member half of the key stays source-qualified so a remote `research` and
a local `research` never share a session.

Co-authored-by: Wenfengcheng <30426178+Wenfengcheng@users.noreply.github.com>
2026-09-15 02:12:21 -05:00
Austin Pickett 1ce9733113 fix(bot-mode): scope stop, harvest and clarify to the originating thread
The co-keyed readers of `room.sessions` had to follow session identity or
they would split from it: the stop interrupt targeted a member's only
session regardless of which thread issued the stop, and the stranded-reply
harvest resumed by bare member key even though its marker already carries
`{before, thread}`.

The clarify mirror moves too — with per-thread sessions one member can be
blocked in two threads at once, and a room-and-member key let thread B's
question silently replace thread A's card. The sweep's owner lookup stays
a MEMBER question, so it reads the member half of the key and keeps the
`::` source-qualifier check on that half alone.

Co-authored-by: Wenfengcheng <30426178+Wenfengcheng@users.noreply.github.com>
2026-09-15 02:12:21 -05:00
Austin Pickett d631ad5f5c fix(bot-mode): key group member sessions by thread
A group member's hidden plumbing session was keyed by member alone, so
every thread in a room collapsed onto one backend transcript: thread B
resumed thread A's session and answered carrying A's context.

Everything else in the room engine is already thread-aware — the delta is
filtered by `groupThreadOf(e) === thread` and the watermark is
`${thread}::${memberKey}` — so only session identity was missing. Key and
title it per thread, and adopt a room's pre-thread pointer into the first
thread that speaks so an upgraded room's history is continued rather than
orphaned.

Co-authored-by: Wenfengcheng <30426178+Wenfengcheng@users.noreply.github.com>
2026-09-15 02:12:21 -05:00
kshitijk4poor db64ddb58e test(desktop): widen the execProbe event-loop test timeout
The 'execProbe keeps the parent event loop available' case asserts nothing
about timing; the 1s spawn budget only exists so a wedged child cannot stall
the run. A cold Windows CI runner can take longer than 1s just to start the
node child, which failed the test for reasons unrelated to what it guards.
5s keeps the safety bound without the flake.
2026-09-15 10:50:53 +05:30
kshitijk4poor 6915e9b4bc fix(desktop): keep the pool slot leased when a spawned start is superseded
runPoolBackendStart reused assertPoolEntryStillOwned at every cancellation
checkpoint, and that helper releases the local backend slot before it throws.
That is right before spawn (no child, nothing to wait for) but wrong at the
post-spawn checkpoints (after the port announcement, waitForHermes, token
adoption, WS probe): the lease was handed back while the superseded child was
still running, so a successor could spawn into an occupied slot, and the
caller's teardownFailedLocalBackend -> releaseLocalBackendSlotAfterExit then
found nothing left to release and became a no-op. This breaks the
pool-spawn-coordinator invariant that a lease is held until the child exits
or the start fails.

Add a `releaseSlot` option to assertPoolEntryStillOwned and pass
`{ releaseSlot: false }` at every site where `entry.process` is set. The
pre-spawn sites keep releasing. The post-spawn release now happens only via
teardownFailedLocalBackend (after the child has provably exited) or the
child's own exit handler.

No vitest case: the helper and its callers live in main.ts alongside the
module-level backendPool / localBackendLifecycle state and are not importable
from a unit test without extracting them; the lease-after-exit ordering
itself is already pinned by pool-spawn-coordinator.test.ts ('a rejected wait
keeps the slot occupied').
2026-09-15 10:50:53 +05:30
kshitijk4poor 844e5f2610 fix(desktop): do not cache a timed-out serve-support probe
The serve-support resolver caches the probe outcome per resolved runtime
for the process lifetime. That is right for a genuine "unknown
subcommand" exit, but a probe that died by timeout says nothing about the
runtime — only that this machine was slow right then (cold AV scan on
Windows, first Python import after boot). Caching that as `false` routed
a modern runtime through the legacy `dashboard` form until the app was
relaunched.

Export isTimeoutError from backend-probes and evict the cache entry on a
timeout so the next check re-probes; non-timeout failures stay cached.
One vitest case covers timeout → re-probe; the existing case still pins
non-timeout failure → cached.

Also replace the last synchronous read on the discovery path
(dashboard.py fast path) with fs.promises.readFile, matching the stack's
goal of keeping runtime discovery off the main event loop.
2026-09-15 10:50:53 +05:30