Commit Graph

4806 Commits

Author SHA1 Message Date
m4 3ca91c55f3 feat(desktop): openExternalFileForIpc opens files via OS handler
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-09-17 22:28:07 +08:00
brooklyn! 1331053fcf style(desktop): anchor composer icon tooltips beside the control
Tips on the composer's row controls open to the left so they never cover
the fanned discs or the input; the model and reasoning pills keep theirs
centred above. A left-anchored Tip rags its wrapped lines toward the
trigger, in the primitive rather than per call site.
2026-09-16 02:56:57 -05:00
brooklyn! bc722dcb85 feat(desktop): fan the composer's voice toggles out of the mic
The docked composer spent three icon buttons on toggles that are set once
and rarely touched. The mic is now the one button in the row; hovering it
fans spoken-replies and the wake word out above it. The mic reports
dictation only — the wake word carries its own on-state on its own disc,
so mirroring it onto the mic read as dictation being on.

The wake disc's tip names the phrase and nothing else; the pressed state
says on/off. The folded HUD/narrow-tile VoiceMenu is unchanged.
2026-09-16 02:56:57 -05:00
brooklyn! c7580f0e06 feat(desktop): FanMenu primitive and a floating Button variant
One hub control that fans its sibling toggles out on hover — a column, a
row split around the hub, or an arc — with the discs portalled past any
overflow-hidden parent and re-anchored on scroll and resize. Hover state,
geometry and open/close timing live in the primitive so a consumer only
re-renders when its own items change.

A document-level pointer watchdog closes the fan when enter/leave is
skipped: a disc that goes disabled under the pointer (a toggle turning
pending right after its click) stops receiving pointer events, so its
pointerleave never fires and the fan stayed open.

Discs wear the new Button `floating` variant off (popover fill +
shadow-md, glyph-only hover) and `default` on, so a toggled state is an
opaque primary disc rather than a translucent tint.
2026-09-16 02:56:57 -05:00
brooklyn! 65b622da0d fix(desktop): transcript directives survive markdown, and the guide keeps its reasoning to itself
A handoff directive whose brief read as markdown (*by week*, a_b c_d, ~/x)
split the paragraph into element children, so the card never claimed it and
the raw ::onboarding{task="…"} line painted as the user's own message.
Directive lines are now backslash-shielded in preprocessMarkdown, next to
the inline-code and math shields, and reach the renderer as one text node.

The guide's reasoning is it reading its own runbook ("Now step 4: offer the
tour with ::ask"); shown under the greeting it breaks the conversation the
guide is trying to have. Reasoning disclosures stay hidden on the guide
thread only.
2026-09-16 02:49:21 -05:00
ouyangbo 012c9c6bed chore: anchor fresh-start history to upstream 2026-09-16 15:06:54 +08:00
brooklyn! 2e320e6d1a fix(desktop): keep inline edit placeholders off entered text 2026-09-16 02:04:58 -05:00
brooklyn! 2b7292ff0f style(desktop): frame inline subagents with a subtle outline 2026-09-16 01:44:33 -05:00
brooklyn! 3bdd4cc5fd fix(desktop): keep background continuations in one visual response 2026-09-16 01:44:33 -05:00
brooklyn! 57d34b14c2 fix(desktop): keep completed subagents from reviving stale activity 2026-09-16 01:44:33 -05:00
teknium1 10652c9345 fix(desktop): classify any 3xx as a redirect and name its Location
fetchJson/fetchPublicJson only reached the redirect diagnostic when the 3xx
carried an HTML body or text/html content-type; an empty-body 302/307 (the
common reverse-proxy / forward-auth shape) still resolved null through the
earlier empty-body check. http.request never follows redirects, so classify
on status first and reject every 3xx with the redirect error regardless of
body.

Pass res.headers.location through so the message says where the request was
sent, and stop blaming credentials when the Location differs from the
requested URL only by scheme or trailing slash -- that is a saved-URL
mismatch, not an authentication proxy. Drop the 404 mention from the
docstring/test: both callers reject >= 400 before this branch, so only the
2xx leg carries the endpoint-missing capability wording.

Part of #112072
2026-09-15 19:10:53 -07:00
teknium1 c902f50efb fix(desktop): send the connection extra gateway headers on remote media streams
createMediaProtocolHandler() set only the session token / bearer on the
/api/files/stream request. The descriptor it resolves now carries the
connection's extra gateway headers, but the handler never read them, so
attachments in session history from a header-gated remote still bounced off
the access proxy even though every fetchJsonForBackend() call got through.

Merge connection.headers into the media request before auth, letting the
forwarded range/cache negotiation headers win, and pin it on both the
token and the OAuth cookie-session legs.

Part of #112072
2026-09-15 19:10:53 -07:00
teknium1 47c3683f23 fix(desktop): stop calling a 3xx HTML reply a missing endpoint
The JSON guard in fetchJson/fetchPublicJson blamed every HTML reply on a
missing backend endpoint. A 3xx HTML body is an access proxy redirecting
to its login page: the endpoint exists, the credentials never arrived.
Besides misleading the user, the wording is the capability signal that
isMissingHealthEndpointError and the renderer's gateway-rpc predicate key
on, so an auth redirect was silently classified as "endpoint missing" and
routed onto compatibility paths instead of surfacing as an error.

Build the error in api-transport.ts (where the sibling httpStatusError
lives) and pick the hint by status: 3xx names the redirect and points at
the saved token/extra headers; everything else keeps the endpoint-missing
wording the predicates rely on.

Part of #112072
2026-09-15 19:10:53 -07:00
teknium1 e65ddf1c7b fix(desktop): keep primary remote gateway extra headers on REST calls
createPrimaryRemoteConnection() rebuilt the primary remote descriptor
field by field and left out `headers`, so every REST call routed through
fetchJsonForBackend() (Settings profiles/config, session history) reached
the gateway without the configured extra headers while chat, the Test
button and the readiness probe -- which read the exact-URL WebSocket header
store or the resolved route directly -- kept working. Behind an access
proxy (e.g. a service-token gate) those REST calls came back as a 302 to
the login page.

Carry `headers` through the descriptor like every other rebuild site
(buildRemoteConnection, the registry pool path and both ensureRegistryBackend
reuse branches already spread it), and pin it with an invariant test on the
existing primary-descriptor seam.

Fixes #112072
Co-authored-by: plluviera <plluviera@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 19:10:53 -07:00
teknium1 66f9f3692a fix(desktop): share the profile-switch freshness latch across font settings and MCP tab
Extract useProfileSwitchLatch(query stamps) and reuse it in ChatFontSetting,
TerminalFontSetting and McpTab instead of three divergent dataUpdatedAt
copies. The MCP tab keeps its release-on-fresh-error behaviour via the
optional errorUpdatedAt stamp; the font settings keep data-only semantics.
2026-09-15 19:10:25 -07:00
teknium1 92a6639e84 refactor(desktop): fold the font reseed latch into one dataUpdatedAt stamp
The previous commit tracks "waiting for the next config fetch" with a
`profilePending` state, a `staleConfigStamp` ref and a second effect that
clears the flag once `dataUpdatedAt` moves. The same guarantee fits in the
seed effect itself: the profile-switch handler stores the query's current
`dataUpdatedAt` as `staleStamp` and the seed effect refuses to reseed while
`dataUpdatedAt === staleStamp`. One state cell instead of state + ref +
effect, no `no-restricted-syntax` ref write in an effect, and the stale
guard is unchanged: the previous profile's cached record carries the
recorded stamp, so it can never repaint the new profile; only a fetch that
lands after the switch bumps the stamp and seeds.

Co-authored-by: danrudy33 <danrudy33@users.noreply.github.com>
2026-09-15 19:10:25 -07:00
KoNit-K df832c86a2 fix(desktop): reseed font controls after config refetch 2026-09-15 19:10:25 -07:00
teknium1 abdb402701 fix(mcp): carry the lazy status across the TUI wire, tests and docs
Follow-up to the ported status fix:

- `tui_gateway/contracts/tools_mcp_plugins.py::McpRuntimeStatus` is a
  closed wire enum; `mcp.servers.status` would raise `ContractViolation`
  on the new `lazy` value. Declare it and regenerate the TS/OpenRPC
  contract files.
- `ui-tui` session panel: an unknown status fell through to the red
  `failed` branch; render `lazy` with its cached tool count (inline
  branch, no component extraction).
- Two invariant tests, both red on origin/main: the real discovery path
  yields `status: lazy` with the cached tool count and a summary without
  `failed` (eager control stays `configured`, live control stays
  `connected`); a lazy-only run neither warns nor re-arms the startup
  retry, while a configured-only run still does.
- Document the per-server `lazy` key (undocumented until now) in
  `cli-config.yaml.example`, the MCP config reference and the MCP guide.
2026-09-15 19:06:54 -07:00
teknium1 837acdc91d fix: share the control-frame opener list and cover [System:/[IMPORTANT:/[PRIOR CONTEXT/[CONTEXT SUMMARY]
Review finding on #112260: the hosted-room member relabel regex was a third
hand-copied opener list that missed the frames context_compressor and
title_generator already treat as harness input, so a member reply starting
with "[System: ..." or "[IMPORTANT: 2 background processes ..." reached peers
in its exact trusted shape. agent.prompt_builder.CONTROL_FRAME_OPENERS is now
the single source; the gateway regex is built from it and the desktop TS
literal mirrors it byte-for-byte.
2026-09-15 19:04:59 -07:00
teknium1 9988545a78 fix(gateway,desktop): one visible relabel for member-quoted control frames on both room surfaces
Follow-up to the two cherry-picked contributor commits (#111571 gateway, #111576 Desktop),
which neutralized the same class with two different mechanisms: an invisible U+200B after
the `[` on the gateway path and a phrase replacement ("RESERVED CONTROL MARKER NEUTRALIZED")
on the Desktop path.

Both room-transcript builders now share one frame set (the mid-turn steer marker open/close,
the compaction handoff, runtime/system notes, planning-state and async-delegation frames —
the same openers agent/title_generator and agent/context_compressor already classify as
harness-authored, case-insensitive) and one visible relabel: the opener `[` becomes
`[member-quoted `. The peer still reads what the member wrote, but the exact trusted shape
the system prompt tells the model to honour is gone and the label says who authored it.
A zero-width space is invisible to a human reading the transcript and easy for a model to
skip over; the visible label is not.

Genuine user lines, stored room events and the displayed message are untouched on both
surfaces (probed live: stored member event byte-identical, `User (user):` line verbatim).
Tests trimmed to one invariant per surface, built from agent.prompt_builder's real marker
constants on the Python side.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
2026-09-15 19:04:59 -07:00
Hukla 498633f4ec fix(desktop): neutralize member control markers 2026-09-15 19:04:59 -07:00
teknium1 8550f9084e docs(profiles): Desktop cron ticker follows profile create/delete live
Also satisfy padding-line-between-statements on the two salvaged hunks.
2026-09-15 18:54:28 -07:00
teknium1 1220491468 chore(desktop): keep the slot-storm fix to the retry backoff
Drop the extra assertLocalProfileCanStart call added ahead of the slot
queue: the deleted-profile fence already runs at the spawn boundary below
and is not the mechanism behind the #111338 retry storm.
2026-09-15 18:54:28 -07:00
KoNit-K 68e4833134 fix(desktop): bound background profile hydration retries 2026-09-15 18:54:28 -07:00
teknium1 05051691c6 test(desktop): trim the archived-view reload coverage to two invariants
Keep the two discriminating cases (a sessions.changed tick reloads the archived
set while the view is open; a failed refresh retains the last good rows) and
drop the closed-view control, which passes on the unfixed base as well.
2026-09-15 18:54:00 -07:00
KoNit-K dbda53b8be fix(desktop): refresh archived sessions on external changes
Co-authored-by: DavidMetcalfe <80915+DavidMetcalfe@users.noreply.github.com>
2026-09-15 18:54:00 -07:00
teknium1 a40d90e9be test(desktop): trim the voice-live toast tests to two invariants
Seven per-case tests collapsed to two: one per behaviour class (our own
close reasons → copy, server reasons verbatim; mic DOMException → recorder
copy, non-mic failures untouched). Source is byte-identical to the
contributor commit. Dropped: the `closed`/blank-reason case, the
`close_requested` filter (pre-existing behaviour), and the per-name
DOMException matrix — the mapping table lives in `micError`, asserting one
name proves the live path routes through it.
2026-09-15 18:53:32 -07:00
finn763 c2625370d7 fix(desktop): stop the voice-live toasts leaking machine strings (#111987)
The session-end toast printed the raw wire reason (`connection_lost (127s)`,
`closed`) and a denied `getUserMedia` reached the toast as bare `DOMException`
text, while the recorder path already had friendly copy for those names.

- `liveEndedMessage` (new, exported for the tests) maps our own close reasons —
  `connection_lost`, `closed`, plus a blank reason — to i18n copy; server-sent
  reasons stay verbatim (unbounded, no redaction claim).
- `micError` is exported and reused on the live start path, so a mic
  DOMException gets the recorder's copy; an unmapped DOMException name now falls
  back to `microphoneStartFailed` instead of its raw text.
- The start catch maps DOMExceptions only: non-mic failures ('GPT-Live session
  already started', 'Missing local SDP offer', API errors) keep their message.

New i18n keys: `notifications.voice.liveEndedConnectionLost` / `liveEndedClosed`
(en base + zh translation; the other locales fall back to en).

Tests: `use-voice-live-conversation.test.ts` drives the real toast store through
a fake transport — reason copy, blank/unknown reasons, `close_requested` filter,
mic DOMException names, and untouched non-mic failures.
2026-09-15 18:53:32 -07:00
0xmosta 758638c64c fix(desktop): trap focus in boot recovery 2026-09-15 18:52:15 -07:00
teknium1 d747a71df2 test(desktop): type the fresh-module MCP health import without an import() annotation
The repo's eslint config forbids `typeof import(...)` type annotations
(@typescript-eslint/consistent-type-imports); derive the module type from the
dynamic-import thunk instead so `npm run check:lint` stays at 0 errors.
2026-09-15 18:51:49 -07:00
KoNit-K c9277d25b9 fix(desktop): honor MCP health snooze after restart 2026-09-15 18:51:49 -07:00
teknium1 b469be8cc3 fix(desktop): a push that lands mid-delivery claims its envelope now — lanes outlive the drain
The previous commit claims every outbox before delivering and runs
deliveries per target profile, but `drainBusy` still spanned the delivery
phase: a `bot_relay.outbox.pending` push that arrived while one lane ran a
long turn (up to RELAY_DELIVER_TIMEOUT_MS) only set `drainRerun`, and the new
envelope was claimed after that turn — the reporter's step 3 (bot C mails D
while A→B runs) still ended in `queued_expired`, because the gateway checks
the TTL at the claim.

Scope `drainBusy` to the claim phase and make the delivery lanes module
state: `relayLanes` maps `target_connection::target_profile` to the tail of
that target's in-flight deliveries, so a later drain appends to the running
lane (same target stays ordered, one turn at a time) or starts a new one
(other targets run now). Lane entries drop once idle; stopBotRelay clears
them so a restart begins fresh.

Tests: keep the contributor's red-on-base test (claims every outbox first,
delivers to different targets concurrently) and replace the ordering-only
test — green on base — with one that pins the mid-delivery claim plus the
same-target ordering (red on base AND on the previous commit alone).

Docs: bot-mode.md states the delivery concurrency contract.

Part of #111587 (with the previous commit: Fixes #111587)
2026-09-15 18:51:23 -07:00
John Paul Soliva 1eb771e2ff fix(desktop): the bot relay claims every outbox first and delivers per target, not one envelope at a time
drainRelayOutboxes drained one gateway's outbox and delivered its envelopes
before draining the next gateway, and delivered every envelope one after
another. One long turn — bot_relay.deliver may take up to
RELAY_DELIVER_TIMEOUT_MS, 25 minutes — therefore held every other bot's mail:
a sibling gateway's envelope sat unclaimed in its outbox, and since the
gateway checks the envelope's age against bot_mode.envelope_ttl_seconds
(15 minutes) at the claim, the sender's waiter received queued_expired for a
message nothing was wrong with; envelopes that did get claimed still waited
their turn behind unrelated deliveries, against a finite waiter.

Claim every gateway's outbox first, then deliver in lanes keyed by target
connection and profile: a lane runs its envelopes in order, one turn at a
time (the target gateway serialises that profile's turns behind its turn
lock anyway), and lanes run concurrently.
2026-09-15 18:51:23 -07:00
teknium1 672fd6b96f chore(desktop): drop a restating comment from the one-walk boot path
The WHY (one resolution instead of two) already lives on
resolveRendererIndexWithMissing; the call-site copy only repeated the code.
2026-09-15 18:50:24 -07:00
John Paul Soliva ab4bfda360 perf(desktop): resolve the renderer bundle once per window, not twice
createMainWindow walked the entire renderer generation twice before it
could call loadWindowUrl:

    const rendererIndex = DEV_SERVER ? null : resolveRendererIndex()
    const tornAssets = rendererIndex ? missingRendererAssets(rendererIndex) : []

resolveRendererIndex already computes exactly that list while choosing the
copy — it needs it to decide whether a copy is torn — and then throws it
away. missingRendererAssets is a BFS that readFileSync's every present
chunk whole and regex-scans it for the inline __vite__mapDeps table, so on
a release tree it is not a stat walk: measured against the real
apps/desktop/dist (252 chunks, 28.7 MiB of JS), one walk is 162
readFileSync calls reading 28.23 MiB, 576 existsSync calls, and 56.6 ms
median (min 55.8, 9 reps, warm page cache, darwin-arm64). Both walks run
synchronously on the main thread before the window gets its URL.

Return the list alongside the index. resolveRendererIndexWithMissing()
carries the existing body and hands back { index, missing }; the
path-only resolveRendererIndex() stays as a one-line wrapper so the nine
other call sites are untouched. The primary-window path takes one
resolution.

Semantics are unchanged in every branch: the same candidate is chosen, the
same log lines are emitted, and the missing list always describes the copy
actually returned. The all-copies-torn branch now reuses the first
candidate's list, captured on the first loop iteration, rather than
recomputing it for present[0] — recomputing there would have reintroduced
the second walk in exactly the case that matters most, and using the
loop's last value would have described a bundle we do not load.

Net effect on every primary-window boot: one fewer full walk, so 162 fewer
readFileSync calls, 28.23 MiB less synchronous reading, 288 fewer
existsSync calls, and ~57 ms of main-thread blocking removed before
loadURL. The win lands on packaged and --prod launches; DEV_SERVER skips
the walk entirely, so `vite dev` is unaffected.
2026-09-15 18:50:24 -07:00
teknium1 a22c731744 fix(desktop): refresh stale 4px scrollbar-width comments to 8px 2026-09-15 18:49:57 -07:00
teknium1 00c66225d4 fix(desktop): widen the portaled-menu scrollbar to match the app theme
`.dt-portal-scrollbar` is the same themed bar as `.scrollbar-dt`, applied
to overlays that portal under document.body (dropdown/context menus, the
command palette, the session and connection switchers). Widening only the
#root theme (#111634) would have left those lists on the 4px bar that was
too thin to grab; keep the two variants on one width (0.5rem = 8px).
2026-09-15 18:49:57 -07:00
kvnloo 22dc293f2a fix(desktop): widen themed scrollbars from 0.25rem to 0.5rem
The app-wide .scrollbar-dt theme (on #root) rendered 4px scrollbars
everywhere, including the conversation window, making them nearly
impossible to see or grab (#111634). Bump the webkit track size to 8px
so the thumb is actually hittable while staying a slim themed bar.

Fixes #111634
2026-09-15 18:49:57 -07:00
teknium1 b74f158b0d fix(desktop): do not prune the desktop half of a package folder the app cannot read
The ghost-prune loop in reconcileUnifiedDesktopHalves used existsSync on
the marker's source, which answers false for EACCES/EPERM as well as
ENOENT, so a mode-000 / ACL-denied package folder was treated as an
uninstall and its materialized half rm -rf'd with no warning.
Distinguish a genuinely missing source (ENOENT/ENOTDIR) from one the app
is not allowed to stat: keep the half and warn. materializeDesktopHalf
now also warns on a non-ENOENT stat failure instead of swallowing it.

Part of #111804
2026-09-15 18:48:59 -07:00
teknium1 b915405f92 fix(desktop): reconcile of unified plugin halves survives one unreadable package
`reconcileUnifiedDesktopHalves` let a stat/copy failure on one package
(`EPERM: lstat` on Windows in #111804; EACCES on a mode-000 file here)
reject the whole pass. The `hermes:fs:desktopPluginsRoot` IPC runs that
reconcile before returning the root, so the renderer never got a root and
every disk desktop plugin silently stopped loading. Warn about the one
package and keep materializing the siblings.

Part of #111804
2026-09-15 18:48:59 -07:00
teknium1 1cfa892db1 fix(desktop): make the zsh probe test legs visible and run them on CI
The zsh login-shell legs in remote-lifecycle.test.ts and
ssh-connection.test.ts silently returned when zsh was missing, and the
js-tests runner image ships no zsh, so the #111949 coverage never ran on
CI and a wrapper regression stayed green.

- js-tests.yml: install zsh on the Linux runner before the checks.
- Both legs now report vitest skips ('zsh not installed') instead of
  passing; the ssh-connection leg is its own test so the skip is visible.
- Docs: note the zsh degraded mode (no process-group kill for a hung
  probe's grandchildren) in the SSH connection guide.
2026-09-15 18:46:32 -07:00
teknium1 2388d401c9 test(desktop): capability probe through a real zsh login shell (#111949)
The user-visible failure was the ownership capability probe reporting a
current remote as unsupported when sshd ran it under a non-interactive zsh.
Pin the probe itself, not only the wrapper: run remoteSupportsSshOwnership
through `zsh -c` against a fixture CLI that advertises both flags; skipped
where no zsh is installed. Fails on the pre-fix wrapper (empty capture).

Co-authored-by: the-repeter <50600051+the-repeter@users.noreply.github.com>
2026-09-15 18:46:32 -07:00
KoNit-K 1267d7fdd0 fix(desktop): support zsh SSH probe watchdogs 2026-09-15 18:46:32 -07:00
teknium1 d30c05bebf fix(desktop): blank-line padding in the foreground-dial test 2026-09-15 18:46:08 -07:00
teknium1 5eb0ed4531 fix(desktop): Vault settings tab dials its scoped profile foreground
requestGatewayForProfile always dialed with the 'background' spawn priority,
so the Vault tab under the Settings 'Applies to' selector kept the #111651
infinite-spinner path on a cold profile even after the REST side moved to
scopedDialPriority. Let requestGatewayForProfile take a spawnPriority option
and pass 'foreground' from the Vault panel; ambient callers are unchanged.

Part of #111651
2026-09-15 18:46:08 -07:00
teknium1 d3ba5b307a fix(desktop): one scoped-dial priority helper; Capabilities selector dials foreground too
Fold the salvaged per-call ternaries in api/config.ts into a single
scopedDialPriority(scope) helper on api/client.ts and apply it to the
Capabilities scope selector's cold-start reads (getSkills / getToolsets /
getMcpCatalog), which hit the same background-capped pool queue when the
selector targets a stopped profile.

Drop the salvaged main.ts source-text change-detector test; keep only the
string update the existing #90812 wiring test needs. The renderer seam test
(hermes-capability-scope) is the invariant: an explicit scope carries
priority 'foreground', the ambient path stays untagged.

Part of #111651 (salvage #111672)
2026-09-15 18:46:08 -07:00
KoNit-K d6ff2777b7 fix(desktop): prioritize scoped settings backend dials 2026-09-15 18:46:08 -07:00
teknium1 2134950069 fix(desktop): keep the original boot error when post-spawn cleanup cannot prove ownership
In connect()'s post-spawn catch, a cleanupStale rejection (indeterminate
ownership probe) replaced the boot failure's message and kind. Catch it,
attach it as error.cleanupCause, and rethrow the original error.
2026-09-15 18:45:47 -07:00
teknium1 e368c01e28 fix(desktop): a lost SSH probe answer no longer kills or orphans a live remote backend
The Desktop SSH bootstrap proves the remote `hermes serve --isolated` alive
(`kill -0 … && echo ALIVE || echo DEAD`) and owned (argv probe printing
OWNED/FOREIGN) over the same SSH channel that is often mid-teardown right
after the served token was resolved. Both probes treated ANY answer that was
not the positive sentinel as the negative verdict, so an exec that resolved
with empty output was read as death: the boot failed with "remote dashboard
exited while its served token was being resolved", the post-spawn cleanup ran
the ownership probe on the same channel, read the lost answer as FOREIGN,
skipped the kill and removed the lockfile — one orphaned ~144 MB backend per
failed attempt, ten in one session on a 1 GB host (#111810).

`execProbeVerdict` now runs both probes: an answer that is neither sentinel is
indeterminate and is retried over a short bounded window (3 attempts, 500 ms);
with no definite answer it fails closed with a `transient-transport-error`, so
`cleanupStale` keeps the ownership record for the next connect to reap instead
of leaving a lockless orphan. The post-spawn catch probes liveness first so a
child that genuinely died at startup skips the ownership proof and the
original error ("exited before announcing") stays visible.

Slimmer redo of #111828 by @kokhlo: same mechanism, without the parallel
`verifyRemotePidAlive`/`verifyPidOwnership`/`readOwnershipVerdict` layer that
left `remotePidAlive` and the original probe as dead duplicates.

Fixes #111810

Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
2026-09-15 18:45:47 -07:00
teknium1 d4d9cf16d9 docs(desktop): record the partition-name invariant and the one-time re-sign-in
The salvaged comment promised the path component stays within
[A-Za-z0-9._-], but sanitizePartitionComponent() also emits the other
characters encodeURIComponent leaves alone (!~*'()). State the real
invariant — nothing Electron percent-escapes, pinned by the test as
"no ':' and no '%'" — and record the migration decision: one partition
name on every platform, so a non-primary cookie-auth remote signed in
on macOS/Linux under the old `:conn:` name is re-prompted once, and the
old `%3Aconn%3A` folder stays on disk, inert. apps/desktop/AGENTS.md gets
the rule so the next partition does not repeat #92183's folder name.
2026-09-15 18:45:13 -07:00