fetchJson/fetchPublicJson only reached the redirect diagnostic when the 3xx
carried an HTML body or text/html content-type; an empty-body 302/307 (the
common reverse-proxy / forward-auth shape) still resolved null through the
earlier empty-body check. http.request never follows redirects, so classify
on status first and reject every 3xx with the redirect error regardless of
body.
Pass res.headers.location through so the message says where the request was
sent, and stop blaming credentials when the Location differs from the
requested URL only by scheme or trailing slash -- that is a saved-URL
mismatch, not an authentication proxy. Drop the 404 mention from the
docstring/test: both callers reject >= 400 before this branch, so only the
2xx leg carries the endpoint-missing capability wording.
Part of #112072
createMediaProtocolHandler() set only the session token / bearer on the
/api/files/stream request. The descriptor it resolves now carries the
connection's extra gateway headers, but the handler never read them, so
attachments in session history from a header-gated remote still bounced off
the access proxy even though every fetchJsonForBackend() call got through.
Merge connection.headers into the media request before auth, letting the
forwarded range/cache negotiation headers win, and pin it on both the
token and the OAuth cookie-session legs.
Part of #112072
The JSON guard in fetchJson/fetchPublicJson blamed every HTML reply on a
missing backend endpoint. A 3xx HTML body is an access proxy redirecting
to its login page: the endpoint exists, the credentials never arrived.
Besides misleading the user, the wording is the capability signal that
isMissingHealthEndpointError and the renderer's gateway-rpc predicate key
on, so an auth redirect was silently classified as "endpoint missing" and
routed onto compatibility paths instead of surfacing as an error.
Build the error in api-transport.ts (where the sibling httpStatusError
lives) and pick the hint by status: 3xx names the redirect and points at
the saved token/extra headers; everything else keeps the endpoint-missing
wording the predicates rely on.
Part of #112072
createPrimaryRemoteConnection() rebuilt the primary remote descriptor
field by field and left out `headers`, so every REST call routed through
fetchJsonForBackend() (Settings profiles/config, session history) reached
the gateway without the configured extra headers while chat, the Test
button and the readiness probe -- which read the exact-URL WebSocket header
store or the resolved route directly -- kept working. Behind an access
proxy (e.g. a service-token gate) those REST calls came back as a 302 to
the login page.
Carry `headers` through the descriptor like every other rebuild site
(buildRemoteConnection, the registry pool path and both ensureRegistryBackend
reuse branches already spread it), and pin it with an invariant test on the
existing primary-descriptor seam.
Fixes#112072
Co-authored-by: plluviera <plluviera@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Extract useProfileSwitchLatch(query stamps) and reuse it in ChatFontSetting,
TerminalFontSetting and McpTab instead of three divergent dataUpdatedAt
copies. The MCP tab keeps its release-on-fresh-error behaviour via the
optional errorUpdatedAt stamp; the font settings keep data-only semantics.
The previous commit tracks "waiting for the next config fetch" with a
`profilePending` state, a `staleConfigStamp` ref and a second effect that
clears the flag once `dataUpdatedAt` moves. The same guarantee fits in the
seed effect itself: the profile-switch handler stores the query's current
`dataUpdatedAt` as `staleStamp` and the seed effect refuses to reseed while
`dataUpdatedAt === staleStamp`. One state cell instead of state + ref +
effect, no `no-restricted-syntax` ref write in an effect, and the stale
guard is unchanged: the previous profile's cached record carries the
recorded stamp, so it can never repaint the new profile; only a fetch that
lands after the switch bumps the stamp and seeds.
Co-authored-by: danrudy33 <danrudy33@users.noreply.github.com>
Follow-up to the ported status fix:
- `tui_gateway/contracts/tools_mcp_plugins.py::McpRuntimeStatus` is a
closed wire enum; `mcp.servers.status` would raise `ContractViolation`
on the new `lazy` value. Declare it and regenerate the TS/OpenRPC
contract files.
- `ui-tui` session panel: an unknown status fell through to the red
`failed` branch; render `lazy` with its cached tool count (inline
branch, no component extraction).
- Two invariant tests, both red on origin/main: the real discovery path
yields `status: lazy` with the cached tool count and a summary without
`failed` (eager control stays `configured`, live control stays
`connected`); a lazy-only run neither warns nor re-arms the startup
retry, while a configured-only run still does.
- Document the per-server `lazy` key (undocumented until now) in
`cli-config.yaml.example`, the MCP config reference and the MCP guide.
Review finding on #112260: the hosted-room member relabel regex was a third
hand-copied opener list that missed the frames context_compressor and
title_generator already treat as harness input, so a member reply starting
with "[System: ..." or "[IMPORTANT: 2 background processes ..." reached peers
in its exact trusted shape. agent.prompt_builder.CONTROL_FRAME_OPENERS is now
the single source; the gateway regex is built from it and the desktop TS
literal mirrors it byte-for-byte.
Follow-up to the two cherry-picked contributor commits (#111571 gateway, #111576 Desktop),
which neutralized the same class with two different mechanisms: an invisible U+200B after
the `[` on the gateway path and a phrase replacement ("RESERVED CONTROL MARKER NEUTRALIZED")
on the Desktop path.
Both room-transcript builders now share one frame set (the mid-turn steer marker open/close,
the compaction handoff, runtime/system notes, planning-state and async-delegation frames —
the same openers agent/title_generator and agent/context_compressor already classify as
harness-authored, case-insensitive) and one visible relabel: the opener `[` becomes
`[member-quoted `. The peer still reads what the member wrote, but the exact trusted shape
the system prompt tells the model to honour is gone and the label says who authored it.
A zero-width space is invisible to a human reading the transcript and easy for a model to
skip over; the visible label is not.
Genuine user lines, stored room events and the displayed message are untouched on both
surfaces (probed live: stored member event byte-identical, `User (user):` line verbatim).
Tests trimmed to one invariant per surface, built from agent.prompt_builder's real marker
constants on the Python side.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
Drop the extra assertLocalProfileCanStart call added ahead of the slot
queue: the deleted-profile fence already runs at the spawn boundary below
and is not the mechanism behind the #111338 retry storm.
Keep the two discriminating cases (a sessions.changed tick reloads the archived
set while the view is open; a failed refresh retains the last good rows) and
drop the closed-view control, which passes on the unfixed base as well.
Seven per-case tests collapsed to two: one per behaviour class (our own
close reasons → copy, server reasons verbatim; mic DOMException → recorder
copy, non-mic failures untouched). Source is byte-identical to the
contributor commit. Dropped: the `closed`/blank-reason case, the
`close_requested` filter (pre-existing behaviour), and the per-name
DOMException matrix — the mapping table lives in `micError`, asserting one
name proves the live path routes through it.
The session-end toast printed the raw wire reason (`connection_lost (127s)`,
`closed`) and a denied `getUserMedia` reached the toast as bare `DOMException`
text, while the recorder path already had friendly copy for those names.
- `liveEndedMessage` (new, exported for the tests) maps our own close reasons —
`connection_lost`, `closed`, plus a blank reason — to i18n copy; server-sent
reasons stay verbatim (unbounded, no redaction claim).
- `micError` is exported and reused on the live start path, so a mic
DOMException gets the recorder's copy; an unmapped DOMException name now falls
back to `microphoneStartFailed` instead of its raw text.
- The start catch maps DOMExceptions only: non-mic failures ('GPT-Live session
already started', 'Missing local SDP offer', API errors) keep their message.
New i18n keys: `notifications.voice.liveEndedConnectionLost` / `liveEndedClosed`
(en base + zh translation; the other locales fall back to en).
Tests: `use-voice-live-conversation.test.ts` drives the real toast store through
a fake transport — reason copy, blank/unknown reasons, `close_requested` filter,
mic DOMException names, and untouched non-mic failures.
The repo's eslint config forbids `typeof import(...)` type annotations
(@typescript-eslint/consistent-type-imports); derive the module type from the
dynamic-import thunk instead so `npm run check:lint` stays at 0 errors.
The previous commit claims every outbox before delivering and runs
deliveries per target profile, but `drainBusy` still spanned the delivery
phase: a `bot_relay.outbox.pending` push that arrived while one lane ran a
long turn (up to RELAY_DELIVER_TIMEOUT_MS) only set `drainRerun`, and the new
envelope was claimed after that turn — the reporter's step 3 (bot C mails D
while A→B runs) still ended in `queued_expired`, because the gateway checks
the TTL at the claim.
Scope `drainBusy` to the claim phase and make the delivery lanes module
state: `relayLanes` maps `target_connection::target_profile` to the tail of
that target's in-flight deliveries, so a later drain appends to the running
lane (same target stays ordered, one turn at a time) or starts a new one
(other targets run now). Lane entries drop once idle; stopBotRelay clears
them so a restart begins fresh.
Tests: keep the contributor's red-on-base test (claims every outbox first,
delivers to different targets concurrently) and replace the ordering-only
test — green on base — with one that pins the mid-delivery claim plus the
same-target ordering (red on base AND on the previous commit alone).
Docs: bot-mode.md states the delivery concurrency contract.
Part of #111587 (with the previous commit: Fixes#111587)
drainRelayOutboxes drained one gateway's outbox and delivered its envelopes
before draining the next gateway, and delivered every envelope one after
another. One long turn — bot_relay.deliver may take up to
RELAY_DELIVER_TIMEOUT_MS, 25 minutes — therefore held every other bot's mail:
a sibling gateway's envelope sat unclaimed in its outbox, and since the
gateway checks the envelope's age against bot_mode.envelope_ttl_seconds
(15 minutes) at the claim, the sender's waiter received queued_expired for a
message nothing was wrong with; envelopes that did get claimed still waited
their turn behind unrelated deliveries, against a finite waiter.
Claim every gateway's outbox first, then deliver in lanes keyed by target
connection and profile: a lane runs its envelopes in order, one turn at a
time (the target gateway serialises that profile's turns behind its turn
lock anyway), and lanes run concurrently.
createMainWindow walked the entire renderer generation twice before it
could call loadWindowUrl:
const rendererIndex = DEV_SERVER ? null : resolveRendererIndex()
const tornAssets = rendererIndex ? missingRendererAssets(rendererIndex) : []
resolveRendererIndex already computes exactly that list while choosing the
copy — it needs it to decide whether a copy is torn — and then throws it
away. missingRendererAssets is a BFS that readFileSync's every present
chunk whole and regex-scans it for the inline __vite__mapDeps table, so on
a release tree it is not a stat walk: measured against the real
apps/desktop/dist (252 chunks, 28.7 MiB of JS), one walk is 162
readFileSync calls reading 28.23 MiB, 576 existsSync calls, and 56.6 ms
median (min 55.8, 9 reps, warm page cache, darwin-arm64). Both walks run
synchronously on the main thread before the window gets its URL.
Return the list alongside the index. resolveRendererIndexWithMissing()
carries the existing body and hands back { index, missing }; the
path-only resolveRendererIndex() stays as a one-line wrapper so the nine
other call sites are untouched. The primary-window path takes one
resolution.
Semantics are unchanged in every branch: the same candidate is chosen, the
same log lines are emitted, and the missing list always describes the copy
actually returned. The all-copies-torn branch now reuses the first
candidate's list, captured on the first loop iteration, rather than
recomputing it for present[0] — recomputing there would have reintroduced
the second walk in exactly the case that matters most, and using the
loop's last value would have described a bundle we do not load.
Net effect on every primary-window boot: one fewer full walk, so 162 fewer
readFileSync calls, 28.23 MiB less synchronous reading, 288 fewer
existsSync calls, and ~57 ms of main-thread blocking removed before
loadURL. The win lands on packaged and --prod launches; DEV_SERVER skips
the walk entirely, so `vite dev` is unaffected.
`.dt-portal-scrollbar` is the same themed bar as `.scrollbar-dt`, applied
to overlays that portal under document.body (dropdown/context menus, the
command palette, the session and connection switchers). Widening only the
#root theme (#111634) would have left those lists on the 4px bar that was
too thin to grab; keep the two variants on one width (0.5rem = 8px).
The app-wide .scrollbar-dt theme (on #root) rendered 4px scrollbars
everywhere, including the conversation window, making them nearly
impossible to see or grab (#111634). Bump the webkit track size to 8px
so the thumb is actually hittable while staying a slim themed bar.
Fixes#111634
The ghost-prune loop in reconcileUnifiedDesktopHalves used existsSync on
the marker's source, which answers false for EACCES/EPERM as well as
ENOENT, so a mode-000 / ACL-denied package folder was treated as an
uninstall and its materialized half rm -rf'd with no warning.
Distinguish a genuinely missing source (ENOENT/ENOTDIR) from one the app
is not allowed to stat: keep the half and warn. materializeDesktopHalf
now also warns on a non-ENOENT stat failure instead of swallowing it.
Part of #111804
`reconcileUnifiedDesktopHalves` let a stat/copy failure on one package
(`EPERM: lstat` on Windows in #111804; EACCES on a mode-000 file here)
reject the whole pass. The `hermes:fs:desktopPluginsRoot` IPC runs that
reconcile before returning the root, so the renderer never got a root and
every disk desktop plugin silently stopped loading. Warn about the one
package and keep materializing the siblings.
Part of #111804
The zsh login-shell legs in remote-lifecycle.test.ts and
ssh-connection.test.ts silently returned when zsh was missing, and the
js-tests runner image ships no zsh, so the #111949 coverage never ran on
CI and a wrapper regression stayed green.
- js-tests.yml: install zsh on the Linux runner before the checks.
- Both legs now report vitest skips ('zsh not installed') instead of
passing; the ssh-connection leg is its own test so the skip is visible.
- Docs: note the zsh degraded mode (no process-group kill for a hung
probe's grandchildren) in the SSH connection guide.
The user-visible failure was the ownership capability probe reporting a
current remote as unsupported when sshd ran it under a non-interactive zsh.
Pin the probe itself, not only the wrapper: run remoteSupportsSshOwnership
through `zsh -c` against a fixture CLI that advertises both flags; skipped
where no zsh is installed. Fails on the pre-fix wrapper (empty capture).
Co-authored-by: the-repeter <50600051+the-repeter@users.noreply.github.com>
requestGatewayForProfile always dialed with the 'background' spawn priority,
so the Vault tab under the Settings 'Applies to' selector kept the #111651
infinite-spinner path on a cold profile even after the REST side moved to
scopedDialPriority. Let requestGatewayForProfile take a spawnPriority option
and pass 'foreground' from the Vault panel; ambient callers are unchanged.
Part of #111651
Fold the salvaged per-call ternaries in api/config.ts into a single
scopedDialPriority(scope) helper on api/client.ts and apply it to the
Capabilities scope selector's cold-start reads (getSkills / getToolsets /
getMcpCatalog), which hit the same background-capped pool queue when the
selector targets a stopped profile.
Drop the salvaged main.ts source-text change-detector test; keep only the
string update the existing #90812 wiring test needs. The renderer seam test
(hermes-capability-scope) is the invariant: an explicit scope carries
priority 'foreground', the ambient path stays untagged.
Part of #111651 (salvage #111672)
In connect()'s post-spawn catch, a cleanupStale rejection (indeterminate
ownership probe) replaced the boot failure's message and kind. Catch it,
attach it as error.cleanupCause, and rethrow the original error.
The Desktop SSH bootstrap proves the remote `hermes serve --isolated` alive
(`kill -0 … && echo ALIVE || echo DEAD`) and owned (argv probe printing
OWNED/FOREIGN) over the same SSH channel that is often mid-teardown right
after the served token was resolved. Both probes treated ANY answer that was
not the positive sentinel as the negative verdict, so an exec that resolved
with empty output was read as death: the boot failed with "remote dashboard
exited while its served token was being resolved", the post-spawn cleanup ran
the ownership probe on the same channel, read the lost answer as FOREIGN,
skipped the kill and removed the lockfile — one orphaned ~144 MB backend per
failed attempt, ten in one session on a 1 GB host (#111810).
`execProbeVerdict` now runs both probes: an answer that is neither sentinel is
indeterminate and is retried over a short bounded window (3 attempts, 500 ms);
with no definite answer it fails closed with a `transient-transport-error`, so
`cleanupStale` keeps the ownership record for the next connect to reap instead
of leaving a lockless orphan. The post-spawn catch probes liveness first so a
child that genuinely died at startup skips the ownership proof and the
original error ("exited before announcing") stays visible.
Slimmer redo of #111828 by @kokhlo: same mechanism, without the parallel
`verifyRemotePidAlive`/`verifyPidOwnership`/`readOwnershipVerdict` layer that
left `remotePidAlive` and the original probe as dead duplicates.
Fixes#111810
Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
The salvaged comment promised the path component stays within
[A-Za-z0-9._-], but sanitizePartitionComponent() also emits the other
characters encodeURIComponent leaves alone (!~*'()). State the real
invariant — nothing Electron percent-escapes, pinned by the test as
"no ':' and no '%'" — and record the migration decision: one partition
name on every platform, so a non-primary cookie-auth remote signed in
on macOS/Linux under the old `:conn:` name is re-prompted once, and the
old `%3Aconn%3A` folder stays on disk, inert. apps/desktop/AGENTS.md gets
the rule so the next partition does not repeat #92183's folder name.
Electron escapes ':' in a session partition name as '%3A' for the on-disk
profile folder. On Windows, a profile folder whose name contains '%3A' gets a
cookie store the network stack can neither read nor write: a jar placed there
reads back zero cookies, `cookies.set()` never reaches disk, and every request
goes out with no cookies at all — the gateway answers 401 `no_cookie` and the
desktop concludes the user is signed out.
Per-connection partitions (#92183) are the only ones carrying a colon beyond
the 'persist:' prefix, so every NON-primary cookie-auth remote hits this: the
dial mints no ticket, the reauth copy fires ("Remote Hermes gateway uses
OAuth, but you are not signed in…"), and the login window — riding the same
partition — writes into the same invisible jar. Nothing ever persists, so the
prompt returns on every connect. The registry primary and the v1 remote keep
the legacy `persist:hermes-remote-oauth` jar, whose folder has no '%3A', which
is why a single-gateway setup never shows the symptom.
Verified against the app's own Electron runtime (40.10.2, Windows): identical
jar bytes seeded into two partition folders read 3 cookies and mint a
ws-ticket (200) from `…-conn-…`, and 0 cookies / 401 from `…%3Aconn%3A…`.
Keep the partition path component inside [A-Za-z0-9._-]; a regression test
pins that it never contains ':' or '%'.
- cmd_gui with DESKTOP_STARTUP_ID: the entry is installed only after the fake Electron
writes the reveal byte to HERMES_DESKTOP_READY_FD (spawn precedes install).
- cmd_gui without DESKTOP_STARTUP_ID (control): install still precedes the spawn.
- DeferredDesktopEntryInstall.finish(): an app that exits without revealing heals
exactly once, after the exit.
- linux-launcher-ready: one byte, fd closed, variable consumed; garbage/absent = no-op.
Docs: user-guide/desktop.md explains when the entry write happens for grid launches.
`hermes desktop` used to create/refresh `~/.local/share/applications/hermes.desktop`
synchronously before spawning Electron. When the entry is ABSENT (first run after an
update, deleted by a cleaner, tombstoned by AV) that write lands while gnome-shell still
has the grid-launched ShellApp in STARTING; unpatched shells (before GNOME MR !4428)
drop the app's last strong reference on `installed_changed` and the next idle GC kills
the whole Wayland session, minutes to an hour later (#111906, residual after #111396).
Now a launch that carries `DESKTOP_STARTUP_ID` (app grid / menu) defers the write:
the launcher opens a pipe, hands Electron its write end as `HERMES_DESKTOP_READY_FD`
(pass_fds), and a worker thread installs the entry once Electron reports the main
window revealed (`onRevealed` in createWindow, via the new
`apps/desktop/electron/linux-launcher-ready.ts`), plus a 2 s settle so the compositor
has mapped the surface. If Electron exits without ever revealing a window, `finish()`
installs the entry after the exit (STOPPED app, no STARTING object) — self-heal
semantics survive. Terminal launches, the updater's detached relaunch and
`--build-only` have no DESKTOP_STARTUP_ID / spawn no app and keep writing immediately.
Why not simply write after Electron exits (#111915's approach): a daemon thread
started right before `sys.exit` is killed with the interpreter, so the heal is lost,
and a heal that waits for the user to quit the app leaves the menu entry missing for
the whole session.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
A route tile restored (or opened) before its plugin route registered kept the
humanized-path fallback as its tab title forever: paneMirror recomputes titles
only when one of its atoms changes, and watchRouteTiles listened to $routeTiles
alone. #112147 made the pane CONTENT heal on late registration; the title
still lagged.
routes.ts exposes $routesVersion, an atom bumped from registry.subscribeArea
(ROUTES_AREA) only while listened to, and watchRouteTiles passes it as `also`
so the sync re-runs and the pane re-registers with the contribution's title.
One test, red on origin/main.
A remembered plugin-page route is session-shaped until its route registers,
and disk plugins load through an async IPC chain. When the backend was
already running (remote/URL connection, macOS close-reopen with the process
alive) the session list could arrive first; the restore effect then read
'/html-gallery' as a session id nobody owned, navigated to the last chat and
erased the remembered route, so the page was lost on every later boot too.
The disk door now publishes $diskPluginsScanPending for the duration of its
first scan, and the restore latch holds a session-shaped remembered route
while that scan is pending, mirroring the existing "sessions not loaded yet"
latch. Once the scan settles the route is either a registered page (restored)
or a genuinely stale session (dropped as before). Two tests: the late
registration restores; the stale session still drops.
Dragging a session tab out of the main strip into its own zone left the
workspace alone in main. Auto treated a lone pane as "not a tab", and the
uncloseable workspace is not stranded, so main lost its tab and its "+" —
while the tile it had just split from kept a strip of its own. Two chats
side by side, one with tabs and one without, reads as "the tabs
disappeared"; ⌥⌘T only helped because it wrote an explicit `always`.
The resolver gains one input, `siblingMainZone`: on auto, a lone main tile
keeps its strip whenever another zone in the layout also hosts a main tile.
A chat that is the whole window stays chromeless, side chrome is untouched,
and an explicit `never` still wins so Hide tabs keeps working.
Both callers read it from `$mainTileZoneCount`, a computed over the tree,
the hidden set and the registry version. It is a number, so zones re-render
only when a main zone appears or goes — never per sash-drag frame — and the
store's answer (what the toggle command flips against) and the strip on
screen cannot disagree.
Same class as the plugin-route bug: KeybindSettings subscribed to
useContributions(KEYBINDS_AREA) but discarded the snapshot and called the
impure allKeybindActions() in render, so the React Compiler memoized the
action list without the subscription as a visible input. A plugin keybind
registered after the tab mounted never appeared, despite the comment
promising "appear/disappear live".
allKeybindActions()/contributedKeybinds() take an optional contribution
snapshot (default: registry read, so the store/imperative callers are
unchanged) and the settings tab passes its subscription through. One test,
red on origin/main.
Drop the contributedRoutes() helper-contract tests (they pin the helper's
signature, not the user-visible failure) and the hot-replace/dispose
variants of the route-tile test: each surface now carries exactly one test
that is red on origin/main and green with the fix.
Regression test from #111541: mounts ChatRoutesSurface at a not-yet-registered
path through the real useContributions + registry under the compiler-enabled
vitest ui project, registers the route after mount, and asserts the page
replaces the `:sessionId` chat fallback. Red on origin/main, green with the
snapshot-fed contributedRoutes(). Replaces the equivalent surface test from
#109418 so the workspace surface carries one invariant test.