* feat(desktop): give Button a loading prop that swaps label for spinner without layout shift
The label stays in the box, invisible, and the spinner is absolutely
centred over it, so a Connect or Approve button keeps its width while it
works instead of collapsing to a spinner. The approval bar had the same
thrash and moves onto it.
* refactor(desktop): one consent card for connectors and MCP setup
McpSetupTool rendered its own copy of the connector card's markup. It now
renders ConnectorCard for the pending question and ConnectorSummary once
settled, and the card gains what MCP needed: keyboard accelerators, a
source line, a question heading. The card also gets an avatar variant
(40px mark in the left gutter, text and buttons on one column) and a
collapseWhenSettled switch so a connector can stay a full card with a
green Connected pill in the action slot while MCP keeps its one-line
summary. Brand marks for Gmail, Calendar, Drive, Discord, Telegram and
Spotify; Slack via Tabler because simple-icons dropped the mark.
* feat(desktop): connector card drives the agent through manage_connections wait
The offer used to end in a Continue in chat button, and the agent, seeing
an unconnected status, would improvise around the app. Now the card does
what the TUI does. Clicking Connect opens the browser and sends one hidden
line telling the agent to park in manage_connections action=wait for that
slug and to never call connect again (a second link cancels the one being
signed into). Not now sends its own line. A hidden request that lands
while the turn is busy steers it, or queues if the turn just ended.
Which call owns the live card changes too: consecutive calls naming the
same apps are one exchange (connect, the wait, the status that follows),
and the first of the last exchange is the card, so the agent's wait no
longer demotes the card mid-authorization and mints a fresh one below it.
A targeted ask renders one or two bare cards; only a real catalog gets the
header, search and refresh.
* feat(desktop): onboarding connects apps in chat and keeps tasks finishable without them
The welcome chat knew connectors only as preferences to pick and wire up
later, so asked to connect Gmail it invented a Settings page that does not
exist. Both scripts now carry one rule set: status once, one batched
connect for every app named, the card is the ask so write a line and end
the turn, never route around a declined app with another client or
credential. The build handoff checks real connection status instead of
asserting none are connected, and the first task must be finishable, not
free of, the apps they picked. The connectors card explains what
connecting means and reports the count on its Continue button.
* fix(tools): resolve the Nous identity for share_auth profiles in the connector gate
A profile created with share_auth has no auth.json of its own and signs
in through the root store. Every other credential reader falls back to
the global root; the connector gate read HERMES_HOME/auth.json directly,
saw nothing, and stripped manage_connections from the profile's tool
list, so the welcome chat's agent truthfully reported the tool missing.
The gate now goes through get_provider_auth_state.
* fix(agent): name a provider retry backoff on the live status line
The retry status is buffered and replays only when every retry fails, so
during a 60s backoff after a 5xx the user saw a bare spinner. Right after
a connector sign-in landed this read as the agent going silent. The
backoff now also rewrites the live wait notice, which the desktop already
renders in the thread status row; it is transient and clears on recovery.
* test(desktop): connector rehearsal launcher and flagged connector spec
connector-rehearsal.mjs starts the real desktop and backend under a fresh
HERMES_HOME with no copied credentials, a fixed Vite port and CDP on 9344,
so the onboarding connector flow can be driven end to end by hand or from
outside. The Playwright spec covers the flagged connector step.
* fix(desktop): send the agent back into wait when the user keeps waiting after a timeout
The card's Keep waiting re-entered the poll but the agent's own wait had
timed out too and nothing told it to go back in, so it would start
talking mid-authorization. keepWaiting now fires onWaiting like connect
does. Tests also pin that an expired or revoked grant asks the gateway
for reconnect, not connect.
* style(desktop): blank lines in connector-flow test per lint
* feat(desktop): HERMES_SKIP_INTRO=1 / --skip-intro skips the first-run film
The intro is a one-time reveal, so anyone rehearsing the guided chat behind
it sits through it on every fresh HERMES_HOME. The flag rides the existing
launch-flags path (main → preload → renderer) next to guestOnboarding and
only gates isIntroRevealEnabled; the backend never sees it. The rehearsal
launcher sets it.
* fix(desktop): onboarding card Continue stays Done after the transcript rebuilds
The card kept its Done flag in component state. The hidden submit and the
turn-end hydrate both rebuild the message list, so the card remounted with
the flag false and Continue came back live, letting a step be answered
twice. The committed steps now live with the other onboarding answers,
keyed by step, and the first-build chip pick rides the same store.
remember_onboarding projects by key, so the new field never reaches USER.md.
* fix(desktop): no provider picker or free-tier chip over the guided first launch
Two sign-in surfaces leaked into the guide. A credential probe on the
setup profile (a free-tier token mid refresh, a session before its runtime
settled) hit requestDesktopOnboarding and dropped the provider picker over
the chat the user was in; and the statusbar free-tier chip sat there
offering a second sign-in the whole time. Both now yield while the gate
phase is cinematic, guided or handoff. The free tier is the provider for
those phases, and the guide offers sign-in on its own ready screen.
* fix(desktop): onboarding connector picks are real catalog slugs
The picker offered Spotify, GitHub and Stripe, none of which the deployed
connector catalog carries, and spelled Calendar and Drive with hyphens the
gateway does not use. A pick the build chat could not honour ended as
"Spotify isn't in the connector list" after the user had been told to
expect it. The list is now twelve slugs from the live status catalog,
spelled as the gateway spells them; GitHub is out (the terminal has git
and gh), chat channels stay on Messaging. Marks for the new entries; the
Google marks answer both spellings. The build runbook offers the picked
connections in its first turn rather than after the work is underway.
* fix(desktop): the free-tier ready screen never interrupts the guided chat
A readiness round fires when the layout pick assembles the window, and it
raised the free-tier ready screen over the conversation: the user was
dropped into the main app, dismissed it, and came back to a card they had
already answered. The guide is the introduction. The ready screen now
yields while the gate is cinematic, guided or handoff, and the notice is
acked the moment the guided chat takes the screen, not only when the film
does, so a skipped film no longer leaves it pending.
* feat(desktop): tour options that lead to building, and a fork that follows the tour
"Just the basics" and "Show me around" read as a click-through with no
exit; "I'll figure it out" read as declining help. Now Quick tour, Show me
everything, and Skip, let's build something. The script also folds the
fork into the same turn as the tour, so when the user closes the overlay
the next ask is already waiting instead of a transcript that ends on the
tour call.
* feat(desktop): the onboarding connector picker reads the live catalog
A hardcoded list, however carefully copied from today's catalog, is the
next drift. The picker now asks connectors.list through the same
session-owned RPC the connector cards use and offers exactly what the
gateway carries: a curated lead order puts the everyday apps first, chat
channels stay on Messaging, everything else is reachable by search. The
picks are gateway slugs, handed straight to manage_connections. No
catalog (toolset off, gateway unreachable) ends the step honestly with
Skip instead of inventing apps.
* test(desktop): the guided first launch never forces a sign-in
The acceptance criterion the guided onboarding was built to, as a test:
while the gate is cinematic, guided or handoff, the provider picker does
not open and a credential warning is dropped rather than deferred to the
next send. Outside the guide the picker opens as before. Red against the
tree before the guards landed (6 of 9).
* fix(desktop): a relaunch mid-guide resumes the guide, in the guide's shape
Closing the app during the guided first launch and reopening it booted the
normal shell around the persisted solo layout: the connecting splash, the
stock composer and model picker, a small window whose sidebars would not
open, while the gate still read guided. The gate now queues a kickoff for
the guided phase too (the kickoff adopts the existing guide chat by title),
takes the solo shape before the gateway opens rather than after, and the
connecting overlay yields to the guide's own opening. A typed reply in the
composer now closes an ask card and the first-build chips the same way a
click does; the layout card's Continue comes back Done.
* style(desktop): one answeredAfter helper for the ask card and first-build chips
* fix(desktop): the guide takes its shape on the tick the film ends, not after the window shows
Between the film and the greeting the full-size shell painted for a beat:
finishIntroReveal showed the main window, then the kickoff shrank it once
the setup profile answered. The listener on the intro's hidden edge now
takes the guide's shape (solo layout + small centred window) synchronously,
so the window is already the guide when it is shown. One takeGuideShape
owns the pair; kickoff and the boot gate call it idempotently.
* style(desktop): the 'nothing connects yet' line reads first on the connectors card
fetchJsonViaOauthSession and finalizeGatewayDownload still hand-rolled the
`Error("<status>: <body>")` + statusCode shape the helper was introduced to
consolidate, so the "shared by all three paths" claim was false on landing.
The download-transport source test now asserts the helper call.
Cold-launching against a gated remote gateway whose stored session has been
invalidated server-side alternated between the connecting state and the
recovery overlay; the Sign in button was only intermittently clickable.
Root cause: fetchJson() built a bare Error("401: ...") for HTTP failures on
the native-bearer path, dropping res.statusCode. Every downstream classifier
is shape-based (isGatewayAuthRejection, isServerSideHttpError, the
ensureNativeAccessToken 401 check), so the confirmed rejection looked like a
transport blip: withTransientRetries hammered it, gatewayTicketFailure used
the transport copy, startHermes tagged the boot retryable, and the renderer's
bounded boot-retry loop re-emitted running:true over the overlay on every
attempt. Separately, gatewayTicketFailure only ever set needsOauthLogin,
which isReauthRequiredError ignores, so even a structured 401 from the cookie
path never latched.
- api-transport: httpStatusError() is the one HTTP-error shape; fetchJson and
fetchPublicJson use it, matching fetchJsonViaOauthSession.
- mintGatewayWsTicket: the gateway never rotates a native bearer server-side
(dashboard_auth/middleware.py), so a bearer 401 gets ONE forced
/auth/native/refresh; a live refresh token retries the mint once, a dead one
drops the stored tokens and the rejection is confirmed.
- gatewayTicketFailure: a confirmed 401/403 is tagged isReauthRequired so
startHermes latches it and marks the boot non-retryable.
- startHermes: latches are set before the first await in the failure path;
updateBootProgress holds every update that is not a re-emit of the latched
failure until a recovery path clears the latch.
Tests: composition + main.ts source pins (remote-reauth-latch.test.ts), unit
coverage for the new helpers, and a Playwright e2e that boots the real app
against a fake gateway with a dead session and proves the overlay latches
once (one mint, one refresh, retryable:false, Sign in stays clickable).
Playwright 1.58 accepts a Promise-valued waitForFunction predicate before its false result. Await each zoom read explicitly; cover delayed responses and fresh-install startup.
Share zoom preparation across both launchers and stage the helper with each driver. Use the Appearance preference bridge and verify page zoom rather than display DPR.
Real Electron regressions fail with transient zoom on focus/navigation and pass with persistence. Repeated click-throughs, onboarding unit tests, E2E typecheck and lint passed. The historical onboarding timeout and full install/update matrix remain unverified.
One conflict: upstream 6e7c7c7da9 replaced bot-mode-closed-chat-stays-closed.spec.ts with bot-mode-row-click-mirrors-registry.spec.ts while our side had rewired its mock-server import. Kept upstream's replacement and rewired the three new specs importing ./mock-server to the consolidated tests-js copy (symbols verified present).
Follow-up on @fortun8te's user-made roster sections:
- Sections start empty: no seeded General/Workforce/Clients. With no
sections created the roster renders exactly as before.
- New section and Rename go through one Dialog + Input + Cancel/Save
(the app's session-rename shape) instead of an inline caret; the row
menu's "New section…" files the bot as it creates.
- Delete needs no confirmation: bots return to Unassigned and the toast
offers Undo (restores the section in its slot and refiles its bots).
- Drag: single-row drag under a private MIME type, every valid target
shows a faint outline while a drag is live, the hovered target lights
up, the source section refuses the drop, Escape cancels, and the moved
row no longer stays faded after it remounts under its new section.
- Multi-select (cmd/shift-click, querySelectorAll shift-range) dropped:
the roster has no selection model. Per-bot saveBotMeta writes run in
sequence, one per profile (membership IS a field on each profile).
- Section heading reuses RosterSectionHeader (gains `action` /
`onDoubleClick`), so user sections fold and look like the gateway
headings; ⋯ menu and right-click drive the same Rename / Move up /
Move down / Delete. Empty sections show a dashed "Drag bots here" slot.
- Composes with gateway buckets: sections nest INSIDE each connection
bucket, indented under a hairline rail (membership lives in the bot's
profile on that gateway); empty sections repeat there only mid-drag.
- Full i18n parity (en / ja / zh / zh-hant) for every new string; icon
toggle and the storage-async plumbing removed.
- Tests trimmed to the three invariants (membership persists through
saveBotMeta + reload, remainder = Unassigned, delete returns bots +
undo) plus a live Electron e2e covering the whole flow.
- Docs: "Organize bots into sections" in user-guide/bot-mode.md.
Every bot's canonical chat is stored under the same title ("Bot Chat" — the
name the gateway resolves it by, and an invariant roster-actions.ts's stale-tile
probe and #90102 rely on), so the main tab strip captioned every open bot chat
identically and two bots' tabs were indistinguishable (#99152).
Fix at the presentation layer, leaving the stored title and tabTitle untouched:
- workspace-scope.ts gains `$workspaceOwnerLabels` + `workspaceOwnerTitle()`:
a bots-mode tab whose resolved title still equals its registered placeholder
reads its owner's label instead. Side threads / Sessions tabs are untouched.
- session-tile.tsx captions tiles through it (and the drag payload); the main
`workspace` tab (controller.tsx) does the same via `$botChatScopes`, the
bot-mode scope the main tab was last opened under (it has no tile).
- The hermes-bots roster publishes displayName() per owner key through the new
`host.setWorkspaceOwnerLabel` (feature-detected), so renames follow.
Supersedes #99177, which set tabTitle at open time — that reverts after mount
because tileTitle() prefers the stored row's title once the hidden row is
upserted, and breaks the `workspaceTabTitle === 'Bot Chat'` invariant.
Tests: one unit test on workspaceOwnerTitle() (bot chat → bot name; side
thread / sessions tab / unlabeled owner untouched) and one Electron e2e
(tab strip reads "Alpha", not "Bot Chat"); both fail on main, pass here.
Closes#99152
Supersedes #99177
Co-authored-by: twotnguyen <nguyenngoctinh011258@gmail.com>
A plain roster click fronted whatever bots-workspace tab the user last had
active for that bot (#96649). A '+' side thread persists in Local Storage
across restarts, so it won every click forever while the row kept previewing
the canonical Bot Chat (profiles.list canonical_session) — sidebar and center
described two different conversations; a message typed there landed in the
side thread and the row never moved. Support thread "[Bots] - Sessions is not
in sync again" (bundle 7dfff039), reproduced live on origin/main.
- roster-actions: the open-tab shortcut may front only the canonical chat
(registry id or lineage tip, via a new onlyStoredIds allowlist on
focusWorkspaceOwnerSessionTile); anything else resolves the registry and
opens in place. Side tabs stay open beside it. "Open Bot Chat" in the row
menu is the same action; the `canonical` option goes away.
- roster-actions: when the FOCUSED Bot Chat's canonical session advances on
the gateway (cron bot-chat delivery, message_agent, group round, CLI turn —
none reach this window's stream), re-open it in place so the transcript
refreshes instead of waiting for an app restart (#99393 class).
Tests: the fronting-shortcut unit file and its e2e spec pinned the reversed
behavior; replaced by one unit file (5 tests) and one e2e spec that fails on
main and passes here. group-to-local-bot-handoff e2e still passes.
The scripted turns execute real terminal commands, and the sidebar
sentinel-wait loop trips the dangerous-command guard: the turn parks
behind a Run/Reject approval card, and the default 'smart' mode fires an
aux LLM approval call at the same mock provider — consuming a
scripted-turn index and never resolving. On the slower CI runner this
stalled the sidebar-dot family (sidebar-states 157/245, tile-unread 166)
until spec timeout; run 33543723331's error-context snapshots show the
approval card blocking each stalled turn. Locally the race usually won
the other way, which is why these passed on dev machines.
Fix: fixtures write 'approvals: mode: "off"' into the mock provider
config by default (specs supplying their own approvals: section own it),
mirroring the auto-title default. Also drop the DOT-DEBUG diagnostics
from tile-unread-bug now that the root cause is identified.
Local: sidebar-states + tile-unread + correction-session-switch all
green in seconds (3-9s vs 90s timeouts); full suite 62 passed /
11 skipped / 1 flaky-passed.
The Desktop E2E lane was disabled Aug 2 – Sep 1; the app and gateway kept
moving, so 16 specs rotted against current main. All failures traced to
spec/harness drift, not product regressions:
- fixtures.ts: title generation now rides the main model (#83636), firing a
background completion at the mock after every turn — it contains the whole
conversation (trigger keywords included), advancing scripted-turn indices
and tripping hold-for-prompt matchers. Disabled by default in the mock
provider config; specs supplying their own `auxiliary:` section own it.
- chat/interim-messages/session-compression/correction-session-switch/
hidden-history-messages: busy-state and transcript assertions updated to
the current composer aria-labels, interim-message semantics, and
verify-on-stop continuation behavior on main.
- bot-mode-closed-chat-stays-closed/group-to-local-bot-handoff: Bot Chat tab
selectors updated for the Bot Mode rework (tabs keyed by
connection+profile, renamed tab triggers).
- glyph-spinner: assertions made compositor-honest for the CI runner
(steps() keyframes + layer promotion probed via the animation registry
instead of GPU-dependent screenshots).
- sidebar-states/tile-unread-bug: event-driven waits with mock-server
release handles replace wall-clock polls that lost races on loaded
runners.
- warm-resume-jitter/image-attachment-resume: real-session-builder harness
waits for the thread viewport before evaluating; failure path now dumps
per-surface pane state.
Local full-suite run on the CI-equivalent xvfb setup: 62 passed,
11 skipped, 1 flaky-passed (correction-session-switch live-correction spec,
passes on retry). No product code changed.
Conflicts, three, resolved:
- scripts/desktop-update.ps1: upstream's side taken whole. Upstream moved
the hand-off to scripts/desktop-update/windows.ps1 (this file is now a
one-line compat forwarder) and the new implementation already drains
both pipes asynchronously with bounded abandonment, which supersedes
this branch's stderr-drain fix for the same deadlock.
- apps/desktop/e2e/fixtures.ts: kept upstream's resolveElectronBinary
import alongside this branch's consolidated mock-server path.
- tests-js/scripts/mock-server.ts: kept upstream's task-panel trigger
addition inside the consolidated file; rewired the five upstream specs
still importing './mock-server' to the consolidated path (export sets
verified identical) and dropped the superseded apps/desktop/e2e copy.
findElectron() probed exactly one path, and got three things wrong at
once for anyone not on a hoisted POSIX install:
* It looked only under the REPO ROOT. This is an npm workspaces repo and
npm hoists a dependency only when nothing conflicts, so `electron`
installing into apps/desktop/node_modules is an ordinary outcome, not a
broken tree.
* It joined a bare `electron`. On Windows the dist file is
`electron.exe`, so the probe could never match there.
* Its PATH fallback spawned `which`, which is not a command on Windows,
so the fallback failed for a reason unrelated to whether electron is on
PATH.
The three combine into a misleading error: the suite refuses to start
with 'Run "npm install" from the repo root' on a tree that has electron
installed. Reproduced on Windows 11 against this repo, where
apps/desktop/node_modules/electron/dist/electron.exe exists and the old
body throws that message; the reporter on #88036 hit the same thing on
Linux and had to hand-symlink the package before the suite would run.
Resolution now asks the installed `electron` package for its own path
first (its main export IS the absolute executable, resolved from
path.txt and honouring ELECTRON_OVERRIDE_DIST_PATH), then falls back to
explicit dist probes for each root, then to PATH with the platform's
lookup command. The error message lists what was searched.
The rules live in e2e/electron-binary.ts so they can be unit-tested
without importing the Playwright runner, with the platform passed in
rather than read from process.platform: reading it would leave every
Windows rule untested on the Linux CI runner.
Wiring: the vitest `electron` project picks up e2e/**/*.unit.test.ts and
Playwright ignores the same pattern, so helper unit tests run in exactly
one runner and the specs are untouched.
Verified: 5 unit tests pass; mutation-checked one rule at a time
(hardcoding the binary name fails 2, reversing the probe order fails 1,
hardcoding `which` fails 1). tsc -p . and tsc -p tsconfig.e2e.json
clean.
This is the environment blocker called out in #88036, not its rendering
bug, so it is deliberately a subset.
Refs #88036
The packaged app crashed at launch with 'No QueryClient set, use
QueryClientProvider to set one': useQuery in a lazy chunk (session-list-density)
read a second @tanstack/react-query runtime whose QueryClientContext was never
populated by the entry's QueryClientProvider. The source tree was correct — the
duplication happened at build time, because react-query was the one
context-bearing runtime not pinned to a shared vendor chunk, and rolldown's
merge heuristics inline the spare copy into a lazy chunk depending on toolchain
version.
- vite.config.ts: add @tanstack/react-query to the vendor-react
advancedChunks group + dev dedupe list, mirroring the react-router fix.
- assert-dist-built.mjs: fail the build when the 'No QueryClient set'
invariant appears in more than one JS asset (launch-smoke guard).
- assert-dist-built.test.mjs: unit tests for the new invariant check.
- launch-packaged-app.spec.ts: e2e smoke test asserting the packaged app
boots to real UI, not the QueryClient error boundary.
In Bot Mode every roster click resolved the bot's canonical "Bot Chat" by
name and opened it as a tab. Nothing records a tab close (the plugin keeps no
closed set; core's tile bucket only forgets), so a Bot Chat the user had
closed came back beside every newer thread on every bot switch — close it,
start a new thread, visit another bot, come back: two tabs again, forever.
A row click is now "go to this bot": when the bot's workspace already holds
tabs, the one the user last had active is fronted and no chat is resolved or
opened. The canonical chat is opened only when the bot has nothing open, or
on the explicit asks — a new "Open Bot Chat" row-menu item and the Bots home
"Open chat" button (`openRosterBot(bot, { canonical: true })`).
- session-states: `focusWorkspaceOwnerSessionTile(ownerKey)` fronts the
owner's remembered-active tile (else its most recent) and reports it.
- sdk: `host.focusOpenWorkspaceSession(ownerKey)` exposes it to plugins;
feature-detected in the plugin so older shells keep the canonical open.
- hermes-bots: `focusExistingBotTab` short-circuits `openRosterBot`; the
claim it records carries only the fronted tab, and the session.reclaimed
re-resume now skips such claims so it cannot resurrect the closed chat.
Tests: vitest for the core helper, a node test for the click path (open tabs
win, nothing open → canonical, explicit canonical, older shell, throwing
host), and an e2e that seeds two bots with real "Bot Chat" rows, closes one,
starts a thread, switches bots and back, and asserts the Bot Chat stays
closed until asked for explicitly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
With several gateways registered, the Sessions profile rail only ever showed
the active gateway's profiles; reaching a bot on another machine meant a
gateway switch first, then a click on the rail that appeared afterwards. Bot
Mode (#91134) and Capabilities already read the union agent roster; the rail
is now its third consumer.
- Every registered gateway's profiles sit on the one strip, in registry order
(This device first, then by label), each group headed by that gateway's
kind glyph. The active gateway's squares are unchanged; the others are
"at rest" (dimmed) with tooltips/accessible names qualified by machine
(`inbox · Homelab`), so same-named profiles never read alike.
- Clicking an at-rest square performs the same dial → commit → re-home as
the statusbar switcher, landing on that exact (gateway, profile):
`selectConnection(id, { profile })`. The spinner sits on the clicked
square; the previous source stays painted until the target answers.
Groups keep their slots whichever gateway is active, so a square never
moves under the pointer that clicked it.
- Right-click on an at-rest square: Switch to / Color / Rename / Edit
SOUL.md / Delete, executed on the owning gateway (renameProfile,
getProfileSoul and updateProfileSoul accept the same scope deleteProfile
already had); the delete confirmation names the machine. The legacy
per-profile "Connect to a remote host…" item is hidden on multi-gateway
setups, where the rail shows machines directly.
- Unreachable gateways keep their squares with an amber dot on the glyph;
two registrations of one backend collapse to one group; past thirteen
squares across the fleet the strip condenses into a menu sectioned by
gateway. Roster is fetched on mount / focus / registry change only — no
periodic fleet polling.
- Single-gateway Desktops render exactly as before: no roster fetch, same DOM.
Also fixes a boot race the e2e surfaced: initializeConnectionsRegistry()
"restored" the launch-mode source over a switch the user had already made
while boot was settling (same class as #91047). The restore now yields when
a switch is pending or already landed.
Tests: pure grouping (fleet-rail.test.ts), rail component fleet mode
(profile-rail-fleet.test.tsx), store (explicit profile pick; restore yields),
and a Playwright e2e (fleet-profile-rail.spec.ts) that boots Desktop with two
REAL backends — the local one plus a second `hermes serve` registered as a
remote URL connection — and verifies layout, a real re-home, gateway-scoped
actions, and order stability.
Docs: multi-connection-desktop.md describes the fleet rail.
Refs #89304, #92384, #91047, #94724
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Electron safeStorage parks a per-app key ('Hermes Key') in the macOS login
keychain; on machines with a locked/missing/corrupted default keychain that
turned every Hermes Desktop launch into a blocking 'Keychain Not Found' /
password dialog. Keychain-backed encryption is now an explicit opt-in:
- electron/secret-storage-policy.ts: standalone policy seam (default OFF,
strict === true coercion, one-shot migration flag) + unit tests
- default path never calls any safeStorage API (including
isEncryptionAvailable, which itself touches the keychain)
- one-shot legacy migration decrypts existing safeStorage blobs to plain
0600 files at first launch; undecryptable blobs are kept but read as
absent afterward (classify 'drop') so a dead keychain prompts at most once
- Settings -> Gateway toggle (all 5 locales) re-encodes every stored secret
store in place when flipped (v1 connection.json, v2 connections.json,
native-oauth-tokens.json)
- e2e: at-rest spec now covers both postures (opted-in unchanged contract,
default saves without secure storage, owner-only bits, restart round-trip)
- docs: multi-connection-desktop + desktop-native-signin updated
Drives the reported path rather than the helper: set a non-default
scale, then navigate to routes Chromium holds no zoom record for, which
is what opening a new session looks like to the per-URL store. Keeps the
Cmd/Ctrl+N case alongside it.
Co-authored-by: Clark Vines <38430798+clarkvines@users.noreply.github.com>
Review-round residuals: the spinner's user-select guard now beats
[data-selectable-text] regardless of stylesheet order; the will-change
layer hint clears under the global renderer pause and reduced motion so
parked spinners hold no compositor layer; the e2e travel assertion reads
the engine's keyframes (a computed transform always serializes to a
matrix, so the old '%' check could never fail); inline import() type
hoisted for the lint gate.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Replaces the deleted stylesheet-text assertions with tests that run the
thing they claim to cover.
e2e/glyph-spinner.spec.ts drives a real browser, where the CSS actually
executes: the strip's animation resolves to steps(N) for N frames, runs
infinitely, and travels a resolved length rather than a percentage (a
percentage translate is layout-dependent and Chromium refuses to
composite it). Both pause gates are covered — the per-spinner
`data-paused` attribute and the global renderer-pause attribute that
window blur / minimize / document-hidden arm — along with the layer
promotion being scoped to running spinners. A sampling test confirms the
transform visits a bounded number of distinct values across one cycle
(steps, not a linear sweep) and that nothing mutates the DOM while it
animates, which is the property the whole change exists to deliver.
status-invalidation-scope.test.tsx pins the scoping itself as a render
count. `useTapbackDoubleClick` is called by AssistantMessageBody and by
nothing else in the tree, which makes it an exact render counter for the
message root without exporting internals. A settle and a delta flush must
both leave that count untouched while the leaves update. Verified by
mutation: reinstating a root-level status subscription fails the settle
test (2 renders where 1 is required).
It also pins node identity across the settle transition, so the
inter-agent collapse cannot go back to swapping element types at the
message-root position and remounting the row under the scroll anchor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QwTc9XqUjhbay446VjugHZ
The same AST sweep over specs found fixtures maintained in parallel
across suites that have no reason to know about each other.
Twenty-one specs each mounted useMessageStream themselves and ten of the
harnesses were byte-identical; twenty now take renderMessageStream, with
overrides for the seams that genuinely vary. The SessionInfo builder was
spelled out field-by-field in seven specs, so a new backend field broke
seven files instead of one. Twenty specs carried their own inert
ResizeObserver and eleven repeated the animation-frame, CSS.escape,
scrollTo and WAAPI stubs the transcript needs to mount at all — split by
scope into src/test/jsdom for what any component might need and the
assistant-ui folder's own kit for the transcript. Plus the window-state
bridge, deferred, the external-store thread runtime, the manual
createRoot harness, and the per-folder caret, env-var, provider and
session fixtures.
Left alone on purpose: the store suites' makePrimary, where the vi.mock
harness around it is the actual duplication and cannot be hoisted out of
a hoisted factory; electron's deferred, where reaching into src/ from
the main process would invert the layering for eight lines; and the two
suites that compose another hook alongside the stream.
Deciding whether the OS can back glass needs os.release(), but every
Hermes window runs its preload with sandbox: true, where require is a
polyfill limited to electron, events, timers and url. The node:os import
threw before contextBridge ran, so window.hermesDesktop was never defined
and the app booted straight into "Desktop IPC bridge is unavailable".
Main already computes both verdicts, so preload asks for them over a
synchronous channel instead. No reply degrades to no glass, which is an
ordinary opaque window rather than a page thinned over nothing.
One renderer coordinator owns every right-click and replaces the
native Electron menus. Menus are assembled from what the click
landed on:
- Links and images get open/copy/save sections; chat links add the
reach-aware resolved-URL copy on remote gateways.
- Editables get spell-check suggestions (async-appended when
Chromium's facts arrive from main), cut/copy/paste, and select
all. Cut and copy need a selection; paste needs a non-empty
clipboard; select all needs field content. The verbs show their
accelerators instead of icons and dispatch a frame after the menu
closes, so the radix focus trap cannot steal the target. Select
all runs renderer-side, scoped to the field, because main's
selectAll acts on the focused frame and could grab the transcript.
- Terminals answer through registered xterm handles; the read-only
agent terminal hides paste.
- The in-app browser guest builds the same menus from the webview
tag's context-menu event: Chromium's editFlags gate the edit
verbs, spell-check rides the event, and Inspect element closes
every menu. Coordinates arrive as window-relative device pixels,
so the handler divides by the window zoom factor; guest edit
commands focus the webview first, because they act on the focused
webContents.
- Bare app chrome falls back to the window verbs.
Labels come from the locale files in all five languages. Main keeps
thin IPC verbs: edit commands, copy-image-at-gesture, spell-check
actions, and dictionary-add for guests (the tag has no session API).
The e2e spec exercises the real focus trap; it is blocked today by
the gateway-checking stall that also fails e2e/chat.spec.ts.
The mock server gains a batch clarify trigger that scripts a
two-question clarify turn. The scripted turn fires only while the
conversation has no tool result, so the answered batch falls through
to the canned reply instead of a repeat of the quiz.
The spec runs the real chain from composer to renderer and asserts
one batch card, the staged-answers confirm gate, and the settled
card. The local harness cannot boot the packaged app in this
environment (the pre-existing chat spec fails the same way), so the
proof for this spec is the CI run.
The batch card previously locked each answer with its own Continue
press. Now picks and typed answers stage locally, and one Confirm and
continue button (enabled when every question has an answer) submits
the whole batch. Staged answers stay editable until that confirm.
The wire protocol is unchanged. The confirm sends the per-question
locks in sequence, because the last lock resolves the blocked tool and
each earlier lock must already be accepted when it lands. Replayed
locked answers from a reconnect pre-stage their questions so restored
progress stays visible. The TUI and CLI keep incremental per-question
locks, so a timeout there still returns partial answers.
A batch clarify rendered as two identical interactive cards. The
tool.start row carries the model tool_call_id and the clarify.request
row carries a gateway request_id. The hydration-race merge correlates
the two rows with the top-level question text. A batch payload has no
top-level question, so the rows never matched and the card mounted
twice.
The correlation key for a batch now comes from the joined per-question
texts. The NUL separator cannot occur in real question text, so a batch
key cannot collide with a single-question key.
New coverage: two hydration-race tests for the batch shape (tool.start
first and clarify.request first), a mock-server batch clarify trigger,
and an E2E spec that runs the full chain and asserts exactly one card,
the per-question locks, the Confirm and continue relabel, and the
settled card.
The green "finished — unread" session dot lived only in the transient
$unreadFinishedSessionIds atom, written by a live busy->idle edge the
renderer had to witness. Closing and reopening the app grayed out every
dot, and a session that finished while the app was closed could never
be flagged at all.
Add a persisted layer (session-unread.ts), ported from the webui's
proven design:
- Seen watermarks (hermes.desktop.sessionSeenCounts): the message_count
last acknowledged per session, keyed by the durable lineage id (same
rule as session colors). A row whose live count exceeds its watermark
paints unread on every list refresh - this reconstructs dots after a
restart AND surfaces sessions that finished while the app was closed.
First sight of an unknown session seeds the watermark so a fresh
install doesn't light up every row.
- Explicit finish markers (hermes.desktop.unreadFinishedSessions): the
live edge, persisted, covering the gap before the sidebar list
refreshes its counts.
Opening a session acks both; the selected session's watermark tracks
its live count so on-screen activity never reads as unread. Chat and
cron rows get full watermark treatment; messaging rows keep explicit
markers only, so inbound messages don't paint false completion dots.
Profile switches keep persisted markers (keyed by durable id) and only
wipe the transient paint layer, so a round-trip repaints them.
Covered by store unit tests and an e2e spec that boots the app three
times: dot appears on a background finish, survives a restart, clears
on open, and stays cleared after another restart.
The helpers were tested; nothing proved main.ts called them. Reverting both
call sites and both imports in readDesktopConnectionConfig /
writeDesktopConnectionConfig left the whole suite green (947 passed / 2
skipped, tsc 0, eslint clean, e2e 1 passed 1 skipped) while connection.json
went back to 0644 — the user-visible fix this PR promises was untested.
The e2e spec could not catch it by construction: it asserts the ENCRYPTION
contract with a raw-bytes scan, and safeStorage keeps the token opaque
regardless of the file's mode, so a 0644 file passes that scan every time.
There was no mode assertion anywhere in e2e/.
Adds the missing third contract — unreadable by other local accounts — on all
three paths that can produce the file:
- write: assert the mode of the artifact test 1 already proves the app wrote.
- read, valid file: seed the app's own encrypted connection.json back to 0644
and assert launch tightens it. Scoped to the MODE only, so it is independent
of the still-deferred plaintext migration — the fixture's token is already
ciphertext, so nothing re-encrypts, no #62319 opt-in marker is involved, and
no rotation guidance is owed.
- read, corrupt file: a truncated file still holds the token bytes and throws
into the swallowing catch, so it would be the one file never tightened. This
is the only test that distinguishes the chmod's placement relative to the
parse.
Also moves the tighten above JSON.parse for exactly that reason, and pins the
cache invariant the placement depends on: the tighten must be a chmod, not a
rewrite, because it sits inside the function whose cache keys on mtimeMs.
Asserted as `mode & 0o077 === 0` rather than `=== 0o600` to avoid a
change-detector, and skipped on win32, where chmod maps to the read-only bit
and the fix deliberately no-ops (ACLs are PR #77527).
Every assertion was mutation-tested: reverting the full wiring fails all three;
reverting only the write path fails only the write test; deleting only the
tighten-on-read fails only the two read tests; moving the tighten below the
parse fails only the corrupt test; making the tighten a rewrite instead of a
chmod fails the mtime assertions. Bundle greps confirmed each mutation reached
dist/electron-main.mjs before the run.
(cherry picked from commit 99cfc16e7cdb759b674d890563f6a82113326547)
`connection.json` under the desktop app's Electron `userData` was written with no
file mode, so it landed at the `0644` umask default — while its two
credential-bearing neighbours in the same directory, `desktop-installation.json`
and `native-oauth-tokens.json`, were already `0600`. That file holds the
safeStorage-encrypted gateway token plus the fields that are NOT encrypted: the
gateway URL and the SSH host, user, and key path.
- Route the single write choke point through a helper that creates the file
owner-only and atomically.
- Tighten an already-existing `0644` file once per launch on the read path, so
installs that already have one do not stay world-readable until the next save.
- Refuse to act on a path that is a symlink or not owned by the current user,
matching the guards `desktop-installation.ts` already applies to its sibling.
The symlink guard alone turned out to be insufficient, and that is worth
recording: `writeSecretFileAtomic` tightens its *temp* path, so a symlink planted
at `connection.json.tmp` meant `writeFileSync` followed it, the guard correctly
bailed, and `renameSync` then moved the link onto `connection.json` permanently.
Measured, guard-only vs. as-landed:
guards only token leaked: true config is a symlink: true 755
guards + temp unlink token leaked: false config is a symlink: false 600
So the temp path is unlinked before the write.
Issue #77486's headline claim — that a dashboard session token is persisted in
plaintext — does not hold against main. The token has been safeStorage-encrypted
since the desktop app reached mainline in 51c68d4ab, and `encryptDesktopSecret`
aborts with an actionable message rather than degrading to plaintext when
safeStorage is unavailable. The `{ encoding: 'plain', value }` literal does exist
at main.ts:7084, but only on the `persistToken: false` branch, whose sole caller
is the connection-test handler, which never writes. So no mainline path *writes*
a plaintext token. The commits that did contain a plaintext-writing fallback
(d3d177283, d208f2c2c) are not ancestors of main — they live only on
upstream/bb/gui-* and the desktop-pr20059-installers pre-release tag.
At-rest migration of legacy non-safeStorage payloads is deliberately NOT included.
An earlier revision of this branch implemented it and it was removed after review
reproduced two token-loss paths: it force-converts the opt-in plaintext choice
PR #62319 adds (silently reverting the user's decision, then destroying the token
on the next launch without the `--password-store=basic` flag), and it converts a
portable credential into a keychain-bound one with no consent — destroying the
only recoverable copy while not remediating the real exposure, since every
existing backup still holds the plaintext and the true remedy is rotation. It also
persisted raw `parsed`, bypassing `sanitizeConnectionProfiles`. A comment at the
read path records the three preconditions any future attempt needs.
`decryptDesktopSecret`'s non-safeStorage read fallback is untouched — it is what
lets a pre-release or hand-edited config work at all.
Windows still inherits the userData directory ACL rather than an explicit
owner-only one; mode bits are advisory there, so that half is deferred to
PR #77527 rather than growing a second ACL implementation here.
e2e: `at-rest-connection-token.spec.ts` asserts the at-rest contract
implementation-independently — the token's plaintext value (and its base64 form)
must not appear in a raw-bytes scan of any file under userData or HERMES_HOME,
AND the app must still put the exact original token on the wire after a restart,
so a fix that simply drops the token cannot pass. Proven non-vacuous by mutation:
writing `{ encoding: 'plain', value }` still fails the scan while the
file-exists and gateway-URL guards pass. The migration case is a documented
`test.fixme` naming its three blockers.
Electron project 928 -> 924 tests (-9 migration, +5 new guard and
mechanism-isolation). Two of those five exist because reverting either owner-only
mechanism alone initially scored zero failures — they were masking each other, so
either could have been deleted green.
(cherry picked from commit 6e01add6578f08f015a567d3a7a7378f2ec3e768)
The dev:mock script duplicated the e2e mock server. The copy had only
the plain chat reply; every scripted path lived only in the e2e version.
A single mock server now lives in tests-js/scripts/mock-server.ts.
The e2e suite imports it as a library. Running the file directly
starts the server, writes a mock config, and launches the desktop app.
The dev:mock script now runs that file.
The e2e tsconfig lists tests-js/scripts in its include, because the
composite project rule requires every imported file to be listed.
Extend the packaged-app HUD geometry test from horizontal-only to full
containment: both axes for the dock and the input, plus an explicit
assertion that no percentage translate survives on the composer dock.
The vertical clipping reported on Windows (#82203) and macOS (#82214)
is the same escape class on the other axis, and the computed-translate
probe makes a future optimizer regression fail with a diagnosis instead
of a bare coordinate mismatch.
Session, instance, HUD, quick-entry and pet-overlay windows all open with
show: false and are revealed only by ready-to-show, so the Electron 40 bug
strands them exactly the way it stranded the primary window — and none of
them have the second-launch workaround that made the main-window case
recoverable.
Generalize the controller to any window and wire all six through one
wireWindowReveal helper. Callers pass their own reveal action (showInactive
for the pet overlay, show + focus for the HUD and quick entry) and their own
post-visible work, so whichever path wins runs them exactly once.
Quick entry now reveals the window the call created rather than whatever
`quickEntryWindow` points at when the event lands.
Every CodingStatusRow mounted its own WorktreeDialog and subscribed to the
same global `$newWorktreeRequest` token, so a single ⌘⇧B with two composers on
screen opened two stacked dialogs — dismissing the front one revealed an
identical empty dialog behind it, which read as the dialog "staying open" after
creating a worktree.
Mount it exactly once in the sidebar (beside ProjectDialog) and drive it from a
`$worktreeDialog` atom, mirroring how the project dialog already works. One
mount cannot double-open. Every entry point (⌘⇧B, the rail's kebab, the
sidebar's + button) now publishes intent instead of rendering its own copy; the
rail and the button pin their own repo so a tile's kebab still targets that
tile's worktree.
The target is resolved at open time by `resolveWorktreeRepoPath`, which walks
the focused surface's cwd then the entered project's root, validating each
candidate against the repo-status probe cache — a project's root folder is not
necessarily a git repo, so existence alone isn't proof. That makes the resolver
the sole authority, so the hotkey no longer pre-gates on `$repoStatus` and now
works from a detached session that sits inside a project. When nothing in reach
is a repo it is a silent no-op: a worktree only exists inside a repo, so there
is nothing to report.
Also adds a project picker to the dialog so the repo can be retargeted before
naming the branch.
E2E: extends worktree-branch-status.spec.ts with a 10-branch repo, visual
snapshots of the base-branch picker and the convert-branch view, a geometry
assertion that the picker isn't clipped by the dialog (fails headlessly on
regression rather than waiting for a human to compare diff images), and a
two-composer test asserting one keypress opens exactly one dialog. Tests 1 and
4 fail against the previous code and pass now.
Fresh installs now default to ~91%, but Playwright hit-testing and
visual baselines still assume 100%. Seed zoom-state.json so isolated
E2E profiles don't inherit the product default.