Follow-up on @fortun8te's user-made roster sections:
- Sections start empty: no seeded General/Workforce/Clients. With no
sections created the roster renders exactly as before.
- New section and Rename go through one Dialog + Input + Cancel/Save
(the app's session-rename shape) instead of an inline caret; the row
menu's "New section…" files the bot as it creates.
- Delete needs no confirmation: bots return to Unassigned and the toast
offers Undo (restores the section in its slot and refiles its bots).
- Drag: single-row drag under a private MIME type, every valid target
shows a faint outline while a drag is live, the hovered target lights
up, the source section refuses the drop, Escape cancels, and the moved
row no longer stays faded after it remounts under its new section.
- Multi-select (cmd/shift-click, querySelectorAll shift-range) dropped:
the roster has no selection model. Per-bot saveBotMeta writes run in
sequence, one per profile (membership IS a field on each profile).
- Section heading reuses RosterSectionHeader (gains `action` /
`onDoubleClick`), so user sections fold and look like the gateway
headings; ⋯ menu and right-click drive the same Rename / Move up /
Move down / Delete. Empty sections show a dashed "Drag bots here" slot.
- Composes with gateway buckets: sections nest INSIDE each connection
bucket, indented under a hairline rail (membership lives in the bot's
profile on that gateway); empty sections repeat there only mid-drag.
- Full i18n parity (en / ja / zh / zh-hant) for every new string; icon
toggle and the storage-async plumbing removed.
- Tests trimmed to the three invariants (membership persists through
saveBotMeta + reload, remainder = Unassigned, delete returns bots +
undo) plus a live Electron e2e covering the whole flow.
- Docs: "Organize bots into sections" in user-guide/bot-mode.md.
Every bot's canonical chat is stored under the same title ("Bot Chat" — the
name the gateway resolves it by, and an invariant roster-actions.ts's stale-tile
probe and #90102 rely on), so the main tab strip captioned every open bot chat
identically and two bots' tabs were indistinguishable (#99152).
Fix at the presentation layer, leaving the stored title and tabTitle untouched:
- workspace-scope.ts gains `$workspaceOwnerLabels` + `workspaceOwnerTitle()`:
a bots-mode tab whose resolved title still equals its registered placeholder
reads its owner's label instead. Side threads / Sessions tabs are untouched.
- session-tile.tsx captions tiles through it (and the drag payload); the main
`workspace` tab (controller.tsx) does the same via `$botChatScopes`, the
bot-mode scope the main tab was last opened under (it has no tile).
- The hermes-bots roster publishes displayName() per owner key through the new
`host.setWorkspaceOwnerLabel` (feature-detected), so renames follow.
Supersedes #99177, which set tabTitle at open time — that reverts after mount
because tileTitle() prefers the stored row's title once the hidden row is
upserted, and breaks the `workspaceTabTitle === 'Bot Chat'` invariant.
Tests: one unit test on workspaceOwnerTitle() (bot chat → bot name; side
thread / sessions tab / unlabeled owner untouched) and one Electron e2e
(tab strip reads "Alpha", not "Bot Chat"); both fail on main, pass here.
Closes#99152
Supersedes #99177
Co-authored-by: twotnguyen <nguyenngoctinh011258@gmail.com>
A plain roster click fronted whatever bots-workspace tab the user last had
active for that bot (#96649). A '+' side thread persists in Local Storage
across restarts, so it won every click forever while the row kept previewing
the canonical Bot Chat (profiles.list canonical_session) — sidebar and center
described two different conversations; a message typed there landed in the
side thread and the row never moved. Support thread "[Bots] - Sessions is not
in sync again" (bundle 7dfff039), reproduced live on origin/main.
- roster-actions: the open-tab shortcut may front only the canonical chat
(registry id or lineage tip, via a new onlyStoredIds allowlist on
focusWorkspaceOwnerSessionTile); anything else resolves the registry and
opens in place. Side tabs stay open beside it. "Open Bot Chat" in the row
menu is the same action; the `canonical` option goes away.
- roster-actions: when the FOCUSED Bot Chat's canonical session advances on
the gateway (cron bot-chat delivery, message_agent, group round, CLI turn —
none reach this window's stream), re-open it in place so the transcript
refreshes instead of waiting for an app restart (#99393 class).
Tests: the fronting-shortcut unit file and its e2e spec pinned the reversed
behavior; replaced by one unit file (5 tests) and one e2e spec that fails on
main and passes here. group-to-local-bot-handoff e2e still passes.
The scripted turns execute real terminal commands, and the sidebar
sentinel-wait loop trips the dangerous-command guard: the turn parks
behind a Run/Reject approval card, and the default 'smart' mode fires an
aux LLM approval call at the same mock provider — consuming a
scripted-turn index and never resolving. On the slower CI runner this
stalled the sidebar-dot family (sidebar-states 157/245, tile-unread 166)
until spec timeout; run 33543723331's error-context snapshots show the
approval card blocking each stalled turn. Locally the race usually won
the other way, which is why these passed on dev machines.
Fix: fixtures write 'approvals: mode: "off"' into the mock provider
config by default (specs supplying their own approvals: section own it),
mirroring the auto-title default. Also drop the DOT-DEBUG diagnostics
from tile-unread-bug now that the root cause is identified.
Local: sidebar-states + tile-unread + correction-session-switch all
green in seconds (3-9s vs 90s timeouts); full suite 62 passed /
11 skipped / 1 flaky-passed.
The Desktop E2E lane was disabled Aug 2 – Sep 1; the app and gateway kept
moving, so 16 specs rotted against current main. All failures traced to
spec/harness drift, not product regressions:
- fixtures.ts: title generation now rides the main model (#83636), firing a
background completion at the mock after every turn — it contains the whole
conversation (trigger keywords included), advancing scripted-turn indices
and tripping hold-for-prompt matchers. Disabled by default in the mock
provider config; specs supplying their own `auxiliary:` section own it.
- chat/interim-messages/session-compression/correction-session-switch/
hidden-history-messages: busy-state and transcript assertions updated to
the current composer aria-labels, interim-message semantics, and
verify-on-stop continuation behavior on main.
- bot-mode-closed-chat-stays-closed/group-to-local-bot-handoff: Bot Chat tab
selectors updated for the Bot Mode rework (tabs keyed by
connection+profile, renamed tab triggers).
- glyph-spinner: assertions made compositor-honest for the CI runner
(steps() keyframes + layer promotion probed via the animation registry
instead of GPU-dependent screenshots).
- sidebar-states/tile-unread-bug: event-driven waits with mock-server
release handles replace wall-clock polls that lost races on loaded
runners.
- warm-resume-jitter/image-attachment-resume: real-session-builder harness
waits for the thread viewport before evaluating; failure path now dumps
per-surface pane state.
Local full-suite run on the CI-equivalent xvfb setup: 62 passed,
11 skipped, 1 flaky-passed (correction-session-switch live-correction spec,
passes on retry). No product code changed.
findElectron() probed exactly one path, and got three things wrong at
once for anyone not on a hoisted POSIX install:
* It looked only under the REPO ROOT. This is an npm workspaces repo and
npm hoists a dependency only when nothing conflicts, so `electron`
installing into apps/desktop/node_modules is an ordinary outcome, not a
broken tree.
* It joined a bare `electron`. On Windows the dist file is
`electron.exe`, so the probe could never match there.
* Its PATH fallback spawned `which`, which is not a command on Windows,
so the fallback failed for a reason unrelated to whether electron is on
PATH.
The three combine into a misleading error: the suite refuses to start
with 'Run "npm install" from the repo root' on a tree that has electron
installed. Reproduced on Windows 11 against this repo, where
apps/desktop/node_modules/electron/dist/electron.exe exists and the old
body throws that message; the reporter on #88036 hit the same thing on
Linux and had to hand-symlink the package before the suite would run.
Resolution now asks the installed `electron` package for its own path
first (its main export IS the absolute executable, resolved from
path.txt and honouring ELECTRON_OVERRIDE_DIST_PATH), then falls back to
explicit dist probes for each root, then to PATH with the platform's
lookup command. The error message lists what was searched.
The rules live in e2e/electron-binary.ts so they can be unit-tested
without importing the Playwright runner, with the platform passed in
rather than read from process.platform: reading it would leave every
Windows rule untested on the Linux CI runner.
Wiring: the vitest `electron` project picks up e2e/**/*.unit.test.ts and
Playwright ignores the same pattern, so helper unit tests run in exactly
one runner and the specs are untouched.
Verified: 5 unit tests pass; mutation-checked one rule at a time
(hardcoding the binary name fails 2, reversing the probe order fails 1,
hardcoding `which` fails 1). tsc -p . and tsc -p tsconfig.e2e.json
clean.
This is the environment blocker called out in #88036, not its rendering
bug, so it is deliberately a subset.
Refs #88036
The packaged app crashed at launch with 'No QueryClient set, use
QueryClientProvider to set one': useQuery in a lazy chunk (session-list-density)
read a second @tanstack/react-query runtime whose QueryClientContext was never
populated by the entry's QueryClientProvider. The source tree was correct — the
duplication happened at build time, because react-query was the one
context-bearing runtime not pinned to a shared vendor chunk, and rolldown's
merge heuristics inline the spare copy into a lazy chunk depending on toolchain
version.
- vite.config.ts: add @tanstack/react-query to the vendor-react
advancedChunks group + dev dedupe list, mirroring the react-router fix.
- assert-dist-built.mjs: fail the build when the 'No QueryClient set'
invariant appears in more than one JS asset (launch-smoke guard).
- assert-dist-built.test.mjs: unit tests for the new invariant check.
- launch-packaged-app.spec.ts: e2e smoke test asserting the packaged app
boots to real UI, not the QueryClient error boundary.
In Bot Mode every roster click resolved the bot's canonical "Bot Chat" by
name and opened it as a tab. Nothing records a tab close (the plugin keeps no
closed set; core's tile bucket only forgets), so a Bot Chat the user had
closed came back beside every newer thread on every bot switch — close it,
start a new thread, visit another bot, come back: two tabs again, forever.
A row click is now "go to this bot": when the bot's workspace already holds
tabs, the one the user last had active is fronted and no chat is resolved or
opened. The canonical chat is opened only when the bot has nothing open, or
on the explicit asks — a new "Open Bot Chat" row-menu item and the Bots home
"Open chat" button (`openRosterBot(bot, { canonical: true })`).
- session-states: `focusWorkspaceOwnerSessionTile(ownerKey)` fronts the
owner's remembered-active tile (else its most recent) and reports it.
- sdk: `host.focusOpenWorkspaceSession(ownerKey)` exposes it to plugins;
feature-detected in the plugin so older shells keep the canonical open.
- hermes-bots: `focusExistingBotTab` short-circuits `openRosterBot`; the
claim it records carries only the fronted tab, and the session.reclaimed
re-resume now skips such claims so it cannot resurrect the closed chat.
Tests: vitest for the core helper, a node test for the click path (open tabs
win, nothing open → canonical, explicit canonical, older shell, throwing
host), and an e2e that seeds two bots with real "Bot Chat" rows, closes one,
starts a thread, switches bots and back, and asserts the Bot Chat stays
closed until asked for explicitly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
With several gateways registered, the Sessions profile rail only ever showed
the active gateway's profiles; reaching a bot on another machine meant a
gateway switch first, then a click on the rail that appeared afterwards. Bot
Mode (#91134) and Capabilities already read the union agent roster; the rail
is now its third consumer.
- Every registered gateway's profiles sit on the one strip, in registry order
(This device first, then by label), each group headed by that gateway's
kind glyph. The active gateway's squares are unchanged; the others are
"at rest" (dimmed) with tooltips/accessible names qualified by machine
(`inbox · Homelab`), so same-named profiles never read alike.
- Clicking an at-rest square performs the same dial → commit → re-home as
the statusbar switcher, landing on that exact (gateway, profile):
`selectConnection(id, { profile })`. The spinner sits on the clicked
square; the previous source stays painted until the target answers.
Groups keep their slots whichever gateway is active, so a square never
moves under the pointer that clicked it.
- Right-click on an at-rest square: Switch to / Color / Rename / Edit
SOUL.md / Delete, executed on the owning gateway (renameProfile,
getProfileSoul and updateProfileSoul accept the same scope deleteProfile
already had); the delete confirmation names the machine. The legacy
per-profile "Connect to a remote host…" item is hidden on multi-gateway
setups, where the rail shows machines directly.
- Unreachable gateways keep their squares with an amber dot on the glyph;
two registrations of one backend collapse to one group; past thirteen
squares across the fleet the strip condenses into a menu sectioned by
gateway. Roster is fetched on mount / focus / registry change only — no
periodic fleet polling.
- Single-gateway Desktops render exactly as before: no roster fetch, same DOM.
Also fixes a boot race the e2e surfaced: initializeConnectionsRegistry()
"restored" the launch-mode source over a switch the user had already made
while boot was settling (same class as #91047). The restore now yields when
a switch is pending or already landed.
Tests: pure grouping (fleet-rail.test.ts), rail component fleet mode
(profile-rail-fleet.test.tsx), store (explicit profile pick; restore yields),
and a Playwright e2e (fleet-profile-rail.spec.ts) that boots Desktop with two
REAL backends — the local one plus a second `hermes serve` registered as a
remote URL connection — and verifies layout, a real re-home, gateway-scoped
actions, and order stability.
Docs: multi-connection-desktop.md describes the fleet rail.
Refs #89304, #92384, #91047, #94724
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Electron safeStorage parks a per-app key ('Hermes Key') in the macOS login
keychain; on machines with a locked/missing/corrupted default keychain that
turned every Hermes Desktop launch into a blocking 'Keychain Not Found' /
password dialog. Keychain-backed encryption is now an explicit opt-in:
- electron/secret-storage-policy.ts: standalone policy seam (default OFF,
strict === true coercion, one-shot migration flag) + unit tests
- default path never calls any safeStorage API (including
isEncryptionAvailable, which itself touches the keychain)
- one-shot legacy migration decrypts existing safeStorage blobs to plain
0600 files at first launch; undecryptable blobs are kept but read as
absent afterward (classify 'drop') so a dead keychain prompts at most once
- Settings -> Gateway toggle (all 5 locales) re-encodes every stored secret
store in place when flipped (v1 connection.json, v2 connections.json,
native-oauth-tokens.json)
- e2e: at-rest spec now covers both postures (opted-in unchanged contract,
default saves without secure storage, owner-only bits, restart round-trip)
- docs: multi-connection-desktop + desktop-native-signin updated
Drives the reported path rather than the helper: set a non-default
scale, then navigate to routes Chromium holds no zoom record for, which
is what opening a new session looks like to the per-URL store. Keeps the
Cmd/Ctrl+N case alongside it.
Co-authored-by: Clark Vines <38430798+clarkvines@users.noreply.github.com>
Review-round residuals: the spinner's user-select guard now beats
[data-selectable-text] regardless of stylesheet order; the will-change
layer hint clears under the global renderer pause and reduced motion so
parked spinners hold no compositor layer; the e2e travel assertion reads
the engine's keyframes (a computed transform always serializes to a
matrix, so the old '%' check could never fail); inline import() type
hoisted for the lint gate.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Replaces the deleted stylesheet-text assertions with tests that run the
thing they claim to cover.
e2e/glyph-spinner.spec.ts drives a real browser, where the CSS actually
executes: the strip's animation resolves to steps(N) for N frames, runs
infinitely, and travels a resolved length rather than a percentage (a
percentage translate is layout-dependent and Chromium refuses to
composite it). Both pause gates are covered — the per-spinner
`data-paused` attribute and the global renderer-pause attribute that
window blur / minimize / document-hidden arm — along with the layer
promotion being scoped to running spinners. A sampling test confirms the
transform visits a bounded number of distinct values across one cycle
(steps, not a linear sweep) and that nothing mutates the DOM while it
animates, which is the property the whole change exists to deliver.
status-invalidation-scope.test.tsx pins the scoping itself as a render
count. `useTapbackDoubleClick` is called by AssistantMessageBody and by
nothing else in the tree, which makes it an exact render counter for the
message root without exporting internals. A settle and a delta flush must
both leave that count untouched while the leaves update. Verified by
mutation: reinstating a root-level status subscription fails the settle
test (2 renders where 1 is required).
It also pins node identity across the settle transition, so the
inter-agent collapse cannot go back to swapping element types at the
message-root position and remounting the row under the scroll anchor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QwTc9XqUjhbay446VjugHZ
The same AST sweep over specs found fixtures maintained in parallel
across suites that have no reason to know about each other.
Twenty-one specs each mounted useMessageStream themselves and ten of the
harnesses were byte-identical; twenty now take renderMessageStream, with
overrides for the seams that genuinely vary. The SessionInfo builder was
spelled out field-by-field in seven specs, so a new backend field broke
seven files instead of one. Twenty specs carried their own inert
ResizeObserver and eleven repeated the animation-frame, CSS.escape,
scrollTo and WAAPI stubs the transcript needs to mount at all — split by
scope into src/test/jsdom for what any component might need and the
assistant-ui folder's own kit for the transcript. Plus the window-state
bridge, deferred, the external-store thread runtime, the manual
createRoot harness, and the per-folder caret, env-var, provider and
session fixtures.
Left alone on purpose: the store suites' makePrimary, where the vi.mock
harness around it is the actual duplication and cannot be hoisted out of
a hoisted factory; electron's deferred, where reaching into src/ from
the main process would invert the layering for eight lines; and the two
suites that compose another hook alongside the stream.
Deciding whether the OS can back glass needs os.release(), but every
Hermes window runs its preload with sandbox: true, where require is a
polyfill limited to electron, events, timers and url. The node:os import
threw before contextBridge ran, so window.hermesDesktop was never defined
and the app booted straight into "Desktop IPC bridge is unavailable".
Main already computes both verdicts, so preload asks for them over a
synchronous channel instead. No reply degrades to no glass, which is an
ordinary opaque window rather than a page thinned over nothing.
One renderer coordinator owns every right-click and replaces the
native Electron menus. Menus are assembled from what the click
landed on:
- Links and images get open/copy/save sections; chat links add the
reach-aware resolved-URL copy on remote gateways.
- Editables get spell-check suggestions (async-appended when
Chromium's facts arrive from main), cut/copy/paste, and select
all. Cut and copy need a selection; paste needs a non-empty
clipboard; select all needs field content. The verbs show their
accelerators instead of icons and dispatch a frame after the menu
closes, so the radix focus trap cannot steal the target. Select
all runs renderer-side, scoped to the field, because main's
selectAll acts on the focused frame and could grab the transcript.
- Terminals answer through registered xterm handles; the read-only
agent terminal hides paste.
- The in-app browser guest builds the same menus from the webview
tag's context-menu event: Chromium's editFlags gate the edit
verbs, spell-check rides the event, and Inspect element closes
every menu. Coordinates arrive as window-relative device pixels,
so the handler divides by the window zoom factor; guest edit
commands focus the webview first, because they act on the focused
webContents.
- Bare app chrome falls back to the window verbs.
Labels come from the locale files in all five languages. Main keeps
thin IPC verbs: edit commands, copy-image-at-gesture, spell-check
actions, and dictionary-add for guests (the tag has no session API).
The e2e spec exercises the real focus trap; it is blocked today by
the gateway-checking stall that also fails e2e/chat.spec.ts.
The mock server gains a batch clarify trigger that scripts a
two-question clarify turn. The scripted turn fires only while the
conversation has no tool result, so the answered batch falls through
to the canned reply instead of a repeat of the quiz.
The spec runs the real chain from composer to renderer and asserts
one batch card, the staged-answers confirm gate, and the settled
card. The local harness cannot boot the packaged app in this
environment (the pre-existing chat spec fails the same way), so the
proof for this spec is the CI run.
The batch card previously locked each answer with its own Continue
press. Now picks and typed answers stage locally, and one Confirm and
continue button (enabled when every question has an answer) submits
the whole batch. Staged answers stay editable until that confirm.
The wire protocol is unchanged. The confirm sends the per-question
locks in sequence, because the last lock resolves the blocked tool and
each earlier lock must already be accepted when it lands. Replayed
locked answers from a reconnect pre-stage their questions so restored
progress stays visible. The TUI and CLI keep incremental per-question
locks, so a timeout there still returns partial answers.
A batch clarify rendered as two identical interactive cards. The
tool.start row carries the model tool_call_id and the clarify.request
row carries a gateway request_id. The hydration-race merge correlates
the two rows with the top-level question text. A batch payload has no
top-level question, so the rows never matched and the card mounted
twice.
The correlation key for a batch now comes from the joined per-question
texts. The NUL separator cannot occur in real question text, so a batch
key cannot collide with a single-question key.
New coverage: two hydration-race tests for the batch shape (tool.start
first and clarify.request first), a mock-server batch clarify trigger,
and an E2E spec that runs the full chain and asserts exactly one card,
the per-question locks, the Confirm and continue relabel, and the
settled card.
The green "finished — unread" session dot lived only in the transient
$unreadFinishedSessionIds atom, written by a live busy->idle edge the
renderer had to witness. Closing and reopening the app grayed out every
dot, and a session that finished while the app was closed could never
be flagged at all.
Add a persisted layer (session-unread.ts), ported from the webui's
proven design:
- Seen watermarks (hermes.desktop.sessionSeenCounts): the message_count
last acknowledged per session, keyed by the durable lineage id (same
rule as session colors). A row whose live count exceeds its watermark
paints unread on every list refresh - this reconstructs dots after a
restart AND surfaces sessions that finished while the app was closed.
First sight of an unknown session seeds the watermark so a fresh
install doesn't light up every row.
- Explicit finish markers (hermes.desktop.unreadFinishedSessions): the
live edge, persisted, covering the gap before the sidebar list
refreshes its counts.
Opening a session acks both; the selected session's watermark tracks
its live count so on-screen activity never reads as unread. Chat and
cron rows get full watermark treatment; messaging rows keep explicit
markers only, so inbound messages don't paint false completion dots.
Profile switches keep persisted markers (keyed by durable id) and only
wipe the transient paint layer, so a round-trip repaints them.
Covered by store unit tests and an e2e spec that boots the app three
times: dot appears on a background finish, survives a restart, clears
on open, and stays cleared after another restart.
The helpers were tested; nothing proved main.ts called them. Reverting both
call sites and both imports in readDesktopConnectionConfig /
writeDesktopConnectionConfig left the whole suite green (947 passed / 2
skipped, tsc 0, eslint clean, e2e 1 passed 1 skipped) while connection.json
went back to 0644 — the user-visible fix this PR promises was untested.
The e2e spec could not catch it by construction: it asserts the ENCRYPTION
contract with a raw-bytes scan, and safeStorage keeps the token opaque
regardless of the file's mode, so a 0644 file passes that scan every time.
There was no mode assertion anywhere in e2e/.
Adds the missing third contract — unreadable by other local accounts — on all
three paths that can produce the file:
- write: assert the mode of the artifact test 1 already proves the app wrote.
- read, valid file: seed the app's own encrypted connection.json back to 0644
and assert launch tightens it. Scoped to the MODE only, so it is independent
of the still-deferred plaintext migration — the fixture's token is already
ciphertext, so nothing re-encrypts, no #62319 opt-in marker is involved, and
no rotation guidance is owed.
- read, corrupt file: a truncated file still holds the token bytes and throws
into the swallowing catch, so it would be the one file never tightened. This
is the only test that distinguishes the chmod's placement relative to the
parse.
Also moves the tighten above JSON.parse for exactly that reason, and pins the
cache invariant the placement depends on: the tighten must be a chmod, not a
rewrite, because it sits inside the function whose cache keys on mtimeMs.
Asserted as `mode & 0o077 === 0` rather than `=== 0o600` to avoid a
change-detector, and skipped on win32, where chmod maps to the read-only bit
and the fix deliberately no-ops (ACLs are PR #77527).
Every assertion was mutation-tested: reverting the full wiring fails all three;
reverting only the write path fails only the write test; deleting only the
tighten-on-read fails only the two read tests; moving the tighten below the
parse fails only the corrupt test; making the tighten a rewrite instead of a
chmod fails the mtime assertions. Bundle greps confirmed each mutation reached
dist/electron-main.mjs before the run.
(cherry picked from commit 99cfc16e7cdb759b674d890563f6a82113326547)
`connection.json` under the desktop app's Electron `userData` was written with no
file mode, so it landed at the `0644` umask default — while its two
credential-bearing neighbours in the same directory, `desktop-installation.json`
and `native-oauth-tokens.json`, were already `0600`. That file holds the
safeStorage-encrypted gateway token plus the fields that are NOT encrypted: the
gateway URL and the SSH host, user, and key path.
- Route the single write choke point through a helper that creates the file
owner-only and atomically.
- Tighten an already-existing `0644` file once per launch on the read path, so
installs that already have one do not stay world-readable until the next save.
- Refuse to act on a path that is a symlink or not owned by the current user,
matching the guards `desktop-installation.ts` already applies to its sibling.
The symlink guard alone turned out to be insufficient, and that is worth
recording: `writeSecretFileAtomic` tightens its *temp* path, so a symlink planted
at `connection.json.tmp` meant `writeFileSync` followed it, the guard correctly
bailed, and `renameSync` then moved the link onto `connection.json` permanently.
Measured, guard-only vs. as-landed:
guards only token leaked: true config is a symlink: true 755
guards + temp unlink token leaked: false config is a symlink: false 600
So the temp path is unlinked before the write.
Issue #77486's headline claim — that a dashboard session token is persisted in
plaintext — does not hold against main. The token has been safeStorage-encrypted
since the desktop app reached mainline in 51c68d4ab, and `encryptDesktopSecret`
aborts with an actionable message rather than degrading to plaintext when
safeStorage is unavailable. The `{ encoding: 'plain', value }` literal does exist
at main.ts:7084, but only on the `persistToken: false` branch, whose sole caller
is the connection-test handler, which never writes. So no mainline path *writes*
a plaintext token. The commits that did contain a plaintext-writing fallback
(d3d177283, d208f2c2c) are not ancestors of main — they live only on
upstream/bb/gui-* and the desktop-pr20059-installers pre-release tag.
At-rest migration of legacy non-safeStorage payloads is deliberately NOT included.
An earlier revision of this branch implemented it and it was removed after review
reproduced two token-loss paths: it force-converts the opt-in plaintext choice
PR #62319 adds (silently reverting the user's decision, then destroying the token
on the next launch without the `--password-store=basic` flag), and it converts a
portable credential into a keychain-bound one with no consent — destroying the
only recoverable copy while not remediating the real exposure, since every
existing backup still holds the plaintext and the true remedy is rotation. It also
persisted raw `parsed`, bypassing `sanitizeConnectionProfiles`. A comment at the
read path records the three preconditions any future attempt needs.
`decryptDesktopSecret`'s non-safeStorage read fallback is untouched — it is what
lets a pre-release or hand-edited config work at all.
Windows still inherits the userData directory ACL rather than an explicit
owner-only one; mode bits are advisory there, so that half is deferred to
PR #77527 rather than growing a second ACL implementation here.
e2e: `at-rest-connection-token.spec.ts` asserts the at-rest contract
implementation-independently — the token's plaintext value (and its base64 form)
must not appear in a raw-bytes scan of any file under userData or HERMES_HOME,
AND the app must still put the exact original token on the wire after a restart,
so a fix that simply drops the token cannot pass. Proven non-vacuous by mutation:
writing `{ encoding: 'plain', value }` still fails the scan while the
file-exists and gateway-URL guards pass. The migration case is a documented
`test.fixme` naming its three blockers.
Electron project 928 -> 924 tests (-9 migration, +5 new guard and
mechanism-isolation). Two of those five exist because reverting either owner-only
mechanism alone initially scored zero failures — they were masking each other, so
either could have been deleted green.
(cherry picked from commit 6e01add6578f08f015a567d3a7a7378f2ec3e768)
Extend the packaged-app HUD geometry test from horizontal-only to full
containment: both axes for the dock and the input, plus an explicit
assertion that no percentage translate survives on the composer dock.
The vertical clipping reported on Windows (#82203) and macOS (#82214)
is the same escape class on the other axis, and the computed-translate
probe makes a future optimizer regression fail with a diagnosis instead
of a bare coordinate mismatch.
Session, instance, HUD, quick-entry and pet-overlay windows all open with
show: false and are revealed only by ready-to-show, so the Electron 40 bug
strands them exactly the way it stranded the primary window — and none of
them have the second-launch workaround that made the main-window case
recoverable.
Generalize the controller to any window and wire all six through one
wireWindowReveal helper. Callers pass their own reveal action (showInactive
for the pet overlay, show + focus for the HUD and quick entry) and their own
post-visible work, so whichever path wins runs them exactly once.
Quick entry now reveals the window the call created rather than whatever
`quickEntryWindow` points at when the event lands.
Every CodingStatusRow mounted its own WorktreeDialog and subscribed to the
same global `$newWorktreeRequest` token, so a single ⌘⇧B with two composers on
screen opened two stacked dialogs — dismissing the front one revealed an
identical empty dialog behind it, which read as the dialog "staying open" after
creating a worktree.
Mount it exactly once in the sidebar (beside ProjectDialog) and drive it from a
`$worktreeDialog` atom, mirroring how the project dialog already works. One
mount cannot double-open. Every entry point (⌘⇧B, the rail's kebab, the
sidebar's + button) now publishes intent instead of rendering its own copy; the
rail and the button pin their own repo so a tile's kebab still targets that
tile's worktree.
The target is resolved at open time by `resolveWorktreeRepoPath`, which walks
the focused surface's cwd then the entered project's root, validating each
candidate against the repo-status probe cache — a project's root folder is not
necessarily a git repo, so existence alone isn't proof. That makes the resolver
the sole authority, so the hotkey no longer pre-gates on `$repoStatus` and now
works from a detached session that sits inside a project. When nothing in reach
is a repo it is a silent no-op: a worktree only exists inside a repo, so there
is nothing to report.
Also adds a project picker to the dialog so the repo can be retargeted before
naming the branch.
E2E: extends worktree-branch-status.spec.ts with a 10-branch repo, visual
snapshots of the base-branch picker and the convert-branch view, a geometry
assertion that the picker isn't clipped by the dialog (fails headlessly on
regression rather than waiting for a human to compare diff images), and a
two-composer test asserting one keypress opens exactly one dialog. Tests 1 and
4 fail against the previous code and pass now.
Fresh installs now default to ~91%, but Playwright hit-testing and
visual baselines still assume 100%. Seed zoom-state.json so isolated
E2E profiles don't inherit the product default.
Playwright closes the app with a turn still in flight, so the new quit
confirmation waited on a click nobody was there to make and the E2E
worker died on a 90s teardown timeout.
The large-session-resume E2E captured initialMockReplyCount immediately
after openSeededSession, which returns once the NEWEST turn is in the
viewport. With FIRST_PAINT_BUDGET=20 (lowered from 60 in this branch),
only the newest ~10 turns mount at first paint; the older turns
backfill in a rAF. The baseline was reading 10 instead of 27, so once
the backfill mounted the full 28 (27 seeded + 1 new), the test saw
"28 ≠ 11" and reported duplicates that were never there.
Wait for the oldest seeded turn to mount before taking the baseline.
This makes the count reflect the fully-mounted transcript regardless
of FIRST_PAINT_BUDGET, so the perf win (smaller first paint) and the
no-duplicate invariant both hold.
Refs #72504
The sidebar "+" now stacks a tab instead of replacing the surface, so the
prior session stays mounted and several chat surfaces can be on the page at
once. Helpers that waited for the old transcript to disappear from the page
timed out, and `.first()` locators / bare `document.querySelector` calls
started resolving against the wrong session (CI's "resolved to 2 elements"
strict-mode violation).
Target the most recently mounted `[data-composer-target]` surface instead,
and assert the NEW surface is empty rather than waiting for the old text to
vanish.
Sentinel-released processes exposed a second bare-sample assertion in the
same file. The unread-dot check ran a synchronous .count() 140ms after the
running dot cleared:
15.91s poll "dot should disappear" -> 2
16.05s ... -> 0 (running dot gone)
16.05s bare .count() for unread dot -> 0 FAIL
"Finished — unread" is an event-driven transition that lands just after the
running dot clears. The old fixed `sleep 5` happened to leave enough slack
between the two that a single sample usually caught it; releasing the
process deterministically removed that incidental slack and made the latent
race deterministic instead.
Poll for it, matching how sidebar-states.spec.ts already asserts this exact
dot. The split-tile assertion at line 234 stays a bare sample on purpose —
it asserts an absence (toBe(0)), where polling would only wait for something
that must never appear.
The cross-session sidebar specs asserted a state that could expire before
they looked at it, making them the flakiest tests in the suite — two reds
on unrelated PRs within three minutes on 2026-07-26.
Root cause, from the failing run's trace: the tests need a background
process that is still RUNNING after the agent turn finishes, but the
process was a fixed `sleep 5` racing two other clocks — the turn itself
(two model round trips plus a real subagent delegation) and the 4s
success linger before a finished task auto-dismisses. On a loaded runner
the "dot should appear" poll took 7.5s to see the dot; by then `sleep 5`
had already exited, `waitForFunction(finalText)` returned in 0.08s
because the turn was long done, and the next line — a bare synchronous
`.count()`, not a wait — sampled 0.
The process lifetime is now test-controlled: `createBackgroundReleaseHandle()`
mints a sentinel path, the scripted command blocks until that file
appears, and the test releases it exactly when it wants the dot to clear.
One clock instead of three, and the "turn done, process still running"
state is stable rather than a window to catch. The wait is bounded (60s)
so a forgotten release can't hang a worker, and `sleep 5` stays as the
default for callers that pass no handle.
No product code touched — E2E harness only.
The unit tests cover each layer in isolation, but nothing exercised the whole
chain the bug lived in: the real gateway persisting an attachment, SessionDB
holding it after the process exits, and the renderer rebuilding a thumbnail
from the stored turn.
Seeds a session through the real gateway with an image attached, then launches
desktop against it — so the first render is already the relaunch case. Pins
native image routing (the majority path, and the one where a text-only persist
override is dropped) and stages the file behind directory and file names with
spaces, mirroring the macOS composer's Application Support path.
The 'queues an Enter-submitted draft while compaction is active' test
pastes a large message to push the session over the fixture's 22k
threshold_tokens. At repeat(500) the payload is only ~4k tokens — the
other ~18k came from the ambient system prompt (tool schemas + skills
index + memory), leaving the trigger margin-less. The hermes-agent
skill hub restructure (e3d524b482) shrank the bundled skills index by
~160 tokens and dropped the total just under threshold: compaction never
started, waitForHeldCompletion() hung, and the test timed out at 90s on
every branch since — including main (first red run: 9a4d1a0130).
Bump the payload to repeat(1500) (~12.4k tokens, total ~30k) so the
test crosses the threshold on its own weight with ~8k tokens of margin,
and document the invariant so the next prompt-weight change doesn't
resurrect this.
* fix(desktop): hide persisted agent-only history scaffolding
Filter verification-stop nudges and context-compaction handoffs at the
stored-history mapper boundary. Preserve a real reply when a compaction
handoff shares its stored message.
* test(desktop): build persisted E2E sessions through the real agent
Drive tui_gateway.entry over its stdio JSON-RPC transport against the mock
provider, wait for real completion events, and persist normal session history
through AIAgent and SessionDB. Migrate resume and hidden-history coverage,
including real compression and live verify-on-stop scaffolding, then remove
the unused direct SessionDB import scripts.
* fix(desktop): use the provisioned Python for real-session E2Es
Run the stdio gateway through uv's synced project environment outside the
Nix dev shell, while retaining the fully provisioned Nix Python when the
shell advertises HERMES_PYTHON_SRC_ROOT.
* fix(nix): expose the provisioned Python environment to uv
Mark the Nix-built Python environment active in the dev shell so the shared
E2E session builder can always run through `uv run --active --no-sync`.
* fix(timeline): persist typed display events
* fix(timeline): strip display-only fields from provider payloads, preserve through rewrites, fix /resume display history
Three review findings from PR #69771:
1. Provider payload leak: display_kind and display_metadata were forwarded
to the provider API as unknown message fields. Strict OpenAI-compatible
backends can reject the next request after a model switch or resumed
typed event. Strip both from the per-request api_msg copy in
conversation_loop alongside the existing api_content pop.
2. Rewrite/import data loss: _insert_message_rows preserved display_kind
but silently dropped display_metadata. After replace_messages,
archive_and_compact, or session import, async-delegation completion
events lost their task counts and fell back to generic display text.
Add display_metadata to the INSERT columns and bind tuple.
3. CLI /resume stale recap: startup --resume A set _resume_display_history
from A's lineage. A subsequent in-session /resume B loaded B only into
conversation_history via get_messages_as_conversation, leaving the stale
A display projection. _display_resumed_history preferentially read the
stale attribute, showing A's recap for B. Switch /resume to
get_resume_conversations and update _resume_display_history alongside
conversation_history.
Tests: 890 Python (5 files), 35 desktop TS — all green.
* feat(tui): render typed display events as ◈ markers in the Ink TUI
The TUI was not handling display_kind at all — model switch markers and
async delegation completions rendered as opaque user messages with the
full [System: ...] text, and hidden compaction handoffs were visible.
Wire display_kind through the full TUI chain:
- _history_to_messages (tui_gateway/server.py) forwards display_kind
and display_metadata to the gateway transcript payload.
- GatewayTranscriptMessage (gatewayTypes.ts) gains both fields.
- Msg.kind (types.ts) gains 'event' value.
- toTranscriptMessages (domain/messages.ts) maps:
- hidden → skip entirely
- model_switch → event "model changed"
- async_delegation_complete → event "N background agents finished"
(or "background agent work finished" without metadata)
- messageGroup (blockLayout.ts) routes event to its own group, with
SELF_SPACED + PAINTS_TRAILING_GAP so it owns its margins.
- messageLine.tsx renders event-kind as a dim ◈ marker with no gutter,
matching the CLI's ◈ event rendering.
- 4 new TUI tests for hidden/model_switch/async_delegation mapping.
TUI typecheck: clean. TUI lint: 0 errors (2 pre-existing warnings).
TUI tests: 9 passed (1 pre-existing failure on main, unrelated).