Commit Graph

72 Commits

Author SHA1 Message Date
Teknium e9dd0bf5d5 feat(desktop): polish bot roster sections — dialog rename, Undo delete, Esc-cancel drag, nested under gateways (salvage #100745)
Follow-up on @fortun8te's user-made roster sections:

- Sections start empty: no seeded General/Workforce/Clients. With no
  sections created the roster renders exactly as before.
- New section and Rename go through one Dialog + Input + Cancel/Save
  (the app's session-rename shape) instead of an inline caret; the row
  menu's "New section…" files the bot as it creates.
- Delete needs no confirmation: bots return to Unassigned and the toast
  offers Undo (restores the section in its slot and refiles its bots).
- Drag: single-row drag under a private MIME type, every valid target
  shows a faint outline while a drag is live, the hovered target lights
  up, the source section refuses the drop, Escape cancels, and the moved
  row no longer stays faded after it remounts under its new section.
- Multi-select (cmd/shift-click, querySelectorAll shift-range) dropped:
  the roster has no selection model. Per-bot saveBotMeta writes run in
  sequence, one per profile (membership IS a field on each profile).
- Section heading reuses RosterSectionHeader (gains `action` /
  `onDoubleClick`), so user sections fold and look like the gateway
  headings; ⋯ menu and right-click drive the same Rename / Move up /
  Move down / Delete. Empty sections show a dashed "Drag bots here" slot.
- Composes with gateway buckets: sections nest INSIDE each connection
  bucket, indented under a hairline rail (membership lives in the bot's
  profile on that gateway); empty sections repeat there only mid-drag.
- Full i18n parity (en / ja / zh / zh-hant) for every new string; icon
  toggle and the storage-async plumbing removed.
- Tests trimmed to the three invariants (membership persists through
  saveBotMeta + reload, remainder = Unassigned, delete returns bots +
  undo) plus a live Electron e2e covering the whole flow.
- Docs: "Organize bots into sections" in user-guide/bot-mode.md.
2026-09-02 06:06:32 -07:00
Teknium 209de12d5f fix(desktop): Bot Mode tabs caption a Bot Chat with the bot's name, not "Bot Chat"
Every bot's canonical chat is stored under the same title ("Bot Chat" — the
name the gateway resolves it by, and an invariant roster-actions.ts's stale-tile
probe and #90102 rely on), so the main tab strip captioned every open bot chat
identically and two bots' tabs were indistinguishable (#99152).

Fix at the presentation layer, leaving the stored title and tabTitle untouched:
- workspace-scope.ts gains `$workspaceOwnerLabels` + `workspaceOwnerTitle()`:
  a bots-mode tab whose resolved title still equals its registered placeholder
  reads its owner's label instead. Side threads / Sessions tabs are untouched.
- session-tile.tsx captions tiles through it (and the drag payload); the main
  `workspace` tab (controller.tsx) does the same via `$botChatScopes`, the
  bot-mode scope the main tab was last opened under (it has no tile).
- The hermes-bots roster publishes displayName() per owner key through the new
  `host.setWorkspaceOwnerLabel` (feature-detected), so renames follow.

Supersedes #99177, which set tabTitle at open time — that reverts after mount
because tileTitle() prefers the stored row's title once the hidden row is
upserted, and breaks the `workspaceTabTitle === 'Bot Chat'` invariant.

Tests: one unit test on workspaceOwnerTitle() (bot chat → bot name; side
thread / sessions tab / unlabeled owner untouched) and one Electron e2e
(tab strip reads "Alpha", not "Bot Chat"); both fail on main, pass here.

Closes #99152
Supersedes #99177

Co-authored-by: twotnguyen <nguyenngoctinh011258@gmail.com>
2026-09-02 05:38:10 -07:00
Teknium 6e7c7c7da9 fix(desktop): a bot row click always lands on the Bot Chat the row previews
A plain roster click fronted whatever bots-workspace tab the user last had
active for that bot (#96649). A '+' side thread persists in Local Storage
across restarts, so it won every click forever while the row kept previewing
the canonical Bot Chat (profiles.list canonical_session) — sidebar and center
described two different conversations; a message typed there landed in the
side thread and the row never moved. Support thread "[Bots] - Sessions is not
in sync again" (bundle 7dfff039), reproduced live on origin/main.

- roster-actions: the open-tab shortcut may front only the canonical chat
  (registry id or lineage tip, via a new onlyStoredIds allowlist on
  focusWorkspaceOwnerSessionTile); anything else resolves the registry and
  opens in place. Side tabs stay open beside it. "Open Bot Chat" in the row
  menu is the same action; the `canonical` option goes away.
- roster-actions: when the FOCUSED Bot Chat's canonical session advances on
  the gateway (cron bot-chat delivery, message_agent, group round, CLI turn —
  none reach this window's stream), re-open it in place so the transcript
  refreshes instead of waiting for an app restart (#99393 class).

Tests: the fronting-shortcut unit file and its e2e spec pinned the reversed
behavior; replaced by one unit file (5 tests) and one e2e spec that fails on
main and passes here. group-to-local-bot-handoff e2e still passes.
2026-09-02 03:41:44 -07:00
Teknium c64054a26b test(desktop-e2e): run the mock-provider suite gate-free (approvals off)
The scripted turns execute real terminal commands, and the sidebar
sentinel-wait loop trips the dangerous-command guard: the turn parks
behind a Run/Reject approval card, and the default 'smart' mode fires an
aux LLM approval call at the same mock provider — consuming a
scripted-turn index and never resolving. On the slower CI runner this
stalled the sidebar-dot family (sidebar-states 157/245, tile-unread 166)
until spec timeout; run 33543723331's error-context snapshots show the
approval card blocking each stalled turn. Locally the race usually won
the other way, which is why these passed on dev machines.

Fix: fixtures write 'approvals: mode: "off"' into the mock provider
config by default (specs supplying their own approvals: section own it),
mirroring the auto-title default. Also drop the DOT-DEBUG diagnostics
from tile-unread-bug now that the root cause is identified.

Local: sidebar-states + tile-unread + correction-session-switch all
green in seconds (3-9s vs 90s timeouts); full suite 62 passed /
11 skipped / 1 flaky-passed.
2026-09-01 12:04:29 -07:00
Teknium eb1b14b952 test(desktop-e2e): fix spec drift accrued while the lane was disabled
The Desktop E2E lane was disabled Aug 2 – Sep 1; the app and gateway kept
moving, so 16 specs rotted against current main. All failures traced to
spec/harness drift, not product regressions:

- fixtures.ts: title generation now rides the main model (#83636), firing a
  background completion at the mock after every turn — it contains the whole
  conversation (trigger keywords included), advancing scripted-turn indices
  and tripping hold-for-prompt matchers. Disabled by default in the mock
  provider config; specs supplying their own `auxiliary:` section own it.
- chat/interim-messages/session-compression/correction-session-switch/
  hidden-history-messages: busy-state and transcript assertions updated to
  the current composer aria-labels, interim-message semantics, and
  verify-on-stop continuation behavior on main.
- bot-mode-closed-chat-stays-closed/group-to-local-bot-handoff: Bot Chat tab
  selectors updated for the Bot Mode rework (tabs keyed by
  connection+profile, renamed tab triggers).
- glyph-spinner: assertions made compositor-honest for the CI runner
  (steps() keyframes + layer promotion probed via the animation registry
  instead of GPU-dependent screenshots).
- sidebar-states/tile-unread-bug: event-driven waits with mock-server
  release handles replace wall-clock polls that lost races on loaded
  runners.
- warm-resume-jitter/image-attachment-resume: real-session-builder harness
  waits for the thread viewport before evaluating; failure path now dumps
  per-surface pane state.

Local full-suite run on the CI-equivalent xvfb setup: 62 passed,
11 skipped, 1 flaky-passed (correction-session-switch live-correction spec,
passes on retry). No product code changed.
2026-09-01 12:04:29 -07:00
Jack Lau 7a1fca6676 fix(desktop): resolve the e2e Electron binary per platform and layout
findElectron() probed exactly one path, and got three things wrong at
once for anyone not on a hoisted POSIX install:

* It looked only under the REPO ROOT. This is an npm workspaces repo and
  npm hoists a dependency only when nothing conflicts, so `electron`
  installing into apps/desktop/node_modules is an ordinary outcome, not a
  broken tree.
* It joined a bare `electron`. On Windows the dist file is
  `electron.exe`, so the probe could never match there.
* Its PATH fallback spawned `which`, which is not a command on Windows,
  so the fallback failed for a reason unrelated to whether electron is on
  PATH.

The three combine into a misleading error: the suite refuses to start
with 'Run "npm install" from the repo root' on a tree that has electron
installed. Reproduced on Windows 11 against this repo, where
apps/desktop/node_modules/electron/dist/electron.exe exists and the old
body throws that message; the reporter on #88036 hit the same thing on
Linux and had to hand-symlink the package before the suite would run.

Resolution now asks the installed `electron` package for its own path
first (its main export IS the absolute executable, resolved from
path.txt and honouring ELECTRON_OVERRIDE_DIST_PATH), then falls back to
explicit dist probes for each root, then to PATH with the platform's
lookup command. The error message lists what was searched.

The rules live in e2e/electron-binary.ts so they can be unit-tested
without importing the Playwright runner, with the platform passed in
rather than read from process.platform: reading it would leave every
Windows rule untested on the Linux CI runner.

Wiring: the vitest `electron` project picks up e2e/**/*.unit.test.ts and
Playwright ignores the same pattern, so helper unit tests run in exactly
one runner and the specs are untouched.

Verified: 5 unit tests pass; mutation-checked one rule at a time
(hardcoding the binary name fails 2, reversing the probe order fails 1,
hardcoding `which` fails 1). tsc -p . and tsc -p tsconfig.e2e.json
clean.

This is the environment blocker called out in #88036, not its rendering
bug, so it is deliberately a subset.

Refs #88036
2026-08-31 11:56:42 -07:00
Finn763 38b93e0abe fix(desktop): keep @tanstack/react-query in one runtime chunk (#95560)
The packaged app crashed at launch with 'No QueryClient set, use
QueryClientProvider to set one': useQuery in a lazy chunk (session-list-density)
read a second @tanstack/react-query runtime whose QueryClientContext was never
populated by the entry's QueryClientProvider. The source tree was correct — the
duplication happened at build time, because react-query was the one
context-bearing runtime not pinned to a shared vendor chunk, and rolldown's
merge heuristics inline the spare copy into a lazy chunk depending on toolchain
version.

- vite.config.ts: add @tanstack/react-query to the vendor-react
  advancedChunks group + dev dedupe list, mirroring the react-router fix.
- assert-dist-built.mjs: fail the build when the 'No QueryClient set'
  invariant appears in more than one JS asset (launch-smoke guard).
- assert-dist-built.test.mjs: unit tests for the new invariant check.
- launch-packaged-app.spec.ts: e2e smoke test asserting the packaged app
  boots to real UI, not the QueryClient error boundary.
2026-08-31 10:10:35 -07:00
Zeus-Deus 7c910793bf fix(desktop): a bot row click returns to its open tabs instead of re-opening a closed Bot Chat
In Bot Mode every roster click resolved the bot's canonical "Bot Chat" by
name and opened it as a tab. Nothing records a tab close (the plugin keeps no
closed set; core's tile bucket only forgets), so a Bot Chat the user had
closed came back beside every newer thread on every bot switch — close it,
start a new thread, visit another bot, come back: two tabs again, forever.

A row click is now "go to this bot": when the bot's workspace already holds
tabs, the one the user last had active is fronted and no chat is resolved or
opened. The canonical chat is opened only when the bot has nothing open, or
on the explicit asks — a new "Open Bot Chat" row-menu item and the Bots home
"Open chat" button (`openRosterBot(bot, { canonical: true })`).

- session-states: `focusWorkspaceOwnerSessionTile(ownerKey)` fronts the
  owner's remembered-active tile (else its most recent) and reports it.
- sdk: `host.focusOpenWorkspaceSession(ownerKey)` exposes it to plugins;
  feature-detected in the plugin so older shells keep the canonical open.
- hermes-bots: `focusExistingBotTab` short-circuits `openRosterBot`; the
  claim it records carries only the fronted tab, and the session.reclaimed
  re-resume now skips such claims so it cannot resurrect the closed chat.

Tests: vitest for the core helper, a node test for the click path (open tabs
win, nothing open → canonical, explicit canonical, older shell, throwing
host), and an e2e that seeds two bots with real "Bot Chat" rows, closes one,
starts a thread, switches bots and back, and asserts the Bot Chat stays
closed until asked for explicitly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 20:30:49 -07:00
Teknium 778feee37a chore: eslint --fix on salvaged e2e spec 2026-08-26 09:47:23 -07:00
funky-xamarin 9ea7a37cc1 fix(desktop): restore live group after bot open failure 2026-08-26 09:47:23 -07:00
funky-xamarin fbd149035b test(desktop): verify group-to-local-bot handoff in Electron 2026-08-26 09:47:23 -07:00
funky-xamarin 35ee27b1c6 test(desktop): add RED group-to-local-bot E2E 2026-08-26 09:47:23 -07:00
Zeus-Deus fd565c80e9 feat(desktop): fleet profile rail — every registered gateway's agents on one strip
With several gateways registered, the Sessions profile rail only ever showed
the active gateway's profiles; reaching a bot on another machine meant a
gateway switch first, then a click on the rail that appeared afterwards. Bot
Mode (#91134) and Capabilities already read the union agent roster; the rail
is now its third consumer.

- Every registered gateway's profiles sit on the one strip, in registry order
  (This device first, then by label), each group headed by that gateway's
  kind glyph. The active gateway's squares are unchanged; the others are
  "at rest" (dimmed) with tooltips/accessible names qualified by machine
  (`inbox · Homelab`), so same-named profiles never read alike.
- Clicking an at-rest square performs the same dial → commit → re-home as
  the statusbar switcher, landing on that exact (gateway, profile):
  `selectConnection(id, { profile })`. The spinner sits on the clicked
  square; the previous source stays painted until the target answers.
  Groups keep their slots whichever gateway is active, so a square never
  moves under the pointer that clicked it.
- Right-click on an at-rest square: Switch to / Color / Rename / Edit
  SOUL.md / Delete, executed on the owning gateway (renameProfile,
  getProfileSoul and updateProfileSoul accept the same scope deleteProfile
  already had); the delete confirmation names the machine. The legacy
  per-profile "Connect to a remote host…" item is hidden on multi-gateway
  setups, where the rail shows machines directly.
- Unreachable gateways keep their squares with an amber dot on the glyph;
  two registrations of one backend collapse to one group; past thirteen
  squares across the fleet the strip condenses into a menu sectioned by
  gateway. Roster is fetched on mount / focus / registry change only — no
  periodic fleet polling.
- Single-gateway Desktops render exactly as before: no roster fetch, same DOM.

Also fixes a boot race the e2e surfaced: initializeConnectionsRegistry()
"restored" the launch-mode source over a switch the user had already made
while boot was settling (same class as #91047). The restore now yields when
a switch is pending or already landed.

Tests: pure grouping (fleet-rail.test.ts), rail component fleet mode
(profile-rail-fleet.test.tsx), store (explicit profile pick; restore yields),
and a Playwright e2e (fleet-profile-rail.spec.ts) that boots Desktop with two
REAL backends — the local one plus a second `hermes serve` registered as a
remote URL connection — and verifies layout, a real re-home, gateway-scoped
actions, and order stability.

Docs: multi-connection-desktop.md describes the fleet rail.

Refs #89304, #92384, #91047, #94724

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 08:39:45 -07:00
Teknium 6a6e16fa5d feat(desktop): OS-keychain encryption for stored secrets is now opt-in — no more macOS Keychain password prompt on every launch
Electron safeStorage parks a per-app key ('Hermes Key') in the macOS login
keychain; on machines with a locked/missing/corrupted default keychain that
turned every Hermes Desktop launch into a blocking 'Keychain Not Found' /
password dialog. Keychain-backed encryption is now an explicit opt-in:

- electron/secret-storage-policy.ts: standalone policy seam (default OFF,
  strict === true coercion, one-shot migration flag) + unit tests
- default path never calls any safeStorage API (including
  isEncryptionAvailable, which itself touches the keychain)
- one-shot legacy migration decrypts existing safeStorage blobs to plain
  0600 files at first launch; undecryptable blobs are kept but read as
  absent afterward (classify 'drop') so a dead keychain prompts at most once
- Settings -> Gateway toggle (all 5 locales) re-encodes every stored secret
  store in place when flipped (v1 connection.json, v2 connections.json,
  native-oauth-tokens.json)
- e2e: at-rest spec now covers both postures (opted-in unchanged contract,
  default saves without secure storage, owner-only bits, restart round-trip)
- docs: multi-connection-desktop + desktop-native-signin updated
2026-08-25 14:10:19 -07:00
Brooklyn Nicholson f8b52e4d80 test(desktop): cover UI scale across recordless hash routes
Drives the reported path rather than the helper: set a non-default
scale, then navigate to routes Chromium holds no zoom record for, which
is what opening a new session looks like to the per-URL store. Keeps the
Cmd/Ctrl+N case alongside it.

Co-authored-by: Clark Vines <38430798+clarkvines@users.noreply.github.com>
2026-08-24 22:01:47 -05:00
Royalaid 76e0ca8826 fix(desktop): order-independent selection guard, layer-hint scoping, honest compositor receipt
Review-round residuals: the spinner's user-select guard now beats
[data-selectable-text] regardless of stylesheet order; the will-change
layer hint clears under the global renderer pause and reduced motion so
parked spinners hold no compositor layer; the e2e travel assertion reads
the engine's keyframes (a computed transform always serializes to a
matrix, so the old '%' check could never fail); inline import() type
hoisted for the lint gate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 04:23:41 -07:00
Royalaid 0269504250 test(desktop): exercise the spinner CSS and the invalidation scope for real
Replaces the deleted stylesheet-text assertions with tests that run the
thing they claim to cover.

e2e/glyph-spinner.spec.ts drives a real browser, where the CSS actually
executes: the strip's animation resolves to steps(N) for N frames, runs
infinitely, and travels a resolved length rather than a percentage (a
percentage translate is layout-dependent and Chromium refuses to
composite it). Both pause gates are covered — the per-spinner
`data-paused` attribute and the global renderer-pause attribute that
window blur / minimize / document-hidden arm — along with the layer
promotion being scoped to running spinners. A sampling test confirms the
transform visits a bounded number of distinct values across one cycle
(steps, not a linear sweep) and that nothing mutates the DOM while it
animates, which is the property the whole change exists to deliver.

status-invalidation-scope.test.tsx pins the scoping itself as a render
count. `useTapbackDoubleClick` is called by AssistantMessageBody and by
nothing else in the tree, which makes it an exact render counter for the
message root without exporting internals. A settle and a delta flush must
both leave that count untouched while the leaves update. Verified by
mutation: reinstating a root-level status subscription fails the settle
test (2 renders where 1 is required).

It also pins node identity across the settle transition, so the
inter-agent collapse cannot go back to swapping element types at the
message-root position and remounting the row under the scroll anchor.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QwTc9XqUjhbay446VjugHZ
2026-08-21 04:23:41 -07:00
brooklyn! 00e5a361b6 Merge pull request #89611 from NousResearch/bb/desktop-godfiles
refactor(desktop): decompose god files into atomic modules
2026-08-19 11:05:01 -05:00
Brooklyn Nicholson 808fde106d test(desktop): share the duplicated fixtures
The same AST sweep over specs found fixtures maintained in parallel
across suites that have no reason to know about each other.

Twenty-one specs each mounted useMessageStream themselves and ten of the
harnesses were byte-identical; twenty now take renderMessageStream, with
overrides for the seams that genuinely vary. The SessionInfo builder was
spelled out field-by-field in seven specs, so a new backend field broke
seven files instead of one. Twenty specs carried their own inert
ResizeObserver and eleven repeated the animation-frame, CSS.escape,
scrollTo and WAAPI stubs the transcript needs to mount at all — split by
scope into src/test/jsdom for what any component might need and the
assistant-ui folder's own kit for the transcript. Plus the window-state
bridge, deferred, the external-store thread runtime, the manual
createRoot harness, and the per-folder caret, env-var, provider and
session fixtures.

Left alone on purpose: the store suites' makePrimary, where the vi.mock
harness around it is the actual duplication and cannot be hoisted out of
a hoisted factory; electron's deferred, where reaching into src/ from
the main process would invert the layering for eight lines; and the two
suites that compose another hook alongside the stream.
2026-08-19 10:51:12 -05:00
Brooklyn Nicholson a8d87ac1d9 fix(desktop): keep the preload bridge alive under the sandbox
Deciding whether the OS can back glass needs os.release(), but every
Hermes window runs its preload with sandbox: true, where require is a
polyfill limited to electron, events, timers and url. The node:os import
threw before contextBridge ran, so window.hermesDesktop was never defined
and the app booted straight into "Desktop IPC bridge is unavailable".

Main already computes both verdicts, so preload asks for them over a
synchronous channel instead. No reply degrades to no glass, which is an
ordinary opaque window rather than a page thinned over nothing.
2026-08-19 10:50:52 -05:00
ethernet 1314a539c2 refactor(desktop): custom translated context menus across the app
One renderer coordinator owns every right-click and replaces the
native Electron menus. Menus are assembled from what the click
landed on:

- Links and images get open/copy/save sections; chat links add the
  reach-aware resolved-URL copy on remote gateways.
- Editables get spell-check suggestions (async-appended when
  Chromium's facts arrive from main), cut/copy/paste, and select
  all. Cut and copy need a selection; paste needs a non-empty
  clipboard; select all needs field content. The verbs show their
  accelerators instead of icons and dispatch a frame after the menu
  closes, so the radix focus trap cannot steal the target. Select
  all runs renderer-side, scoped to the field, because main's
  selectAll acts on the focused frame and could grab the transcript.
- Terminals answer through registered xterm handles; the read-only
  agent terminal hides paste.
- The in-app browser guest builds the same menus from the webview
  tag's context-menu event: Chromium's editFlags gate the edit
  verbs, spell-check rides the event, and Inspect element closes
  every menu. Coordinates arrive as window-relative device pixels,
  so the handler divides by the window zoom factor; guest edit
  commands focus the webview first, because they act on the focused
  webContents.
- Bare app chrome falls back to the window verbs.

Labels come from the locale files in all five languages. Main keeps
thin IPC verbs: edit commands, copy-image-at-gesture, spell-check
actions, and dictionary-add for guests (the tag has no session API).
The e2e spec exercises the real focus trap; it is blocked today by
the gateway-checking stall that also fails e2e/chat.spec.ts.
2026-08-18 23:53:07 -04:00
ethernet 9a9015dfa8 test(desktop): batch clarify E2E spec and mock trigger
The mock server gains a batch clarify trigger that scripts a
two-question clarify turn. The scripted turn fires only while the
conversation has no tool result, so the answered batch falls through
to the canned reply instead of a repeat of the quiz.

The spec runs the real chain from composer to renderer and asserts
one batch card, the staged-answers confirm gate, and the settled
card. The local harness cannot boot the packaged app in this
environment (the pre-existing chat spec fails the same way), so the
proof for this spec is the CI run.
2026-08-18 21:28:53 -04:00
ethernet 25c6516607 feat(desktop): single confirm for the batch clarify card
The batch card previously locked each answer with its own Continue
press. Now picks and typed answers stage locally, and one Confirm and
continue button (enabled when every question has an answer) submits
the whole batch. Staged answers stay editable until that confirm.

The wire protocol is unchanged. The confirm sends the per-question
locks in sequence, because the last lock resolves the blocked tool and
each earlier lock must already be accepted when it lands. Replayed
locked answers from a reconnect pre-stage their questions so restored
progress stays visible. The TUI and CLI keep incremental per-question
locks, so a timeout there still returns partial answers.
2026-08-18 21:28:53 -04:00
ethernet 838c4692bc fix(desktop): merge duplicate batch clarify cards
A batch clarify rendered as two identical interactive cards. The
tool.start row carries the model tool_call_id and the clarify.request
row carries a gateway request_id. The hydration-race merge correlates
the two rows with the top-level question text. A batch payload has no
top-level question, so the rows never matched and the card mounted
twice.

The correlation key for a batch now comes from the joined per-question
texts. The NUL separator cannot occur in real question text, so a batch
key cannot collide with a single-question key.

New coverage: two hydration-race tests for the batch shape (tool.start
first and clarify.request first), a mock-server batch clarify trigger,
and an E2E spec that runs the full chain and asserts exactly one card,
the per-question locks, the Confirm and continue relabel, and the
settled card.
2026-08-18 21:28:53 -04:00
Zeus-Deus d2c2b0b6d4 fix(desktop): persist sidebar unread dots across app restarts
The green "finished — unread" session dot lived only in the transient
$unreadFinishedSessionIds atom, written by a live busy->idle edge the
renderer had to witness. Closing and reopening the app grayed out every
dot, and a session that finished while the app was closed could never
be flagged at all.

Add a persisted layer (session-unread.ts), ported from the webui's
proven design:

- Seen watermarks (hermes.desktop.sessionSeenCounts): the message_count
  last acknowledged per session, keyed by the durable lineage id (same
  rule as session colors). A row whose live count exceeds its watermark
  paints unread on every list refresh - this reconstructs dots after a
  restart AND surfaces sessions that finished while the app was closed.
  First sight of an unknown session seeds the watermark so a fresh
  install doesn't light up every row.
- Explicit finish markers (hermes.desktop.unreadFinishedSessions): the
  live edge, persisted, covering the gap before the sidebar list
  refreshes its counts.

Opening a session acks both; the selected session's watermark tracks
its live count so on-screen activity never reads as unread. Chat and
cron rows get full watermark treatment; messaging rows keep explicit
markers only, so inbound messages don't paint false completion dots.
Profile switches keep persisted markers (keyed by durable id) and only
wipe the transient paint layer, so a round-trip repaints them.

Covered by store unit tests and an e2e spec that boots the app three
times: dot appears on a background finish, survives a restart, clears
on open, and stays cleared after another restart.
2026-08-15 01:21:40 -07:00
Nicky Molina 6f64a2c631 fix(desktop): reanchor transcript on window focus 2026-08-14 20:40:15 -07:00
Michael Huang 2cabeba563 fix(tui,desktop): refresh context usage live during active turns 2026-08-14 20:24:42 -07:00
张豪杰 1535c114c9 test(desktop): cover connection.json owner-only mode end to end
The helpers were tested; nothing proved main.ts called them. Reverting both
call sites and both imports in readDesktopConnectionConfig /
writeDesktopConnectionConfig left the whole suite green (947 passed / 2
skipped, tsc 0, eslint clean, e2e 1 passed 1 skipped) while connection.json
went back to 0644 — the user-visible fix this PR promises was untested.

The e2e spec could not catch it by construction: it asserts the ENCRYPTION
contract with a raw-bytes scan, and safeStorage keeps the token opaque
regardless of the file's mode, so a 0644 file passes that scan every time.
There was no mode assertion anywhere in e2e/.

Adds the missing third contract — unreadable by other local accounts — on all
three paths that can produce the file:

- write: assert the mode of the artifact test 1 already proves the app wrote.
- read, valid file: seed the app's own encrypted connection.json back to 0644
  and assert launch tightens it. Scoped to the MODE only, so it is independent
  of the still-deferred plaintext migration — the fixture's token is already
  ciphertext, so nothing re-encrypts, no #62319 opt-in marker is involved, and
  no rotation guidance is owed.
- read, corrupt file: a truncated file still holds the token bytes and throws
  into the swallowing catch, so it would be the one file never tightened. This
  is the only test that distinguishes the chmod's placement relative to the
  parse.

Also moves the tighten above JSON.parse for exactly that reason, and pins the
cache invariant the placement depends on: the tighten must be a chmod, not a
rewrite, because it sits inside the function whose cache keys on mtimeMs.

Asserted as `mode & 0o077 === 0` rather than `=== 0o600` to avoid a
change-detector, and skipped on win32, where chmod maps to the read-only bit
and the fix deliberately no-ops (ACLs are PR #77527).

Every assertion was mutation-tested: reverting the full wiring fails all three;
reverting only the write path fails only the write test; deleting only the
tighten-on-read fails only the two read tests; moving the tighten below the
parse fails only the corrupt test; making the tighten a rewrite instead of a
chmod fails the mtime assertions. Bundle greps confirmed each mutation reached
dist/electron-main.mjs before the run.

(cherry picked from commit 99cfc16e7cdb759b674d890563f6a82113326547)
2026-08-12 22:38:17 -07:00
张豪杰 7e151bd9d3 fix(desktop): create connection.json owner-only
`connection.json` under the desktop app's Electron `userData` was written with no
file mode, so it landed at the `0644` umask default — while its two
credential-bearing neighbours in the same directory, `desktop-installation.json`
and `native-oauth-tokens.json`, were already `0600`. That file holds the
safeStorage-encrypted gateway token plus the fields that are NOT encrypted: the
gateway URL and the SSH host, user, and key path.

- Route the single write choke point through a helper that creates the file
  owner-only and atomically.
- Tighten an already-existing `0644` file once per launch on the read path, so
  installs that already have one do not stay world-readable until the next save.
- Refuse to act on a path that is a symlink or not owned by the current user,
  matching the guards `desktop-installation.ts` already applies to its sibling.

The symlink guard alone turned out to be insufficient, and that is worth
recording: `writeSecretFileAtomic` tightens its *temp* path, so a symlink planted
at `connection.json.tmp` meant `writeFileSync` followed it, the guard correctly
bailed, and `renameSync` then moved the link onto `connection.json` permanently.
Measured, guard-only vs. as-landed:

    guards only          token leaked: true    config is a symlink: true   755
    guards + temp unlink token leaked: false   config is a symlink: false  600

So the temp path is unlinked before the write.

Issue #77486's headline claim — that a dashboard session token is persisted in
plaintext — does not hold against main. The token has been safeStorage-encrypted
since the desktop app reached mainline in 51c68d4ab, and `encryptDesktopSecret`
aborts with an actionable message rather than degrading to plaintext when
safeStorage is unavailable. The `{ encoding: 'plain', value }` literal does exist
at main.ts:7084, but only on the `persistToken: false` branch, whose sole caller
is the connection-test handler, which never writes. So no mainline path *writes*
a plaintext token. The commits that did contain a plaintext-writing fallback
(d3d177283, d208f2c2c) are not ancestors of main — they live only on
upstream/bb/gui-* and the desktop-pr20059-installers pre-release tag.

At-rest migration of legacy non-safeStorage payloads is deliberately NOT included.
An earlier revision of this branch implemented it and it was removed after review
reproduced two token-loss paths: it force-converts the opt-in plaintext choice
PR #62319 adds (silently reverting the user's decision, then destroying the token
on the next launch without the `--password-store=basic` flag), and it converts a
portable credential into a keychain-bound one with no consent — destroying the
only recoverable copy while not remediating the real exposure, since every
existing backup still holds the plaintext and the true remedy is rotation. It also
persisted raw `parsed`, bypassing `sanitizeConnectionProfiles`. A comment at the
read path records the three preconditions any future attempt needs.

`decryptDesktopSecret`'s non-safeStorage read fallback is untouched — it is what
lets a pre-release or hand-edited config work at all.

Windows still inherits the userData directory ACL rather than an explicit
owner-only one; mode bits are advisory there, so that half is deferred to
PR #77527 rather than growing a second ACL implementation here.

e2e: `at-rest-connection-token.spec.ts` asserts the at-rest contract
implementation-independently — the token's plaintext value (and its base64 form)
must not appear in a raw-bytes scan of any file under userData or HERMES_HOME,
AND the app must still put the exact original token on the wire after a restart,
so a fix that simply drops the token cannot pass. Proven non-vacuous by mutation:
writing `{ encoding: 'plain', value }` still fails the scan while the
file-exists and gateway-URL guards pass. The migration case is a documented
`test.fixme` naming its three blockers.

Electron project 928 -> 924 tests (-9 migration, +5 new guard and
mechanism-isolation). Two of those five exist because reverting either owner-only
mechanism alone initially scored zero failures — they were masking each other, so
either could have been deleted green.

(cherry picked from commit 6e01add6578f08f015a567d3a7a7378f2ec3e768)
2026-08-12 22:38:17 -07:00
Teknium 085a9d332f test(desktop): widen HUD composer containment regression coverage (#82319)
Extend the packaged-app HUD geometry test from horizontal-only to full
containment: both axes for the dock and the input, plus an explicit
assertion that no percentage translate survives on the composer dock.
The vertical clipping reported on Windows (#82203) and macOS (#82214)
is the same escape class on the other axis, and the computed-translate
probe makes a future optimizer regression fail with a diagnosis instead
of a bare coordinate mismatch.
2026-08-09 04:03:46 -05:00
Steve Darlow f2731da4a4 fix(desktop): keep HUD composer within window 2026-08-08 22:54:16 -05:00
Brooklyn Nicholson 41ae3db426 fix(desktop): reveal every window after a missed ready event, not just the main one
Session, instance, HUD, quick-entry and pet-overlay windows all open with
show: false and are revealed only by ready-to-show, so the Electron 40 bug
strands them exactly the way it stranded the primary window — and none of
them have the second-launch workaround that made the main-window case
recoverable.

Generalize the controller to any window and wire all six through one
wireWindowReveal helper. Callers pass their own reveal action (showInactive
for the pet overlay, show + focus for the HUD and quick entry) and their own
post-visible work, so whichever path wins runs them exactly once.

Quick entry now reveals the window the call created rather than whatever
`quickEntryWindow` points at when the event lands.
2026-08-08 16:51:44 -05:00
ethernet b818c427c8 fix(desktop): mount one worktree dialog instead of one per composer
Every CodingStatusRow mounted its own WorktreeDialog and subscribed to the
same global `$newWorktreeRequest` token, so a single ⌘⇧B with two composers on
screen opened two stacked dialogs — dismissing the front one revealed an
identical empty dialog behind it, which read as the dialog "staying open" after
creating a worktree.

Mount it exactly once in the sidebar (beside ProjectDialog) and drive it from a
`$worktreeDialog` atom, mirroring how the project dialog already works. One
mount cannot double-open. Every entry point (⌘⇧B, the rail's kebab, the
sidebar's + button) now publishes intent instead of rendering its own copy; the
rail and the button pin their own repo so a tile's kebab still targets that
tile's worktree.

The target is resolved at open time by `resolveWorktreeRepoPath`, which walks
the focused surface's cwd then the entered project's root, validating each
candidate against the repo-status probe cache — a project's root folder is not
necessarily a git repo, so existence alone isn't proof. That makes the resolver
the sole authority, so the hotkey no longer pre-gates on `$repoStatus` and now
works from a detached session that sits inside a project. When nothing in reach
is a repo it is a silent no-op: a worktree only exists inside a repo, so there
is nothing to report.

Also adds a project picker to the dialog so the repo can be retargeted before
naming the branch.

E2E: extends worktree-branch-status.spec.ts with a 10-branch repo, visual
snapshots of the base-branch picker and the convert-branch view, a geometry
assertion that the picker isn't clipped by the dialog (fails headlessly on
regression rather than waiting for a human to compare diff images), and a
two-composer test asserting one keypress opens exactly one dialog. Tests 1 and
4 fail against the previous code and pass now.
2026-08-05 11:43:09 -04:00
Johnny 1051434325 perf(desktop): isolate right pane layout work 2026-08-03 13:39:33 +05:30
Brooklyn Nicholson 4a9754b3c3 feat(desktop): default UI zoom to Appearance 90% preset
Land on the exact 90% scale so the settings control shows a selection
on fresh installs, instead of five Ctrl+- steps (~91%) between presets.
2026-07-28 04:00:55 -05:00
Brooklyn Nicholson 9ad400580c fix(desktop): pin E2E sandboxes to Chromium zoom 0
Fresh installs now default to ~91%, but Playwright hit-testing and
visual baselines still assume 100%. Seed zoom-state.json so isolated
E2E profiles don't inherit the product default.
2026-07-28 03:33:10 -05:00
Brooklyn Nicholson 579b66336f fix(desktop): let automated teardown quit past the active-work prompt
Playwright closes the app with a turn still in flight, so the new quit
confirmation waited on a click nobody was there to make and the E2E
worker died on a 90s teardown timeout.
2026-07-27 16:19:27 -05:00
brooklyn! 825069004a Merge pull request #72504 from NousResearch/bb/desktop-real-session-perf
perf(desktop): 60fps on real sessions — reflow-gated pins, adaptive flush, stream-aware backfill
2026-07-27 01:50:41 -05:00
Brooklyn Nicholson 437b9b1204 test(desktop): wait for backfill before the duplicate-count baseline
The large-session-resume E2E captured initialMockReplyCount immediately
after openSeededSession, which returns once the NEWEST turn is in the
viewport. With FIRST_PAINT_BUDGET=20 (lowered from 60 in this branch),
only the newest ~10 turns mount at first paint; the older turns
backfill in a rAF. The baseline was reading 10 instead of 27, so once
the backfill mounted the full 28 (27 seeded + 1 new), the test saw
"28 ≠ 11" and reported duplicates that were never there.

Wait for the oldest seeded turn to mount before taking the baseline.
This makes the count reflect the fully-mounted transcript regardless
of FIRST_PAINT_BUDGET, so the perf win (smaller first paint) and the
no-duplicate invariant both hold.

Refs #72504
2026-07-27 01:44:42 -05:00
Gille c7b75a7849 test(desktop): isolate compression from slash completion 2026-07-26 23:34:29 -07:00
Gille b1081c1d22 test(desktop): assert compress argument stage 2026-07-26 23:34:29 -07:00
Gille 76e416052c test(desktop): wait for committed compress directive 2026-07-26 23:34:29 -07:00
Brooklyn Nicholson f2f5b32531 fix(desktop): expect keep-alive tabs not to repaint on reactivation 2026-07-26 05:11:46 -05:00
Brooklyn Nicholson 76e9ac3915 fix(desktop): preserve correction order in session tabs 2026-07-26 05:00:18 -05:00
Brooklyn Nicholson fc39c7ac31 test(desktop): scope e2e transcript helpers to the active chat surface
The sidebar "+" now stacks a tab instead of replacing the surface, so the
prior session stays mounted and several chat surfaces can be on the page at
once. Helpers that waited for the old transcript to disappear from the page
timed out, and `.first()` locators / bare `document.querySelector` calls
started resolving against the wrong session (CI's "resolved to 2 elements"
strict-mode violation).

Target the most recently mounted `[data-composer-target]` surface instead,
and assert the NEW surface is empty rather than waiting for the old text to
vanish.
2026-07-26 03:57:49 -05:00
Hermes Agent 21a2185f86 fix(desktop-e2e): poll for the finished-unread dot instead of sampling once
Sentinel-released processes exposed a second bare-sample assertion in the
same file. The unread-dot check ran a synchronous .count() 140ms after the
running dot cleared:

  15.91s  poll "dot should disappear" -> 2
  16.05s  ... -> 0  (running dot gone)
  16.05s  bare .count() for unread dot -> 0  FAIL

"Finished — unread" is an event-driven transition that lands just after the
running dot clears. The old fixed `sleep 5` happened to leave enough slack
between the two that a single sample usually caught it; releasing the
process deterministically removed that incidental slack and made the latent
race deterministic instead.

Poll for it, matching how sidebar-states.spec.ts already asserts this exact
dot. The split-tile assertion at line 234 stays a bare sample on purpose —
it asserts an absence (toBe(0)), where polling would only wait for something
that must never appear.
2026-07-26 00:12:36 -07:00
Hermes Agent 3a3bc41c7e fix(desktop-e2e): end the sidebar background-dot wall-clock race
The cross-session sidebar specs asserted a state that could expire before
they looked at it, making them the flakiest tests in the suite — two reds
on unrelated PRs within three minutes on 2026-07-26.

Root cause, from the failing run's trace: the tests need a background
process that is still RUNNING after the agent turn finishes, but the
process was a fixed `sleep 5` racing two other clocks — the turn itself
(two model round trips plus a real subagent delegation) and the 4s
success linger before a finished task auto-dismisses. On a loaded runner
the "dot should appear" poll took 7.5s to see the dot; by then `sleep 5`
had already exited, `waitForFunction(finalText)` returned in 0.08s
because the turn was long done, and the next line — a bare synchronous
`.count()`, not a wait — sampled 0.

The process lifetime is now test-controlled: `createBackgroundReleaseHandle()`
mints a sentinel path, the scripted command blocks until that file
appears, and the test releases it exactly when it wants the dot to clear.
One clock instead of three, and the "turn done, process still running"
state is stable rather than a window to catch. The wait is bounded (60s)
so a forgotten release can't hang a worker, and `sleep 5` stays as the
default for callers that pass no handle.

No product code touched — E2E harness only.
2026-07-26 00:12:36 -07:00
Brooklyn Nicholson f71ba11d4c test(desktop): cover attached-image resume end to end
The unit tests cover each layer in isolation, but nothing exercised the whole
chain the bug lived in: the real gateway persisting an attachment, SessionDB
holding it after the process exits, and the renderer rebuilding a thumbnail
from the stored turn.

Seeds a session through the real gateway with an image attached, then launches
desktop against it — so the first render is already the relaunch case. Pins
native image routing (the majority path, and the one where a text-only persist
override is dropped) and stages the file behind directory and file names with
spaces, mirroring the macOS composer's Application Support path.
2026-07-24 21:20:52 -05:00
Teknium b29ee6a650 test(desktop): give the auto-compaction E2E real headroom over threshold_tokens
The 'queues an Enter-submitted draft while compaction is active' test
pastes a large message to push the session over the fixture's 22k
threshold_tokens. At repeat(500) the payload is only ~4k tokens — the
other ~18k came from the ambient system prompt (tool schemas + skills
index + memory), leaving the trigger margin-less. The hermes-agent
skill hub restructure (e3d524b482) shrank the bundled skills index by
~160 tokens and dropped the total just under threshold: compaction never
started, waitForHeldCompletion() hung, and the test timed out at 90s on
every branch since — including main (first red run: 9a4d1a0130).

Bump the payload to repeat(1500) (~12.4k tokens, total ~30k) so the
test crosses the threshold on its own weight with ~8k tokens of margin,
and document the invariant so the next prompt-weight change doesn't
resurrect this.
2026-07-23 21:05:35 -07:00
ethernet a4bc1ca502 fix(timeline): persist typed display events (#69771)
* fix(desktop): hide persisted agent-only history scaffolding

Filter verification-stop nudges and context-compaction handoffs at the
stored-history mapper boundary. Preserve a real reply when a compaction
handoff shares its stored message.

* test(desktop): build persisted E2E sessions through the real agent

Drive tui_gateway.entry over its stdio JSON-RPC transport against the mock
provider, wait for real completion events, and persist normal session history
through AIAgent and SessionDB. Migrate resume and hidden-history coverage,
including real compression and live verify-on-stop scaffolding, then remove
the unused direct SessionDB import scripts.

* fix(desktop): use the provisioned Python for real-session E2Es

Run the stdio gateway through uv's synced project environment outside the
Nix dev shell, while retaining the fully provisioned Nix Python when the
shell advertises HERMES_PYTHON_SRC_ROOT.

* fix(nix): expose the provisioned Python environment to uv

Mark the Nix-built Python environment active in the dev shell so the shared
E2E session builder can always run through `uv run --active --no-sync`.

* fix(timeline): persist typed display events

* fix(timeline): strip display-only fields from provider payloads, preserve through rewrites, fix /resume display history

Three review findings from PR #69771:

1. Provider payload leak: display_kind and display_metadata were forwarded
   to the provider API as unknown message fields. Strict OpenAI-compatible
   backends can reject the next request after a model switch or resumed
   typed event. Strip both from the per-request api_msg copy in
   conversation_loop alongside the existing api_content pop.

2. Rewrite/import data loss: _insert_message_rows preserved display_kind
   but silently dropped display_metadata. After replace_messages,
   archive_and_compact, or session import, async-delegation completion
   events lost their task counts and fell back to generic display text.
   Add display_metadata to the INSERT columns and bind tuple.

3. CLI /resume stale recap: startup --resume A set _resume_display_history
   from A's lineage. A subsequent in-session /resume B loaded B only into
   conversation_history via get_messages_as_conversation, leaving the stale
   A display projection. _display_resumed_history preferentially read the
   stale attribute, showing A's recap for B. Switch /resume to
   get_resume_conversations and update _resume_display_history alongside
   conversation_history.

Tests: 890 Python (5 files), 35 desktop TS — all green.

* feat(tui): render typed display events as ◈ markers in the Ink TUI

The TUI was not handling display_kind at all — model switch markers and
async delegation completions rendered as opaque user messages with the
full [System: ...] text, and hidden compaction handoffs were visible.

Wire display_kind through the full TUI chain:

- _history_to_messages (tui_gateway/server.py) forwards display_kind
  and display_metadata to the gateway transcript payload.
- GatewayTranscriptMessage (gatewayTypes.ts) gains both fields.
- Msg.kind (types.ts) gains 'event' value.
- toTranscriptMessages (domain/messages.ts) maps:
  - hidden → skip entirely
  - model_switch → event "model changed"
  - async_delegation_complete → event "N background agents finished"
    (or "background agent work finished" without metadata)
- messageGroup (blockLayout.ts) routes event to its own group, with
  SELF_SPACED + PAINTS_TRAILING_GAP so it owns its margins.
- messageLine.tsx renders event-kind as a dim ◈ marker with no gutter,
  matching the CLI's ◈ event rendering.
- 4 new TUI tests for hidden/model_switch/async_delegation mapping.

TUI typecheck: clean. TUI lint: 0 errors (2 pre-existing warnings).
TUI tests: 9 passed (1 pre-existing failure on main, unrelated).
2026-07-23 14:46:24 -04:00