The scripted turns execute real terminal commands, and the sidebar
sentinel-wait loop trips the dangerous-command guard: the turn parks
behind a Run/Reject approval card, and the default 'smart' mode fires an
aux LLM approval call at the same mock provider — consuming a
scripted-turn index and never resolving. On the slower CI runner this
stalled the sidebar-dot family (sidebar-states 157/245, tile-unread 166)
until spec timeout; run 33543723331's error-context snapshots show the
approval card blocking each stalled turn. Locally the race usually won
the other way, which is why these passed on dev machines.
Fix: fixtures write 'approvals: mode: "off"' into the mock provider
config by default (specs supplying their own approvals: section own it),
mirroring the auto-title default. Also drop the DOT-DEBUG diagnostics
from tile-unread-bug now that the root cause is identified.
Local: sidebar-states + tile-unread + correction-session-switch all
green in seconds (3-9s vs 90s timeouts); full suite 62 passed /
11 skipped / 1 flaky-passed.
The Desktop E2E lane was disabled Aug 2 – Sep 1; the app and gateway kept
moving, so 16 specs rotted against current main. All failures traced to
spec/harness drift, not product regressions:
- fixtures.ts: title generation now rides the main model (#83636), firing a
background completion at the mock after every turn — it contains the whole
conversation (trigger keywords included), advancing scripted-turn indices
and tripping hold-for-prompt matchers. Disabled by default in the mock
provider config; specs supplying their own `auxiliary:` section own it.
- chat/interim-messages/session-compression/correction-session-switch/
hidden-history-messages: busy-state and transcript assertions updated to
the current composer aria-labels, interim-message semantics, and
verify-on-stop continuation behavior on main.
- bot-mode-closed-chat-stays-closed/group-to-local-bot-handoff: Bot Chat tab
selectors updated for the Bot Mode rework (tabs keyed by
connection+profile, renamed tab triggers).
- glyph-spinner: assertions made compositor-honest for the CI runner
(steps() keyframes + layer promotion probed via the animation registry
instead of GPU-dependent screenshots).
- sidebar-states/tile-unread-bug: event-driven waits with mock-server
release handles replace wall-clock polls that lost races on loaded
runners.
- warm-resume-jitter/image-attachment-resume: real-session-builder harness
waits for the thread viewport before evaluating; failure path now dumps
per-surface pane state.
Local full-suite run on the CI-equivalent xvfb setup: 62 passed,
11 skipped, 1 flaky-passed (correction-session-switch live-correction spec,
passes on retry). No product code changed.
findElectron() probed exactly one path, and got three things wrong at
once for anyone not on a hoisted POSIX install:
* It looked only under the REPO ROOT. This is an npm workspaces repo and
npm hoists a dependency only when nothing conflicts, so `electron`
installing into apps/desktop/node_modules is an ordinary outcome, not a
broken tree.
* It joined a bare `electron`. On Windows the dist file is
`electron.exe`, so the probe could never match there.
* Its PATH fallback spawned `which`, which is not a command on Windows,
so the fallback failed for a reason unrelated to whether electron is on
PATH.
The three combine into a misleading error: the suite refuses to start
with 'Run "npm install" from the repo root' on a tree that has electron
installed. Reproduced on Windows 11 against this repo, where
apps/desktop/node_modules/electron/dist/electron.exe exists and the old
body throws that message; the reporter on #88036 hit the same thing on
Linux and had to hand-symlink the package before the suite would run.
Resolution now asks the installed `electron` package for its own path
first (its main export IS the absolute executable, resolved from
path.txt and honouring ELECTRON_OVERRIDE_DIST_PATH), then falls back to
explicit dist probes for each root, then to PATH with the platform's
lookup command. The error message lists what was searched.
The rules live in e2e/electron-binary.ts so they can be unit-tested
without importing the Playwright runner, with the platform passed in
rather than read from process.platform: reading it would leave every
Windows rule untested on the Linux CI runner.
Wiring: the vitest `electron` project picks up e2e/**/*.unit.test.ts and
Playwright ignores the same pattern, so helper unit tests run in exactly
one runner and the specs are untouched.
Verified: 5 unit tests pass; mutation-checked one rule at a time
(hardcoding the binary name fails 2, reversing the probe order fails 1,
hardcoding `which` fails 1). tsc -p . and tsc -p tsconfig.e2e.json
clean.
This is the environment blocker called out in #88036, not its rendering
bug, so it is deliberately a subset.
Refs #88036
Session, instance, HUD, quick-entry and pet-overlay windows all open with
show: false and are revealed only by ready-to-show, so the Electron 40 bug
strands them exactly the way it stranded the primary window — and none of
them have the second-launch workaround that made the main-window case
recoverable.
Generalize the controller to any window and wire all six through one
wireWindowReveal helper. Callers pass their own reveal action (showInactive
for the pet overlay, show + focus for the HUD and quick entry) and their own
post-visible work, so whichever path wins runs them exactly once.
Quick entry now reveals the window the call created rather than whatever
`quickEntryWindow` points at when the event lands.
Fresh installs now default to ~91%, but Playwright hit-testing and
visual baselines still assume 100%. Seed zoom-state.json so isolated
E2E profiles don't inherit the product default.
Playwright closes the app with a turn still in flight, so the new quit
confirmation waited on a click nobody was there to make and the E2E
worker died on a 90s teardown timeout.
* test(desktop): e2e test for interim assistant message preservation (#65919)
Adds a Playwright E2E test that reproduces the fix from PR #65919 across
all three layers (agent core → tui_gateway → desktop renderer). The mock
inference server is upgraded with a multi-turn scripted response that
exercises several interleaved patterns:
1. text + tool_call → should produce an interim message
2. text + tool_call → another interim message
3. no text + tool_call → NO interim (no visible text alongside tools)
4. text + tool_call → another interim message
5. final answer (stop) → message.complete, different from all interims
Two describe blocks exercise display.interim_assistant_messages both on
(default) and off:
- ON: all interim texts + the final answer visible in the transcript
- OFF: only the final answer visible, all interim texts wiped
Also fixes a footgun: test:e2e now runs `npm run build` as a pretest
hook so the renderer dist/ is always fresh. Previously, running
`npx playwright test` locally would silently load a stale dist/ that
predated renderer fixes — the python backend ran from source (had the
fix) but the renderer was frozen in an old bundle. CI already built
fresh, so the explicit build step there is removed to avoid duplication.
* test(desktop): e2e sidebar states — background dot, subagent, cross-session
Add sidebar-states.spec.ts with three E2E tests exercising the desktop
sidebar's session dot states driven by real gateway events:
1. Background process dot appears during a terminal(background=true)
call and disappears after auto-dismiss; subagent (delegate_task)
runs concurrently; final answer is visible in the transcript.
2. Background dot remains visible while a subagent runs concurrently
(longer sleep 5 background process so the dot is catchable).
3. Cross-session dot transition: start a turn with a background process,
wait for the turn to complete, open a new session, then verify the
original session's dot transitions from 'background running' to
'finished — unread' when the background process exits.
The mock server gains SIDEBAR_SCRIPT and SIDEBAR_CROSS_SCRIPT trigger
keywords that return tool_calls for terminal(background=true) and
delegate_task — the agent executes these for real (real background
process, real subagent), so the tests assert against genuine gateway
events rather than mocked UI state.
Verified: 3 passed (1.2m) under cage headless wlroots.
* test(desktop): e2e tests for tile-unread bug (tab passes, split fails)
Two scenarios for the tile-unread bug where a session that finishes
while visible on-screen gets the green 'finished unread' dot even
though the user is looking right at it.
The unread check in handleTransition (session-states.ts:174) only
compares against $selectedStoredSessionId and ignores $sessionTiles,
so a session visible in a tile gets marked unread even though it's
on screen.
1. TAB (hidden, PASSES): ⌃-click opens the session as a stacked tab
that is NOT visible on screen. The unread dot IS correct here —
the user isn't looking at it.
2. SPLIT (visible, FAILS): drag the session row to the workspace's
right edge to create a side-by-side split tile. Both sessions are
visible on screen. The unread dot is WRONG — the session is visible
in the split tile, so it should not be marked 'unread'. This test
is RED until the fix lands.
Also adds explicit page.screenshot() calls at key assertion points in
sidebar-states.spec.ts so the trace viewer has full-res captures of the
sidebar dot states during the test.
* test(desktop): cover compression and queued stop lifecycle
Add real desktop E2E coverage for session compression continuation and
queue parking after an explicit Stop. Extend the mock server with a
blocking scripted turn and submitted-prompt assertions.
* test(desktop): cover busy composer submit routing
Replace the invalid queued-stop E2E scenario: plain text redirects a busy
turn rather than entering the queue. Add focused submit-routing coverage for
plain text, slash commands, attachments, explicit Stop, and idle submission.
Keep a live session projection from adding its user turn when the latest
persisted row already represents that same inflight prompt. Add real Electron
coverage for fast and cold resume with idle and background-inference sessions.
Exercise the full Electron, gateway, and mock-provider submit path while
same-chat route query tokens churn during session creation. Assert the mock
provider receives the prompt and its streamed response reaches the transcript.
* nix: add `cage` to devShell
* test(desktop): add pre-filled sessions support
Exports createSandbox, writeMockProviderConfig, writeEnvFile,
buildAppEnv, findElectron, and launchDesktop from fixtures.ts so
specs can compose their own seeded-backend fixtures without duplicating
the sandbox/config/launch logic.
* test(desktop): auto-fail e2e tests on error banner
Adds a shared test fixture (e2e/test.ts) that wraps @playwright/test's
page with an error-banner guard. When any [role="alert"] element
(error notification toast) appears in the DOM during a test, the test
fails with the error message text.
The guard uses:
- A MutationObserver (injected via addInitScript) that watches for
[role="alert"] elements appearing at any point during the test
- A final DOM scan in afterEach for alerts still visible at teardown
- Deduplication so the same error text only fires once
All existing e2e specs updated to import { test, expect } from './test'
instead of '@playwright/test'. No per-spec setup needed — the guard is
auto-installed on every page via the extended fixture.
This catches issues like the "resume failed" error banner that can
appear during session loading — previously the test would pass while
an error toast was silently visible on screen.
* fix(state): parse tool_calls JSON string before re-serializing
_insert_message_rows and append_message both do json.dumps(tool_calls)
to serialize the field for SQLite storage. But when tool_calls arrives
as a JSON string (from import_sessions / export_session, which store it
as TEXT), json.dumps double-encodes it — wrapping the already-serialized
string in quotes and escaping the inner quotes.
When _rows_to_conversation later does json.loads(row['tool_calls']),
the double-encoded string parses back to a plain string (not a list).
_history_to_messages then iterates this string character-by-character,
calling tc.get('function', {}) on each char — 'str' object has no
attribute 'get'.
This was a pre-existing bug (on main), but only triggered by the
import_sessions path (the live agent always passes tool_calls as a
Python list). The e2e error-banner guard caught it via the 'Resume
failed' notification toast.
Fix: in both append_message and _insert_message_rows, parse tool_calls
with json.loads first if it's a string, then re-serialize.
* fix(desktop): exempt boot-failure from error guard
- boot-failure: add allowErrorBanners() beforeEach — these tests
deliberately trigger boot errors, so error toasts are expected
- test.ts: export allowErrorBanners() opt-out + reset flag in afterEach
Reverts the Preparing component changes from b2857110b so the progress bar
turns red (bg-destructive) and the error text shows below it when boot.error
is set, instead of bailing out with an early return null.
The corresponding e2e guard in waitForBootFailure (e2e/fixtures.ts) that
rejected any progress bar in the DOM is dropped — it now waits for the
failure dialog (Retry/Repair/Use local gateway/Connection settings) or the
"Desktop boot failed" toast. The boot-failure.spec.ts header comment is
updated to match.
Verified: tsc clean, vitest boot-failure-reauth (21/21) + boot-failure-overlay
(3/3) pass, npm run build clean, playwright e2e/boot-failure.spec.ts 2/2 pass.
The boot-failure screenshot showed a progress bar because of two bugs:
1. waitForBootFailure matched on "Let's get you setup" (the onboarding
header that mounts from frame 1 during normal boot), so the screenshot
fired at ~86% progress while the Preparing component's progress bar was
still painted.
2. The Preparing component kept rendering the progress bar even after
boot.error was set — it just turned the bar red and appended the error
text below it.
Fixes:
- Preparing bails out (returns null) when boot.error is set, so
BootFailureOverlay (z-1400) owns the screen exclusively.
- applyDesktopBootProgress no longer clobbers a previously-set boot.error
when a late progress event arrives with error: null — failDesktopBoot is
terminal for the boot cycle.
- waitForBootFailure guards against progress bars being visible and matches
on actual failure signals (error toast, Retry/Repair buttons), not the
onboarding header.
- setupDeadBackend now accepts { fakeError: true } which injects
HERMES_DESKTOP_BOOT_FAKE_ERROR to trigger a real boot failure — the
previous dead-provider fixture never actually caused a boot failure
(hermes serve starts fine; the dead endpoint only matters at chat time).
- boot-failure.spec.ts updated to use { fakeError: true }.
Verified: e2e test passes with 0 progress bars in the DOM at screenshot
time (confirmed via DOM inspection), 16/16 vitest tests pass, typecheck
clean.
Screenshots were catching the app mid-boot at ~92% with the onboarding
Preparing progress bar still visible. waitForAppReady checked for the
composer (textarea/contenteditable) with state:'visible', but Playwright
considers an element visible even when a z-1300+ fixed overlay covers it
(non-zero bounding box, not display:none).
Now waits for the composer to be attached, then polls
document.elementFromPoint at viewport center — if the topmost element is
inside a position:fixed inset:0 overlay, the app isn't ready yet.
Adds a full desktop Playwright E2E suite that launches the Electron app
against a mock inference server, exercising the full boot chain:
electron -> hermes serve -> mock provider -> renderer
Includes:
- Mock OpenAI-compatible inference server (mock-server.ts)
- Shared fixtures with sandbox isolation (credentials, HERMES_HOME,
userData, fixed window-state.json for reproducible screenshots)
- Test specs: boot, boot-failure, onboarding, mock-backend-setup, chat,
and packaged-app launch
- Visual regression: expectVisualSnapshot() wraps toHaveScreenshot in
try/catch so diffs are reported without failing the test suite
- CI workflow: xvfb at 1280x1024, baseline cache from main
(--update-snapshots on main, compare on PRs), step summary table with
diff/actual/expected image links, dedicated visual-diffs artifact
- dev:mock script for local fake-provider development
- test:e2e:visual + test:e2e:update-snapshots scripts using cage
- .gitignore: *-snapshots/ (baselines cached in CI, not committed)