10f99bc15e
Every Linux lane that does real work ran on a 4-core `ubuntu-latest`. The Python suite and the JS checks were split into many small jobs to make that size usable. Each split job repeated the full setup. In most of the JS jobs the repeated setup cost more than the work. The work lanes move to larger runners. Then the splits that existed only to make small runners usable go away. Python tests: 12 slices become 1 job on a 96-core runner. Slicing cost a matrix job, a duration cache, a per-slice artifact and a merge job. 96 cores clear the floor that the slowest single test file sets, which is about 82s. A second slice divides work that is already at that floor, and adds a second setup. Duration data from run 32522943054 gives the numbers behind this: 3178 files, 11645s in series. The worker count is explicit, because `run_tests.sh` defaults to twice the core count. A later commit sets it from a measurement on this hardware. JS checks: 14 jobs become 1. The matrix paid about 371s of repeated setup to spread about 612s of work. One larger runner installs one time. The three UI shard scripts and `run-ui-shard.mjs` are therefore removed, because the unsharded `test:ui` covers the same tests. The unit of parallel work inside that job is a CHECK, and not a workspace. apps/desktop is most of the payload, and its own `check` is a serial && chain. A spread across workspaces alone therefore leaves that chain as the long pole. A package that declares `check:*` sub-scripts gives one unit for each sub-script. That is the same selection rule the matrix used. The loop lives in `.github/scripts/run-workspace-checks.mjs`, so the same sequence runs on a laptop. It runs 11 units together, buffers the output of each one, and fails at the end with the full list. Children that share one stdout interleave their lines and make a failure hard to read. `npm run --ws check` stops at the first workspace that fails. `check:test:plugins` joins the desktop `check` script. The matrix prefers `check:*` sub-scripts over the plain `check` script, so `check:test:plugins` ran only as its own leg. Without this change the merge drops that suite and the job stays green. node_modules is cached on the lockfile, and `npm ci` is skipped on an exact hit. The `cache: npm` option of `setup-node` caches only the ~/.npm tarball cache, which leaves the extract and the postinstalls to pay again. The arm64 image build stays on a native arm64 runner. A build of linux/arm64 on an x64 host uses emulation. The docker test lane caps its workers at the core count. Each of those tests drives a container, so the docker daemon sets the limit and not the processor. `.github/actionlint.yaml` declares the runner labels. actionlint knows the GitHub-hosted labels only, and an undeclared label reads as an error that hides the real findings. The `detect` job checks out one file through a sparse checkout, and its timeout drops to 1 minute. It reads `scripts/ci/classify_changes.py` and nothing else. Verification: - actionlint reports 9 findings across all workflows. An unmodified HEAD with the same config reports the same 9. This change adds none. - A wrong label still fails. actionlint reports `ubuntu-latest-32-cor` and `ubuntu-latest-32-arm-cores`. - Every changed workflow parses, and `name` parses as a string. - A replay of the `save-durations` merge step against a three-artifact layout returns all 3178 entries. - An expansion of the npm script graph gives the same leaf commands for the parallel units and for a plain `npm run check`, in both directions. Against the 13-leg matrix the count is 13 to 11, and the whole difference is the three UI shards that collapse into one unsharded `check:test:ui`. - `--list` reports the 11 units, and a full local run completes and reports the time of each unit. - The runner labels cannot be verified here. The first real run is the test.
285 lines
12 KiB
YAML
285 lines
12 KiB
YAML
name: E2E Desktop
|
|
|
|
on:
|
|
workflow_call:
|
|
outputs:
|
|
review_status:
|
|
description: Screenshot and visual-diff status for the CI review comment.
|
|
value: ${{ jobs.e2e.outputs.review_status }}
|
|
|
|
permissions:
|
|
contents: read
|
|
|
|
concurrency:
|
|
group: e2e-desktop-${{ github.ref }}
|
|
cancel-in-progress: true
|
|
|
|
jobs:
|
|
e2e:
|
|
name: Playwright E2E (Linux)
|
|
# This job builds the renderer and the electron bundle, then drives a real
|
|
# Electron app under xvfb. vite, tsc and the Playwright workers all scale
|
|
# with the core count.
|
|
runs-on: ubuntu-latest-32-core
|
|
timeout-minutes: 20
|
|
outputs:
|
|
review_status: ${{ steps.review-status.outputs.review_status }}
|
|
steps:
|
|
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
|
|
# ── System deps for Electron on headless Ubuntu ───────────────────
|
|
# Electron needs GTK, NSS,atk, etc. even under xvfb. Playwright's
|
|
# install-deps covers browsers; for Electron we install the apt
|
|
# packages directly.
|
|
- name: Install system dependencies for Electron
|
|
run: |
|
|
sudo apt-get update -qq
|
|
sudo apt-get install -y -qq \
|
|
xvfb \
|
|
libgtk-3-0 libnotify4 libnss3 libxss1 libxtst6 \
|
|
xdg-utils libatspi2.0-0 libdrm2 libgbm1 libasound2t64
|
|
|
|
# ── Node ───────────────────────────────────────────────────────────
|
|
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
|
|
with:
|
|
node-version: 26
|
|
cache: npm
|
|
|
|
- name: grab npm 12
|
|
run: |
|
|
npm i -g npm@12
|
|
|
|
# Full npm ci (not --ignore-scripts): electron's postinstall
|
|
# downloads the binary we launch, and node-pty's native build is
|
|
# needed for the terminal pane.
|
|
- uses: ./.github/actions/retry
|
|
with:
|
|
command: npm ci
|
|
|
|
# ── Python (for the hermes serve backend) ──────────────────────────
|
|
- name: Install uv
|
|
uses: astral-sh/setup-uv@fac544c07dec837d0ccb6301d7b5580bf5edae39 # 8.2.0
|
|
with:
|
|
# Pin the uv version: unpinned, setup-uv resolves "latest" by
|
|
# fetching a manifest from raw.githubusercontent.com on EVERY job —
|
|
# a transient fetch failure fails the whole job (2026-07-28 slice-5
|
|
# incident). Pinned, the binary downloads directly; no manifest hop.
|
|
version: '0.9.28'
|
|
enable-cache: true
|
|
cache-dependency-glob: |
|
|
pyproject.toml
|
|
uv.lock
|
|
|
|
- name: Set up Python 3.11
|
|
uses: ./.github/actions/retry
|
|
with:
|
|
command: uv python install 3.11
|
|
|
|
- name: Install Python dependencies
|
|
uses: ./.github/actions/retry
|
|
with:
|
|
command: uv sync --locked --python 3.11 --extra all --extra dev
|
|
|
|
# ── Build desktop app ─────────────────────────────────────────────
|
|
# The Playwright step below runs `npm run build` before testing so
|
|
# dist/ is always fresh — no separate build step needed here.
|
|
|
|
# ── Restore visual baseline screenshots from main ──────────────────
|
|
# Baselines are generated on main (via --update-snapshots) and cached.
|
|
# On PRs, we restore them so toHaveScreenshot has something to compare
|
|
# against. The cache key is keyed on the desktop source files so a
|
|
# UI change naturally invalidates it — but we fall back to the main
|
|
# cache to avoid cold starts on unrelated PRs.
|
|
- name: Restore visual baseline screenshots
|
|
id: restore-baselines
|
|
uses: actions/cache@0400d5f644dc74513175e3cd8d07132dd4860809 # v4.2.4
|
|
with:
|
|
path: apps/desktop/e2e/*-snapshots
|
|
key: visual-baselines-${{ github.ref_name }}
|
|
restore-keys: |
|
|
visual-baselines-main
|
|
|
|
# ── Run Playwright E2E under xvfb ─────────────────────────────────
|
|
# xvfb runs at a fixed 1280x1024 screen so the 1220x800 Electron
|
|
# window always has a consistent viewport for screenshot comparison.
|
|
# On main, we run with --update-snapshots to generate baselines.
|
|
# `npm run test:e2e` builds dist/ as a pretest hook so the renderer
|
|
# is always fresh — no separate build step needed.
|
|
- name: Run Playwright E2E tests
|
|
working-directory: apps/desktop
|
|
run: |
|
|
if [ "${{ github.ref_name }}" = "main" ]; then
|
|
echo "On main — generating/updating baseline screenshots"
|
|
npm run build && xvfb-run -a --server-args="-screen 0 1280x1024x24" \
|
|
npx playwright test --reporter=list --update-snapshots
|
|
else
|
|
echo "On PR — comparing against cached baselines"
|
|
npm run build && xvfb-run -a --server-args="-screen 0 1280x1024x24" \
|
|
npx playwright test --reporter=list
|
|
fi
|
|
env:
|
|
CI: 'true'
|
|
# Ensure no real API keys leak into the test env.
|
|
OPENROUTER_API_KEY: ''
|
|
OPENAI_API_KEY: ''
|
|
NOUS_API_KEY: ''
|
|
|
|
# ── Save updated baselines to cache (main only) ───────────────────
|
|
- name: Save updated baselines to cache
|
|
if: github.ref_name == 'main' && always()
|
|
uses: actions/cache/save@0400d5f644dc74513175e3cd8d07132dd4860809 # v4.2.4
|
|
with:
|
|
path: apps/desktop/e2e/*-snapshots
|
|
key: visual-baselines-main
|
|
|
|
# ── Upload Playwright report (HTML + traces) ──────────────────────
|
|
- name: Upload Playwright report
|
|
id: upload-report
|
|
if: always()
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
|
with:
|
|
name: playwright-report-${{ github.sha }}
|
|
path: apps/desktop/playwright-report
|
|
retention-days: 14
|
|
overwrite: true
|
|
|
|
# ── Upload test results (screenshots, traces, diffs) ───────────────
|
|
- name: Upload test results
|
|
id: upload-results
|
|
if: always()
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
|
with:
|
|
name: playwright-test-results-${{ github.sha }}
|
|
path: apps/desktop/test-results
|
|
retention-days: 14
|
|
overwrite: true
|
|
|
|
# ── Upload just the visual diffs (small, fast to review) ──────────
|
|
- name: Upload visual diffs
|
|
id: upload-diffs
|
|
if: always() && github.ref_name != 'main'
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
|
with:
|
|
name: visual-diffs-${{ github.sha }}
|
|
path: |
|
|
apps/desktop/test-results/**/*-diff.png
|
|
apps/desktop/test-results/**/*-actual.png
|
|
apps/desktop/test-results/**/*-expected.png
|
|
retention-days: 14
|
|
overwrite: true
|
|
if-no-files-found: ignore
|
|
|
|
- name: Build screenshot review status
|
|
id: review-status
|
|
if: always()
|
|
working-directory: apps/desktop
|
|
env:
|
|
RESULTS_URL: ${{ steps.upload-results.outputs.artifact-url }}
|
|
run: |
|
|
python3 ../../scripts/ci/e2e_screenshot_status.py \
|
|
--results-dir test-results \
|
|
--manifest-output /tmp/e2e-screenshot-manifest.json \
|
|
--evidence-dir /tmp/e2e-evidence \
|
|
--artifact-url "$RESULTS_URL" \
|
|
--output /tmp/e2e-review-status.json
|
|
{
|
|
echo 'review_status<<__E2E_REVIEW_STATUS__'
|
|
cat /tmp/e2e-review-status.json
|
|
echo '__E2E_REVIEW_STATUS__'
|
|
} >> "$GITHUB_OUTPUT"
|
|
cp /tmp/e2e-review-status.json review-status.json
|
|
|
|
- name: Upload review status artifact
|
|
if: always() && steps.review-status.outcome != 'skipped'
|
|
continue-on-error: true
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
|
with:
|
|
name: review-status-e2e-desktop
|
|
path: apps/desktop/review-status.json
|
|
retention-days: 1
|
|
overwrite: true
|
|
if-no-files-found: ignore
|
|
|
|
# The trusted workflow_run publisher consumes only this flat, bounded
|
|
# artifact. It turns selected images into GitHub attachment URLs; it
|
|
# never checks out or runs this PR's code.
|
|
- name: Upload inline E2E evidence
|
|
if: always() && github.ref_name != 'main'
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
|
with:
|
|
name: e2e-evidence-${{ github.sha }}
|
|
path: /tmp/e2e-evidence
|
|
retention-days: 14
|
|
overwrite: true
|
|
if-no-files-found: error
|
|
|
|
# ── Generate step summary with visual diff info ───────────────────
|
|
# Parse the JSON report + scan for diff images, then post a summary
|
|
# to the GitHub Actions step output so reviewers can see what changed
|
|
# without downloading artifacts. Runs AFTER uploads so it can link
|
|
# the artifact download URLs from their step outputs.
|
|
- name: Generate visual diff summary
|
|
if: always()
|
|
working-directory: apps/desktop
|
|
env:
|
|
REPORT_URL: ${{ steps.upload-report.outputs.artifact-url }}
|
|
RESULTS_URL: ${{ steps.upload-results.outputs.artifact-url }}
|
|
DIFFS_URL: ${{ steps.upload-diffs.outputs.artifact-url }}
|
|
run: |
|
|
{
|
|
echo "## Desktop E2E — Visual Diff Report"
|
|
echo ""
|
|
|
|
# Count diff images (playwright writes *-diff.png on mismatch)
|
|
DIFF_COUNT=$(find test-results -name '*-diff.png' 2>/dev/null | wc -l)
|
|
ACTUAL_COUNT=$(find test-results -name '*-actual.png' 2>/dev/null | wc -l)
|
|
|
|
if [ "$DIFF_COUNT" -eq 0 ]; then
|
|
echo "✅ All $ACTUAL_COUNT screenshot(s) matched their baselines (or no baselines existed yet)."
|
|
else
|
|
echo "📸 **$DIFF_COUNT of $ACTUAL_COUNT screenshot(s) differ from baseline:**"
|
|
echo ""
|
|
echo "| Test | Diff | Actual | Expected |"
|
|
echo "|------|------|--------|----------|"
|
|
|
|
# List each diff image with a link to the artifact
|
|
for diff in $(find test-results -name '*-diff.png' 2>/dev/null | sort); do
|
|
base=${diff%-diff.png}
|
|
test_name=$(basename "$base")
|
|
echo "| $test_name | [diff]($diff) | [actual](${base}-actual.png) | [expected](${base}-expected.png) |"
|
|
done
|
|
fi
|
|
|
|
echo ""
|
|
echo "📥 **Artifacts:**"
|
|
echo ""
|
|
if [ -n "$RESULTS_URL" ]; then
|
|
echo "- [playwright-test-results]($RESULTS_URL) — all screenshots (actual + expected + diff) + traces"
|
|
fi
|
|
if [ -n "$REPORT_URL" ]; then
|
|
echo "- [playwright-report]($REPORT_URL) — interactive HTML report"
|
|
fi
|
|
if [ -n "$DIFFS_URL" ]; then
|
|
echo "- [visual-diffs]($DIFFS_URL) — just the diffed screenshots (small, fast to review)"
|
|
fi
|
|
echo ""
|
|
echo "**To update baselines:** merge to main (baselines auto-update on main runs) or run \`npx playwright test --update-snapshots\` locally."
|
|
|
|
# Also parse the JSON report for pass/fail counts
|
|
if [ -f playwright-report/results.json ]; then
|
|
echo ""
|
|
echo "### Test Results"
|
|
echo ""
|
|
node -e "
|
|
const r = require('./playwright-report/results.json');
|
|
const stats = r.stats || {};
|
|
console.log('| Status | Count |');
|
|
console.log('|--------|-------|');
|
|
console.log('| ✅ Passed | ' + (stats.expected || 0) + ' |');
|
|
console.log('| ❌ Failed | ' + (stats.unexpected || 0) + ' |');
|
|
console.log('| ⏭️ Skipped | ' + (stats.skipped || 0) + ' |');
|
|
console.log('| 🔄 Flaky | ' + (stats.flaky || 0) + ' |');
|
|
" 2>/dev/null || true
|
|
fi
|
|
} >> "$GITHUB_STEP_SUMMARY"
|