fix(install-e2e): review findings — input hygiene, chart ranking, cost docs

tag-count now reaches the shell via the environment, validated to 1-10
(an apostrophe in the raw interpolation could terminate quoting; above
~14 tags the expansion exceeds GitHub's 256-job matrix limit).

Result-chart cell ranking matches on the leading token: rendered
success/failure cells carry artifact links, so whole-cell indexOf
ranked them -1 and any skip in the map beat a real outcome.

README documents per-run cost, route slice sizes, the tag-count bound,
and a warning against running the GUI drivers outside a disposable VM.
This commit is contained in:
yoniebans
2026-09-02 18:27:06 +02:00
parent 558d76c401
commit a3e7d6a1c7
3 changed files with 18 additions and 3 deletions
+8 -1
View File
@@ -95,9 +95,16 @@ jobs:
sparse-checkout: scripts/sandbox/pick-release-tags.sh
sparse-checkout-cone-mode: false
- id: pick
env:
# Dispatch inputs never touch shell syntax directly: TAG_COUNT
# arrives via the environment, is validated here, and is capped
# below the 256-job matrix limit (14 tags ~= 252 windows entries).
TAG_COUNT: ${{ inputs.tag-count || 2 }}
run: |
set -euo pipefail
tags="$(scripts/sandbox/pick-release-tags.sh --count '${{ inputs.tag-count || 2 }}')"
[[ "$TAG_COUNT" =~ ^[0-9]+$ ]] || { echo "tag-count must be a positive integer, got: $TAG_COUNT" >&2; exit 1; }
(( TAG_COUNT >= 1 && TAG_COUNT <= 10 )) || { echo "tag-count must be 1-10, got: $TAG_COUNT" >&2; exit 1; }
tags="$(scripts/sandbox/pick-release-tags.sh --count "$TAG_COUNT")"
echo "Testing updates from: $tags"
# Annotate each tag with what its own tree supports, so run
# workflows can natively skip surfaces the starting version does
+5 -1
View File
@@ -317,6 +317,10 @@ export function renderMarkdownResults(jobs, tagAnnotations = [], artifactById =
// outcome always beats a skip, and a bad outcome beats a good one.
const RANK = ['skip', 'TODO', 'pre-desktop', '&#x2705;', 'running', 'cancelled', '&#x274C;'];
const SKIPS = ['skip', 'TODO', 'pre-desktop'];
// Rendered success/failure cells carry artifact links after the glyph;
// rank by the leading token or every such cell would rank as unknown (-1)
// and lose to any skip already in the map.
const rankOf = (cell) => RANK.findIndex((t) => cell === t || cell.startsWith(`${t} `));
/** @type {Map<string, Map<string, string>>} */
const rows = new Map();
/** @type {string[]} */
@@ -355,7 +359,7 @@ export function renderMarkdownResults(jobs, tagAnnotations = [], artifactById =
})();
const byTag = /** @type {Map<string, string>} */ (rows.get(combo));
const prev = byTag.get(tag);
if (prev === undefined || RANK.indexOf(cell) > RANK.indexOf(prev)) {
if (prev === undefined || rankOf(cell) > rankOf(prev)) {
byTag.set(tag, cell);
}
}
+5 -1
View File
@@ -65,7 +65,7 @@ A grey leg is normal. There are two causes:
The result chart on the run summary shows each leg as passed, failed, or skipped.
## Triggers
## Triggers and cost
The matrix does not run on pull requests. One leg installs real toolchains and takes more than 10 minutes. The triggers are:
@@ -77,6 +77,10 @@ The matrix does not run on pull requests. One leg installs real toolchains and t
gh workflow run install-e2e.yml --ref <branch> -f route=both -f tag-count=2
```
Cost per run, so nobody is surprised: the full board at the default 2 sampled tags is up to ~76 legs. A typical green leg finishes in 7-15 minutes; every leg is capped at 60. Route slices for cheaper reads: `update` (linux only, ~8 legs), `windows-desktop` (~18/tag), `macos-desktop` (~15/tag). `tag-count` is validated to 1-10; above ~14 tags the expansion would exceed GitHub's 256-job matrix limit.
Running the drivers locally: don't, except in a disposable VM. The windows driver kills every process named Hermes during teardown and the macos driver operates on `/Applications/Hermes.app`; on a machine with a real Hermes install they will interfere with it.
## Artifacts
Each leg uploads its logs as an artifact. Every leg also records the screen for its whole run: the composite action `.github/actions/e2e-screen-record` installs ffmpeg, records with the OS's capture backend (x11grab on linux, gdigrab on windows, avfoundation on macos), and fails the leg if the recording is missing or has zero frames. Linux runners have no display, so the action starts `Xvfb :99` first and exports `DISPLAY` for every later step — the app under test and the recorder share that display. The windows GUI leg also uploads screenshots and the update result file. Get them with `gh run download <run-id>`.