fix(ci): feed the timeout scaler only healthy durations, and wire the cache in CI

Greptile's two findings on the original PR were both right.

1. The scaler read test_durations.json from the checkout, but CI ran on
   a fresh runner where that file never exists (it is gitignored and the
   slicing-era artifact/merge job that produced it is gone). The feature
   was inert exactly where the false FLAKY kills happen. tests.yml now
   restores the most recent main-saved cache before the run (PRs read
   only) and saves it after a green push to main, mirroring the
   ci-timings-baseline restore/save pattern already in ci.yaml.

2. _save_durations persisted every file's total subprocess wall,
   including the ~cap of a timed-out attempt and the retry-summed wall
   of a FLAKY file. With the scaler that compounds: a hang cached at
   ~300s earns 900s next run, then ~900s cached earns 2700s, until the
   job timeout is the only bound. _clean_pass_durations drops failed and
   FLAKY files from the write so a file's cached duration is always a
   first-attempt-clean measurement; those files keep their previous
   known-good entry.

Tests trimmed to the salvage bar (<=2 invariants for the scaler plus one
for the cache filter) and moved next to the other runner tests under
tests/scripts/.
This commit is contained in:
teknium1
2026-09-14 18:35:56 -07:00
committed by Teknium
parent 9a69785790
commit d228013832
4 changed files with 118 additions and 61 deletions
+25
View File
@@ -87,6 +87,22 @@ jobs:
# re-download, keeping the persisted cache small and fast to restore.
run: uv cache prune --ci
- name: Restore per-file duration cache
# scripts/run_tests_parallel.py raises a file's timeout to
# 3x its last healthy duration (_effective_file_timeout) so a
# known-slow file dilated by load is not SIGKILL'd at the flat cap
# and laundered into a FLAKY retry. The scaler reads
# test_durations.json from the checkout; without this restore the
# file is absent on a fresh runner and the scaler is inert.
# Exact key never matches (run_id differs); restore-keys picks the
# most recent cache saved by a main push. PRs read, only main writes.
uses: actions/cache/restore@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5.0.5
with:
path: test_durations.json
key: test-durations-never-exact
restore-keys: |
test-durations-
- name: Run tests
# Per-file isolation via scripts/run_tests.sh: each test file runs
# in its own freshly-spawned `python -m pytest <file>` subprocess
@@ -125,6 +141,15 @@ jobs:
OPENAI_API_KEY: ""
NOUS_API_KEY: ""
- name: Save per-file duration cache (main only)
# Only green first-attempt durations are written by the runner, so
# a hang on main cannot ratchet its own bound upward.
if: github.event_name == 'push' && github.ref == 'refs/heads/main' && hashFiles('test_durations.json') != ''
uses: actions/cache/save@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5.0.5
with:
path: test_durations.json
key: test-durations-${{ github.run_id }}
e2e:
runs-on: ubuntu-latest
timeout-minutes: 15