d228013832
Greptile's two findings on the original PR were both right. 1. The scaler read test_durations.json from the checkout, but CI ran on a fresh runner where that file never exists (it is gitignored and the slicing-era artifact/merge job that produced it is gone). The feature was inert exactly where the false FLAKY kills happen. tests.yml now restores the most recent main-saved cache before the run (PRs read only) and saves it after a green push to main, mirroring the ci-timings-baseline restore/save pattern already in ci.yaml. 2. _save_durations persisted every file's total subprocess wall, including the ~cap of a timed-out attempt and the retry-summed wall of a FLAKY file. With the scaler that compounds: a hang cached at ~300s earns 900s next run, then ~900s cached earns 2700s, until the job timeout is the only bound. _clean_pass_durations drops failed and FLAKY files from the write so a file's cached duration is always a first-attempt-clean measurement; those files keep their previous known-good entry. Tests trimmed to the salvage bar (<=2 invariants for the scaler plus one for the cache filter) and moved next to the other runner tests under tests/scripts/.