A single urlopen(timeout=5) TimeoutError from the PS runspace listener
failed the test on a loaded runner (run 32440286339) even though the
listener recovered moments later — a transient stall is not the hang
this test guards. /progress sampling now retries until a deadline
(only a persistently unresponsive listener fails), the self-test hold
grows 10s -> 30s so retry time cannot push sampling past the held
stage, and the exit wait gets matching headroom.
test_progress_advances_while_the_orchestrator_blocks raced its subject on
both edges within one hour of PR CI (#90358):
- Run 1: sampled right after the shim URL printed, before the orchestrator
published its stage — caught the page boot default
('Hermes will open once done.' != 'Testing quiet update').
- Run 2 (rerun): with HOLD=4s on a slow runner, the second sample slid past
the hold and caught the cleared terminal state ('' != 'Testing quiet
update').
Fix: wait (<=10s) for the published stage to actually land before starting
the 1.5s stability window, and raise the hold to 10s so both samples land
inside it. Same assertions, same contract — just anchored to the event the
test is about instead of wall-clock luck.
The self-test grew a branch that spawns a Python child so pytest could prove
progress advances during one. It doesn't need to: /progress is answered from
its own runspace, so the existing hold already blocks the main thread, and
the spawn only exercised Invoke-HermesStep, which nothing here changes.
The Windows test now asserts the invariant instead of the self-test's stage
string, and the posix half -- previously untested, and the half that broke --
gets real coverage: serve-ui.py's wire shape, and posix.sh driven end to end
with a stub `hermes` that reports which stage was on screen while it ran.