The first hardening (deterministic worker start) still went red on its
own PR's CI: the 0.1s deadline can ALSO expire between worker start and
handle_function_call — argument parsing, hooks, and cold imports run in
that window, and the timeout interrupt then returns the middleware
without dispatching (first_started unset; observed 0.15s tool error vs
0.1s budget in the slice-4 log).
Deadline raised to a 1.0s floor: still fast, ~7x the worst observed
preamble, and the submit() sync from the first commit keeps the
countdown anchored to worker start. The clarify human-wait test's sleep
rises to 1.3s so it still outlasts the deadline (the relationship the
test exists to pin). Sabotage v2 injects BOTH races (0.3s thread start
+ 0.5s preamble): hardened passes, un-hardened reproduces the exact CI
failure.
test_sequential_tool_timeout_emits_result_and_continues failed twice in
two days on unrelated computer_use PRs (slice 4, xdist): assert
first_started.is_set() -> False. Root cause: the sequential timeout path
computes deadline = now + timeout_s right after executor.submit(); with
the test's 0.1s deadline, a loaded CI worker can take longer than the
whole deadline just to START the pool thread, so the future is cancelled
before the tool ever dispatches.
Fix: an autouse fixture subclasses DaemonThreadPoolExecutor so submit()
blocks (bounded 10s) until the worker callable has begun — the deadline
now races the tool, not the thread scheduler, which is what these tests
mean to pin. Timeouts stay tight (0.05s), so the suite stays fast.
Sabotage-proven: a 0.3s injected thread-start delay reproduces the exact
CI failure without the fixture and passes with it.
Clarify waits on a human for up to 3600s or unlimited. The generic sequential timeout was aborting that wait at 420s and leaving the prompt and worker active.