Two defects in the same path, both measured on the 1,393-agent refactor run.
1. A nested orchestrator (depth > 0) runs delegate_task synchronously by
design: it needs its workers' results inside its own turn. But the
sequential tool runner put every tool call under the generic 420 s
deadline, and delegate_task was not exempt, so every batch longer than
seven minutes returned "timed out after 420.0s" while the children kept
running as orphans. 332 such timeouts in 234 orchestrator sessions; only
89 nested delegate_task calls in the whole run ever returned a real result.
The orchestrators then spent 388 h of wall time polling: 1,526 reads of
the live transcript files, 551 list actions, 242 h of explicit sleep,
about $4k of API turns. delegate_task is now exempt from the sequential
deadline (the batch owns its liveness: per-child heartbeats, the stale
monitor, delegation.child_timeout_seconds).
Live A/B, depth-1 orchestrator dispatching a 75 s leaf with the deadline
set to 40 s (glm-5.3-flash via Nous): main -> "Error executing tool
'delegate_task': timed out after 40.0s", leaf result lost; branch ->
orchestrator blocked 161 s and returned the leaf's LEAF_DONE_MARKER.
2. _parent_summary_char_budget computed the parent's context headroom from
session_prompt_tokens, which is the running SUM of prompt tokens over
every API call in the session. After a few hundred calls it exceeds any
window, headroom goes negative, and every child summary collapses to the
2,000-char floor with the full text spilled to disk. All 1,393 child
summaries in the run were truncated this way; the orchestrator planned
from stubs. The budget now reads the last call's prompt_tokens from
_last_turn_usage.
Tests: delegate_task is in the exempt set and the set is narrow; budget for
a long-lived parent equals the budget for a fresh parent with the same
current prompt, and exceeds the floor.