fix(goal): the judge only sees the goal's own background processes, and a pid/session wait barrier expires after 30 min

In a fan-out run the /goal loop parked for 3 h 22 min at the end on a
grandchild's poller (waiting_on_session=proc_21a6fe2369a1, 00:47 -> 04:09)
while the root itself had nothing running. Every one of the root's 7 logged
judge verdicts was WAIT on a child-owned process.

Two causes. gather_background_processes() was called with no task_id from
both goal-loop callers (CLI cli_loops_mixin, gateway run_goals), and the
registry's task_id is the CONTAINER key, which collapses to one value for
every agent in the process, so the judge's process list was every one of
~1,300 subagents' pollers. And a pid/session barrier had no ceiling: once
the judge said WAIT on a session that never exits, nothing resumed judging;
waiting_since was recorded and never read.

Now list_sessions() reports owner_task_id (the RAW spawning id the
registry already keeps for ownership checks), gather_background_processes
takes owner_task_id and both callers pass their own session id (CLI turns
register processes under self.session_id; gateway turns under
turn_ctx.session_id), and is_waiting() ages out a pid/session barrier
after _MAX_BARRIER_WAIT_S (30 min). Timed barriers keep their own deadline.

Tests: only the owning session's running processes are returned when
owner_task_id is given (unfiltered behaviour unchanged); a live barrier
older than the ceiling clears and judging resumes.
This commit is contained in:
Teknium
2026-09-05 00:30:19 -07:00
parent 5ac75e91e2
commit f8b87f5637
5 changed files with 69 additions and 6 deletions
+1
View File
@@ -1776,6 +1776,7 @@ class ProcessRegistry:
"command": s.command[:200],
"cwd": s.cwd,
"pid": s.pid,
"owner_task_id": s.owner_task_id or s.task_id,
"started_at": time.strftime("%Y-%m-%dT%H:%M:%S", time.localtime(s.started_at)),
"uptime_seconds": int(time.time() - s.started_at),
"status": "exited" if s.exited else "running",