fix(kanban): live-claim guard keys on a live worker process, not on any claim

The first cut refused every claim-less complete of a running+claimed card,
which also refused the flows that have no worker to protect: a library or
CLI claim that never spawned a worker, and a worker whose process is gone
(12 sibling tests exercise exactly that shape). The guard now fires only
when tasks.worker_pid names a process that is still alive under its spawn
fingerprint (_worker_alive), which is the run the issue asked us to keep
open. Test updated to stand in as the live worker via _set_worker_pid.
This commit is contained in:
teknium1
2026-09-15 13:47:34 -07:00
committed by Teknium
parent 0959224313
commit 72916de360
2 changed files with 26 additions and 4 deletions
+11 -3
View File
@@ -2653,15 +2653,23 @@ def complete_task(
return False
if acceptance is not None and not record_acceptance(conn, task_id, acceptance):
return False
trow = conn.execute("SELECT status, claim_lock FROM tasks WHERE id = ?", (task_id,)).fetchone()
trow = conn.execute(
"SELECT status, claim_lock, worker_pid, worker_started_at FROM tasks WHERE id = ?",
(task_id,),
).fetchone()
prior_status = trow["status"] if trow else None
# Refuse to close a live worker's run without proof of ownership
# (expected_run_id) or an explicit human override (force=True).
# Refuse to close a LIVE worker's run without proof of ownership
# (expected_run_id) or an explicit human override (force=True). "Live"
# means the spawned worker process still exists: a claim whose worker
# is gone (or a library claim that never spawned one) has no run to
# protect, so manual completion keeps working there.
if (
expected_run_id is None
and not force
and prior_status == "running"
and trow["claim_lock"] is not None
and trow["worker_pid"]
and _worker_alive(trow["worker_pid"], trow["worker_started_at"])
):
raise LiveClaimError(task_id)
sql = """