Commit Graph

3 Commits

Author SHA1 Message Date
devops 122bfad535 fix(cron): persisted-state recovery re-arms recurring job stuck in stale last_status=error (t_8b5480b3)
The 2026-08-14 incident (t_20e23f84): 4 recurring no_agent interval jobs
EAGAIN-failed at 12:50 and recorded ZERO executions for ~1h47m, surviving a
gateway restart, cleared only by operator `cron resume` / force-run. The
in-memory stale-claim sweep (t_3778a491, already on origin/main) heals a
leaked `_running_job_ids` claim in-process, but a recurring job whose
PERSISTED state shows last_status=error and whose next_run_at was re-armed
into the future by mark_job_run is invisible to that sweep: it is not in the
running set and not due, so it just sits — the restart-surviving half.

cron/jobs.py::_get_due_jobs_locked now re-arms such a recurring job to
next_run_at=now when all hold: persisted last_status==error, last_run_at older
than cadence+grace (so it is a real wedge, not a normal transient-error retry),
next_run_at in the future, and not running in this process. The scheduler then
re-dispatches it on the next tick without force-run/resume. Logs
cron.persisted_error.recovered, bumps a probe-visible counter, appends a JSONL
row. Within-cadence errors are never force-re-armed.

Tests: tests/cron/test_recurring_persisted_error_recovery.py (clean behavioral
RED on unfixed main / GREEN here; 2 consecutive auto-fires; within-cadence not
re-armed). Full tests/cron/: 713 passed, 1 skipped.
2026-08-17 16:55:00 +05:30
Evgenii acaafcc6bb fix(cron): make immediate execution race-safe
- claim_job_for_fire returns the atomically claimed snapshot with a unique
  fire owner; heartbeat_fire_claim renews the lease; mark_job_run fences
  terminal writes by expected_fire_owner so a stale worker cannot record
  over a replacement claim.
- run_one_job heartbeats the fire claim and forwards a combined cancel
  event (ownership loss OR external cancel) into run_job; the agent path
  is interrupted cooperatively and script-based jobs (no_agent + pre-run
  scripts) are hard-stopped with a process-tree kill (POSIX killpg
  SIGTERM then SIGKILL for surviving group members; Windows
  taskkill /T /F), with a bounded pipe drain so a SIGTERM-ignoring
  descendant cannot wedge the worker on communicate() EOF.
- Shutdown interruption is scoped to the exact execution token instead of
  the bare job ID, so a replacement run of the same job never consumes a
  stale interrupted flag.
- fire_claim_fence serializes save/deliver side effects per profile+job
  with a cross-process flock; remove_job prunes the fence-lock entry.
- Preserves upstream BaseException terminal recording (#73973),
  completed one-shot retention (#80624), blocked_config preflight
  (T1-26), and the advance_next_runs batch on top of current main.
2026-08-14 20:46:50 -07:00
devops e14248ac1e fix(cron): self-heal leaked in-flight claim so a wedged recurring job re-dispatches (t_8b5480b3)
Port t_3778a491's in-flight stale-claim guard, absent from origin/main.

_submit_with_guard adds a job id to _running_job_ids before the future
that owns its release exists. Anything that hangs or dies between the
add and pool.submit (EAGAIN thread exhaustion on a substrate spike, or a
wedged SessionDB.__init__ on a stale sqlite flock) leaks the claim; every
later tick short-circuits with 'already running - skipping' silently - no
execution row, no last_error, no counter - until the gateway process
restarts. This wedged 4 recurring no_agent router/watchdog jobs (verdict-
router, wake-scanner, auto-review-router, blocked-task-notifier) for ~1h47m
on 2026-08-14 (t_20e23f84), cleared only by manual force-run.

- Record claim timestamp + pending-future sentinel in the same critical
  section as the add; replace sentinel with the owning future after submit.
- sweep_stale_inflight() runs every tick (even idle) and force-releases
  claims older than max(2*interval, 30m floor) with no live future: WARNING
  cron.inflight.forced_release, get_inflight_guard_stats() counter, JSONL
  record, and mark_job_run(success=False) so the wedge surfaces as last_error.
- Wrap the pre-future init (create_execution/copy_context) so an exception
  there releases the claim immediately instead of leaking it.
- Finite-repeat jobs are released without mark_job_run so a forced release
  never consumes a one-shot budget.

Scheduler-internal only: no provider/model routing, no credentials, no
spend, no guardrail weakening, no cron permission widening.

Tests: tests/cron/test_inflight_stale_guard.py (18), plus regression tests
for the recurring EAGAIN re-dispatch and the create_execution/pool-submit
leak paths. Full tests/cron/: 616 passed.
2026-08-15 02:23:56 +05:30