e14248ac1e
Port t_3778a491's in-flight stale-claim guard, absent from origin/main. _submit_with_guard adds a job id to _running_job_ids before the future that owns its release exists. Anything that hangs or dies between the add and pool.submit (EAGAIN thread exhaustion on a substrate spike, or a wedged SessionDB.__init__ on a stale sqlite flock) leaks the claim; every later tick short-circuits with 'already running - skipping' silently - no execution row, no last_error, no counter - until the gateway process restarts. This wedged 4 recurring no_agent router/watchdog jobs (verdict- router, wake-scanner, auto-review-router, blocked-task-notifier) for ~1h47m on 2026-08-14 (t_20e23f84), cleared only by manual force-run. - Record claim timestamp + pending-future sentinel in the same critical section as the add; replace sentinel with the owning future after submit. - sweep_stale_inflight() runs every tick (even idle) and force-releases claims older than max(2*interval, 30m floor) with no live future: WARNING cron.inflight.forced_release, get_inflight_guard_stats() counter, JSONL record, and mark_job_run(success=False) so the wedge surfaces as last_error. - Wrap the pre-future init (create_execution/copy_context) so an exception there releases the claim immediately instead of leaking it. - Finite-repeat jobs are released without mark_job_run so a forced release never consumes a one-shot budget. Scheduler-internal only: no provider/model routing, no credentials, no spend, no guardrail weakening, no cron permission widening. Tests: tests/cron/test_inflight_stale_guard.py (18), plus regression tests for the recurring EAGAIN re-dispatch and the create_execution/pool-submit leak paths. Full tests/cron/: 616 passed.