481156139d
When an external scheduler (Chronos on hosted deployments) cannot deliver a fire — dead loopback hop at fire time, retry budget exhausted — the job's next_run_at stays parked in the past and nothing ever runs it: external providers have no local tick loop, so the day is silently lost even if the gateway heals minutes later (4 consecutive nightly misses in the field). fire_overdue_jobs() in cron/scheduler_provider.py, called from the gateway housekeeping loop every 5 minutes: - No-op for the built-in ticker (its tick loop already self-heals past-due jobs) and when cron.misfire_grace_minutes <= 0. - Waits out a grace window (default 10 min) so the external scheduler's own retry backoff gets first right to deliver. - Claims via the provider's claim_fire (store CAS — a concurrent late external retry is de-duplicated) and runs fire_claimed in a daemon thread, mirroring the webhook admission pattern, so housekeeping never blocks for the length of an agent run. Provider re-arm logic (Chronos NAS one-shots) runs exactly as for a normal fire. Docs: cron.md section + cron.misfire_grace_minutes reference.