Files
hermes-agent/cron
Teknium 481156139d feat: misfire catch-up for external cron providers
When an external scheduler (Chronos on hosted deployments) cannot
deliver a fire — dead loopback hop at fire time, retry budget exhausted
— the job's next_run_at stays parked in the past and nothing ever runs
it: external providers have no local tick loop, so the day is silently
lost even if the gateway heals minutes later (4 consecutive nightly
misses in the field).

fire_overdue_jobs() in cron/scheduler_provider.py, called from the
gateway housekeeping loop every 5 minutes:

- No-op for the built-in ticker (its tick loop already self-heals
  past-due jobs) and when cron.misfire_grace_minutes <= 0.
- Waits out a grace window (default 10 min) so the external scheduler's
  own retry backoff gets first right to deliver.
- Claims via the provider's claim_fire (store CAS — a concurrent late
  external retry is de-duplicated) and runs fire_claimed in a daemon
  thread, mirroring the webhook admission pattern, so housekeeping
  never blocks for the length of an agent run. Provider re-arm logic
  (Chronos NAS one-shots) runs exactly as for a normal fire.

Docs: cron.md section + cron.misfire_grace_minutes reference.
2026-08-17 11:42:25 -07:00
..