4822156923
Since the managed-cron redesign (#84339, v2026.8.13) the dashboard fire webhook forwards fires to the gateway process and returns 503 when it is unreachable so NAS/QStash retries. Correct for transient windows — but an operator-STOPPED gateway can never be fixed by retrying: every fire on every job burns the full scheduler retry budget, NAS converts each 503 to a retryable 502, and the resulting storms page on-call for a non-incident (OOF-266 and its five duplicate tickets; +93% relay callback failures as the fleet adopted v2026.8.13). Split the unreachable path by durable operator intent: - desired_state == "stopped" (written only by the s6 lifecycle commands; the same intent signal container-boot reconciliation trusts) -> drop the fire with 200 + a structured log line, mirroring NAS's own instance_stopped drop. Jobs are not lost: the Chronos provider reconciles and re-arms every job on the next gateway start. - Anything else (crash loop, scale-to-zero wake, restart, legacy state file without desired_state) -> keep the retryable 503, now stamped with Retry-After: 60 so a scheduler that honors it spaces retries past the wake/restart window instead of exhausting them inside it. The gateway's own pass-through 503s (draining) get the same hint. The intent check fails open (any parse/resolution error -> retryable path) and is only consulted when the gateway is actually unreachable, so a stale state file can never shadow a live gateway.