9de5460c12
* feat(cron): durable failure incidents with signature dedup and ack Introduce a durable cron incident store (cron_incidents in the shared cron/executions.db) that groups "same job + same error signature" across runs, so a known recurring failure stops re-pinging the operator every run once it has been acknowledged. - cron/incidents.py: lazily-created incident table (detected -> alerted -> reviewed -> closed lifecycle; closed is per-signature terminal), sha256 signature dedup over job_id + normalized error, redacted/truncated error storage, failure-type classification, and ack/list/get/count helpers. - cron/scheduler.py: record an incident on the failure delivery path and suppress the per-run failure ping when the exact signature is acked (both the normal failure path and the processing-raised retry path). Best-effort: an incident-store error never breaks the cron run or delivery. Streak nudge, alert-once markers, and delivery-error behavior are untouched. - hermes_cli: add `hermes cron incidents [--state ...]` and `hermes cron incidents ack <id>`. - tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction, classification, lazy-schema, scheduler gating, and CLI coverage. Non-goals deferred to later slices: Discord buttons/review view, HMAC action tokens, owner-agent review launch, approval-gated fixes, incident playbooks. * refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome Follow-ups on top of the salvaged #94692: - Drop the dead 'reviewed' state and the SQLite CHECK (state validity lives in INCIDENT_STATES so future slices can add states without a table rebuild); lifecycle is detected -> alerted -> closed. - Actually mark incidents 'alerted' after a failure ping reaches delivery, on both the normal and exception delivery paths. - Record ack-suppressed runs with a distinct 'suppressed_acked' delivery outcome (registered in cron_health monitoring) instead of the ambiguous generic 'suppressed'. - Drift-skip alerts explicitly bypass the ack gate (they carry the remediation command and alert once via drift_alerted already). - Docs: failure-incidents section in the cron guide. - Tests for the alerted transition + never-resurrect-closed. --------- Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>