Commit Graph

4 Commits

Author SHA1 Message Date
Teknium 7eb8ccc315 refactor(agent/runtime): monitoring — shared OTLP plumbing in otlp_exporter, event base class, EmitterStreamer, dedupe health export
- otlp_exporter hosts SDK loading (symbol table), header/endpoint/resource helpers
  shared with gateway_health_export; export module imports them instead of
  keeping copies (_otlp_config/_resolve_headers/_install_id/_safe_resource_attributes
  unified). _KEEP_BY_KIND attribute allowlist reused for diagnostic log attrs.
- events: _MonitoringEvent base supplies to_dict; field order (wire order) unchanged.
- redaction: redact_bounded() replaces 3 inline redact+truncate try/excepts.
- EmitterStreamer base owns the shared unsubscribe/flush/shutdown for
  OTLPStreamer and GatewayDiagnosticLogStreamer.
- gateway_health: metric construction helpers, _contains_any shared with
  cron_health, dead _allowed_logger/redact_gateway_message removed (0 refs).
- Comment/docstring compaction keeping every stated invariant.
2026-09-02 13:29:47 -07:00
Teknium 9de5460c12 feat(cron): acked failure signatures stop re-pinging — durable incidents + ack CLI (salvage #94692) (#95017)
* feat(cron): durable failure incidents with signature dedup and ack

Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.

- cron/incidents.py: lazily-created incident table (detected -> alerted ->
  reviewed -> closed lifecycle; closed is per-signature terminal), sha256
  signature dedup over job_id + normalized error, redacted/truncated error
  storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
  suppress the per-run failure ping when the exact signature is acked (both
  the normal failure path and the processing-raised retry path). Best-effort:
  an incident-store error never breaks the cron run or delivery. Streak nudge,
  alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
  `hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
  classification, lazy-schema, scheduler gating, and CLI coverage.

Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.

* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome

Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
  lives in INCIDENT_STATES so future slices can add states without a
  table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
  delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
  delivery outcome (registered in cron_health monitoring) instead of
  the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
  remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.

---------

Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>
2026-08-25 14:02:40 -07:00
Victor Kyriazakos a65a647b04 fix(monitoring): correct cron operational signals 2026-07-24 18:54:46 +00:00
Victor Kyriazakos efbf2fd79c feat(monitoring): add cron operational telemetry 2026-07-24 18:54:46 +00:00