Commit Graph

2 Commits

Author SHA1 Message Date
Shannon Sands ba5dc00bb2 fix(memory-status): review follow-ups — incident-keyed dismissal, honest copy, single mobile offset (NS-656)
Addresses the human review findings on the memory-pressure feature:

* [P2] Dismissal hid later incidents of the same kind. The gateway now
  publishes `boot_id` (the lifecycle sentinel's started_at — changes on
  every gateway life) in the /api/status memory block, and the dashboard
  keys OOM-restart dismissal on it: acknowledging one restart no longer
  mutes the NEXT one (the OOM-loop case this banner exists for). Live
  pressure dismissals now also reset once pressure is demonstrably back
  to "ok" — "unknown" (stale heartbeat) is absence of evidence and
  clears nothing. Dismissal storage moved to a JSON list; old bare-string
  entries fail JSON.parse and degrade to a clean reset.

* [P2] suspected_oom is a heuristic (unclean exit + low-memory final
  heartbeat), not proof the OOM killer acted — banner copy now says
  "restarted unexpectedly, most likely because it ran out of memory"
  instead of stating OOM as fact.

* [P3] Mobile header clearance was applied per-banner (mt-14 on both
  MemoryPressureBanner and ProfileScopeBanner) AND on the content
  (pt-14), double/triple-stacking 56px gaps when banners were visible.
  Replaced with a single h-14 spacer above the banner stack.
2026-08-13 20:30:12 -07:00
Shannon Sands e11d1ddc7f feat(status): surface memory pressure and suspected-OOM restarts to users (NS-656)
Hosted agents can be OOM-killed hourly while the dashboard and the NAS
agent card both look perfectly healthy — every memory signal the gateway
already produces (heartbeat mem samples, lifecycle-ledger unclean-exit
verdicts, cache-pressure evictions) dies in server-side log files. The
BlueAtlas incident (NS-608) ran for three days like this.

This is the read-side fix:

* New gateway/memory_status.py distills the existing 30s loop heartbeat
  (gateway RSS + system MemAvailable/MemTotal + swap) and the lifecycle
  sentinel into a compact `memory` block: pressure ok/elevated/critical/
  unknown, coarse MB numbers, and last-boot unclean/suspected-OOM flags.
  Pure file reads, no new sampling, no gateway IPC. Stale (>150s) or
  future-dated heartbeats degrade pressure to "unknown" so a dead
  gateway's final gasp can't render a live "critical" banner forever.
  Critical thresholds mirror the ledger's OOM-suspicion heuristics: if a
  level would make a later unclean death "suspected OOM", warn at that
  level while the process is still alive.

* lifecycle_ledger.record_startup now carries prior_unclean_exit /
  prior_suspected_oom onto the reclaimed sentinel — previously the
  verdict survived only in append-only diag prose. Flags age out on the
  next sentinel rewrite (scoped to the life after the crash).

* /api/status serves the block (profile-aware, executor-offloaded,
  fail-safe to pressure=unknown). Deliberately NOT folded into
  components/overall: memory pressure is advisory, and flipping overall
  to "degraded" on it would page NAS's availability sweep for a
  condition the eviction valve is already handling. Public-safety:
  coarse numbers/enums/booleans only — same disclosure class as the
  existing nous_session_valid field, added for the same NAS-sweep
  audience.

* Dashboard: new MemoryPressureBanner (app-shell, next to
  ProfileScopeBanner) with worst-first trigger precedence
  (critical > suspected-OOM restart > elevated), per-trigger
  session-scoped dismissal, and escalation re-opening past a dismissal.
  i18n keys optional with English fallbacks, matching the
  managingProfileBanner convention.

Tests: gateway/test_memory_status.py (classification bands, staleness,
clock skew, corrupt files, bool-is-not-int), lifecycle sentinel
carry-forward, /api/status contract (block always present, collector
crash degrades instead of 500), and 7 banner component tests.

NAS-side ingestion (agent-card notice + memory-tier upsell) ships
separately.

Refs NS-656; context: NS-608, NS-657, OOF-77.
2026-08-13 20:30:12 -07:00