Addresses the human review findings on the memory-pressure feature:
* [P2] Dismissal hid later incidents of the same kind. The gateway now
publishes `boot_id` (the lifecycle sentinel's started_at — changes on
every gateway life) in the /api/status memory block, and the dashboard
keys OOM-restart dismissal on it: acknowledging one restart no longer
mutes the NEXT one (the OOM-loop case this banner exists for). Live
pressure dismissals now also reset once pressure is demonstrably back
to "ok" — "unknown" (stale heartbeat) is absence of evidence and
clears nothing. Dismissal storage moved to a JSON list; old bare-string
entries fail JSON.parse and degrade to a clean reset.
* [P2] suspected_oom is a heuristic (unclean exit + low-memory final
heartbeat), not proof the OOM killer acted — banner copy now says
"restarted unexpectedly, most likely because it ran out of memory"
instead of stating OOM as fact.
* [P3] Mobile header clearance was applied per-banner (mt-14 on both
MemoryPressureBanner and ProfileScopeBanner) AND on the content
(pt-14), double/triple-stacking 56px gaps when banners were visible.
Replaced with a single h-14 spacer above the banner stack.
Hosted agents can be OOM-killed hourly while the dashboard and the NAS
agent card both look perfectly healthy — every memory signal the gateway
already produces (heartbeat mem samples, lifecycle-ledger unclean-exit
verdicts, cache-pressure evictions) dies in server-side log files. The
BlueAtlas incident (NS-608) ran for three days like this.
This is the read-side fix:
* New gateway/memory_status.py distills the existing 30s loop heartbeat
(gateway RSS + system MemAvailable/MemTotal + swap) and the lifecycle
sentinel into a compact `memory` block: pressure ok/elevated/critical/
unknown, coarse MB numbers, and last-boot unclean/suspected-OOM flags.
Pure file reads, no new sampling, no gateway IPC. Stale (>150s) or
future-dated heartbeats degrade pressure to "unknown" so a dead
gateway's final gasp can't render a live "critical" banner forever.
Critical thresholds mirror the ledger's OOM-suspicion heuristics: if a
level would make a later unclean death "suspected OOM", warn at that
level while the process is still alive.
* lifecycle_ledger.record_startup now carries prior_unclean_exit /
prior_suspected_oom onto the reclaimed sentinel — previously the
verdict survived only in append-only diag prose. Flags age out on the
next sentinel rewrite (scoped to the life after the crash).
* /api/status serves the block (profile-aware, executor-offloaded,
fail-safe to pressure=unknown). Deliberately NOT folded into
components/overall: memory pressure is advisory, and flipping overall
to "degraded" on it would page NAS's availability sweep for a
condition the eviction valve is already handling. Public-safety:
coarse numbers/enums/booleans only — same disclosure class as the
existing nous_session_valid field, added for the same NAS-sweep
audience.
* Dashboard: new MemoryPressureBanner (app-shell, next to
ProfileScopeBanner) with worst-first trigger precedence
(critical > suspected-OOM restart > elevated), per-trigger
session-scoped dismissal, and escalation re-opening past a dismissal.
i18n keys optional with English fallbacks, matching the
managingProfileBanner convention.
Tests: gateway/test_memory_status.py (classification bands, staleness,
clock skew, corrupt files, bool-is-not-int), lifecycle sentinel
carry-forward, /api/status contract (block always present, collector
crash degrades instead of 500), and 7 banner component tests.
NAS-side ingestion (agent-card notice + memory-tier upsell) ships
separately.
Refs NS-656; context: NS-608, NS-657, OOF-77.