9c76c133b7
Hosted agents that die uncleanly (kernel OOM kill, SIGKILL, whole-VM death) leave no trace: shutdown_forensics only covers graceful signals, gateway-exit-diag.log only covers exit paths that actually run, and the VM reboot wipes dmesg before anyone can capture it. NS-608 (BlueAtlas hourly crash cycle, July 12-15) took days of manual log correlation to classify because nothing recorded 'the previous life ended violently'. Add gateway/lifecycle_ledger.py — a sentinel state machine persisted to <HERMES_HOME>/state/gateway.lifecycle.json: - start_gateway() claims the sentinel (phase=running) right after the PID-file/runtime-lock claim, and reports any prior life that never reached an exit path as gateway.previous_unclean_exit in gateway-exit-diag.log + a WARNING log line. - Every exit funnel marks the sentinel exited with a reason: _exit_after_graceful_shutdown (graceful_shutdown), the shutdown watchdog (shutdown_watchdog), and the loop-liveness watchdog (loop_liveness_watchdog). - Ownership-guarded for --replace takeovers: a live matching owner is never reported dead, and the old life cannot clobber the replacement's freshly claimed sentinel on its way out. The 30s loop heartbeat now embeds a cheap /proc memory sample (own RSS, MemAvailable, swap used) so every unclean-death report carries a 'memory N seconds before death' snapshot; the detector flags suspected_oom when the last sample shows <64MiB or <5% available. container-boot.log lines gain prior_exit=clean|unclean|unknown per profile, stamping unclean container deaths into the volume-persisted boot log where support can grep for them. Tests: tests/gateway/test_lifecycle_ledger.py (16 cases) + 4 new container-boot annotation cases. Existing watchdog/forensics/boot suites all green; ruff clean.