c0d974b19f
A gateway session whose summary model keeps timing out no longer retries compaction on the same fixed interval forever. The in-agent compressor already escalates repeat summary timeouts 60 -> 300 -> 900s (ContextCompressor.record_timeout_failure), but that ladder reads the in-memory _consecutive_timeout_failures counter and bind_session_state() zeroes it (context_compressor.py:1645). Session hygiene constructs a FRESH AIAgent for every run (gateway/run.py:16820) and re-binds state each time, so from the gateway that streak is structurally always 0 -- only the flat hygiene_failure_cooldown_seconds (300s) could ever be recorded. Issue #79624 reported exactly that steady state: an oversized session (1053 messages, ~119.5k tokens) whose aux model always timed out, re-attempting compaction every 300s across five days until the reporter deleted the session by hand. Track the streak on PersistentState instead, which outlives the per-run agent and is not cleared by turn/boundary resets, so consecutive hygiene failures climb 300 -> 900 -> 2700s and then saturate. Both failure sites (progress timeout and aborted compression) feed it; a real compression resets it, so a session that recovers starts from the first rung again. The ladder multiplies the configured base, so operators who tuned hygiene_failure_cooldown_seconds keep their first rung. Per-session, so one wedged chat cannot penalize other conversations. Deliberately NOT changed, since each is a maintainer policy call rather than a defect (all three are written up on #79624): - no durable failure-streak column, so escalation still resets on restart - the gateway 30s / in-agent 120s / aux-client 300s-floor timeout mismatch - no `hermes doctor` check or `hermes sessions list` marker for a session stuck in a compression-failure cooldown Note the reported exit(1) is NOT a crash: it is the deliberate _signal_initiated_shutdown path (gateway/run.py:26746-26751, #5646) that lets systemd Restart=on-failure revive the gateway after a bare SIGTERM, and it fires on every `systemctl restart` independently of compaction. The compaction log lines appear after the shutdown line because the gateway-owned executor is torn down with shutdown(wait=False, cancel_futures=True) (run.py:21164), so an in-flight turn keeps logging during teardown. Full analysis on the issue. Post-review hardening (Phase 2c + /simplify-code found five real defects in the first cut): - the recovery gate hand-rolled `_new_tokens < _approx_tokens` when a canonical predicate already existed: `compression_made_progress` (agent/turn_context.py, #39548). They disagree on 3 of 5 cases -- the hand-rolled form misses a row-count win when the summary keeps the token estimate flat, misses one where the summary is slightly MORE verbose (so a genuinely recovered session would keep escalating forever), and counts a sub-5% wobble as recovery. Now reuses the shared predicate, promoted from `_compression_made_progress` to a public name with the old private name kept as a back-compat alias so the existing importer (tests/agent/test_protected_tail_pressure_61932.py) and any patcher of that symbol keep working. - the reset was gated on "not aborted", but the degenerate "did not rotate or compact in place" branch (#21301) is NOT aborted and yields zero reduction, so a session wedged there reset its streak every run and could never escalate -- silently defeating the fix. Now gated on real progress. - no absolute ceiling: base * 9 reaches 9h at an operator base of 3600s, indistinguishable from "compaction switched off". Added _HYGIENE_COOLDOWN_MAX_SECONDS = 3600, mirroring the in-file _RECONNECT_BACKOFF_CAP precedent. - the reset used the get-or-create accessor to write a 0 that was already 0, materialising a _sessions entry (never evicted). Now peeks. - the abort verdict was probed twice, leaving the reset/record mutual exclusion implicit; a future await between the probes would have broken it silently. Computed once into _hyg_aborted. Tests: 19 new in tests/gateway/test_hygiene_failure_cooldown_ladder.py -- ladder escalation, saturation, the absolute cap, per-session isolation, reset-on-recovery, custom/zero base, PersistentState scoping (a mutation moving the field to TurnState fails), degraded runners, the progress gate, the exact progress-predicate semantics the gate depends on, and end-to-end that the escalated value is what reaches the state DB. All 12 mutations caught, including ones that restore the flat cooldown (the original bug), ungate the reset, swap the canonical predicate back for the hand-rolled comparison, remove the cap, and share the streak globally; the harness hard-errors when a mutation cannot be applied, since a silently no-op mutation check is worse than none -- an earlier version of it WAS silently no-opping after a refactor. The gate's contract test slices by AST node span rather than a fixed character count, which had already truncated once as the block grew. gateway hygiene + session-state + the three touched agent compression suites: 50 passed; ruff clean. E2E with real imports demonstrates the premise rather than asserting it: bind_session_state zeroes the in-agent counter, and the recorded deadlines go 300 -> 900 -> 2700 -> 2700 -> 2700s where they were previously a flat 300s. Reported by @yucezerey (#79624), whose state.db column dump and "deleting the session fixed it" datapoint made the real mechanism findable.