6bbe55dd09
The per-session agent cache is capped at 128 entries with a 1h idle TTL, and neither bound knows how many bytes it holds. Each cached agent pins _session_messages -- the full transcript including tool output, tens of MB on a session with 100+ tool calls -- so a gateway serving many chats keeps every warm transcript resident: agents that took a turn inside the TTL are never idle-swept, and the idle sweep additionally defers finalizable sessions until they expire. RSS climbs until the cgroup throttles and SIGTERM can no longer flush inside systemd's stop timeout. Add the missing bound. Each session-expiry watcher tick compares the process's anonymous RSS against a budget and, when over, sheds LRU agents through the same soft-eviction path the cap enforcer uses, then runs malloc_trim so the freed arenas actually return to the OS. Evicted sessions rebuild their transcript from the persisted session on the next turn. Three classes of session are never shed: agents mid-turn, the most recently used ones, and any session whose transcript has not finished reaching disk (_last_flushed_db_idx vs len(_session_messages) -- the same divergence the FTS write-corruption guard reacts to when it preserves live history). memory_high_mb defaults to "auto", deriving the budget from the cgroup limit the gateway runs under, so a MemoryHigh/MemoryMax on the unit is respected without a second number to keep in sync. The two existing bounds become configurable alongside it under agent.agent_cache. protect_recent is clamped to half the cache: a couple of sessions can exhaust the budget on their own, and a fixed MRU guard would then protect everything and leave the gateway climbing with nothing it would shed. Fixes #80764