6e5413844e
The legacy tail budget scales as threshold×target_ratio, which was designed around 128K windows at a 50% trigger (~13K tail). On modern big-window models with raised thresholds it silently hoards: a 1M-window session at threshold 0.85 keeps a 170K-token verbatim tail (255K soft ceiling) out of EVERY compaction, so a 540K manual /compress lands at ~290K and every subsequent turn re-ships the hoard. Nobody chooses this; it is an artifact of the formula outside its design envelope. Lean mode (#87326, compaction-v2) was built for exactly this and its recall was validated in the before/after eval (evals/compaction/results/): clamped 2.5%-of-window tail (10K floor / 25K cap), continuity carried by the upgraded summary (digests, anchor index, verbatim user messages, session_search recovery pointers). This flips the DEFAULT to lean; explicit 'tail_mode: legacy' in config keeps the old behavior exactly. Also fixes a latent bug the flip exposed: update_model() re-assigned the LEGACY formula directly when recomputing budgets, silently reverting a lean compressor to the hoard on every mid-session model switch. The recompute now routes through the mode-aware tail_token_budget property (regression test included). Surfaces: context_compressor.py defaults + getattr fallbacks, agent_init parse default, DEFAULT_CONFIG, gateway _CACHE_BUSTING_CONFIG_KEYS gains compression.tail_mode (mode changes now evict cached gateway agents like target_ratio changes do), user + developer docs. Tests: 3 new default contracts, legacy tests pinned explicitly, feasibility-skip scenario pinned to legacy (under lean its payloads correctly become compressible). E2E counterfactual (real imports, 1M window @ 0.85): main default: legacy, tail 170,000 (ceiling 255,000) head default: lean, tail 25,000 (ceiling 37,500) head legacy: 170,000 (opt-out intact) update_model to 400K: 10,000 (lean preserved across switch)