c49e2a496e
Follow-up to the mid-turn pre-API guard fix (#96995 / #97602): sweep the remaining call sites that derive automatic compression pressure from a generic estimate over the assembled durable history, which on a compacted native-Codex session overstates the wire payload by orders of magnitude. - agent/turn_context.py idle-triggered compaction: use _preflight_request_tokens (anchor -> native pruned -> generic) instead of the raw generic request estimate, so resuming a compacted codex session after an idle gap does not fire a compaction the next request never needed. - agent/turn_context.py uncompressed-session overflow-warn RE-ARM: match the warn site's route-aware figure so the dedup re-arms correctly on native sessions. - agent/conversation_loop.py post-response should_compress fallback (last_prompt_tokens==0, i.e. no provider usage after a disconnect or gateway restart — the unanchored case in #97602's repro): route through _midturn_request_pressure_tokens instead of the generic figure. Left alone deliberately: provider-proven overflow recovery paths (413 / context-length errors — the provider already proved the request does not fit, figures there only arm recovery and score progress), compression progress before/after pairs (relative deltas on the same scale), manual /compress display estimates (gateway/CLI/ACP feedback, not automatic triggers), MoA advisor budget trimming (not a codex-native wire payload), and context_compressor internals (measure local durable-history shrink).