4252aecc2e
The MINIMUM_CONTEXT_LENGTH floor in _compute_threshold_tokens only degraded to the 85% trigger when it met or exceeded the effective window exactly (#14690). Near-minimum windows slipped through: at context_length=65536 the threshold passed through at 64,000 — 97.7% of the window, ~1.5K tokens of output room — so pre-API compaction effectively could not fire. Providers that silently truncate over-window prompts instead of rejecting them (e.g. ollama's OpenAI-compatible /v1 endpoint) never deliver the reactive context-overflow backstop either. Observed live on a 65,536-token local model: the session rode into the window ceiling and each length-continuation retry re-sent a window-filling prompt (65,120 -> 65,273 prompt tokens, 263 output tokens of room) until the turn died with "Response remained truncated after 4 continuation attempts" — every retry paying a full multi-minute prefill. Cap the floored threshold at _MIN_CTX_TRIGGER_RATIO (85%) of the effective input budget whenever the floor is the binding term. An explicit threshold_percent above 85% is user intent and stays uncapped; windows where the floor lands at/below the cap are unchanged.