CLI, gateway, TUI and ACP each re-sequenced the same chain (partial split -> estimate ->
_compress_context(force=True) -> lock-skip detection -> rejoin tail -> summary), and the
flag set differed per surface: TUI treated `--preview` as a focus topic, ACP ignored
arguments entirely. For the one command that legitimately breaks the prompt cache that
divergence is a correctness problem, not a style one.
`agent/conversation_compression_manual.py::compress_now` owns the sequence; surfaces parse
their own argv, install `after_messages`, re-anchor session ids and render. TUI and ACP gain
`--preview`, `--aggressive` refusal and `here [N]` parity.
A bare `astra: 0.85` in compression.model_thresholds was written for the Codex
OAuth route, where Astra is capped at 272K and 50% would compact at ~136K. The
key is substring-matched on the model name alone, so it also fired on
openai/gpt-6-astra via OpenRouter and Nous, where the window is 1.1M: the user's
0.5 global threshold was silently replaced by 0.85 and the session sat at 620K
(~59%) without compacting.
Keys may now carry a provider prefix: `"openai-codex:astra": 0.85` applies only
when the session's provider is openai-codex; bare keys keep their route-agnostic
behaviour. Ranking is by model-substring length with scope as the tie-break, so
`astra-900k` still outranks `openai-codex:astra` for the 900K picker. The
provider flows through ContextCompressor (ctor + update_model), the ContextEngine
base class and the TUI hot-reload path, so a /model switch between routes
re-scopes the override.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.