feat(model): Codex GPT slugs default back to 272K; explicit -900k picker variants opt into the verified large window
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the live-verified 900K burned through subscription usage for users who never asked for the larger window (bigger window = more input tokens per request). - Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the advertised 272K again — the cheaper limit is the default. - The model picker synthesizes explicit <slug>-900k variants (e.g. gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into the 900K window. Slugs that genuinely enforce 272K (gpt-5.5, gpt-5.4-mini) get no variant. - The -900k suffix is Hermes-side only: stripped before the model id hits the wire (main transport + auxiliary Responses adapter), and pricing aliases the variants onto the base entries. - Docs: new opt-in section in context-compression-and-caching.md.
This commit is contained in:
@@ -994,11 +994,16 @@ _OFFICIAL_DOCS_PRICING: Dict[tuple[str, str], PricingEntry] = {
|
||||
|
||||
# GPT-5.6 "-pro" high-effort variants bill at the same per-token rates as
|
||||
# their base tiers (more tokens per task, not a higher rate). Alias them
|
||||
# onto the base entries so the snapshot stays single-source.
|
||||
# onto the base entries so the snapshot stays single-source. The Hermes-side
|
||||
# "-900k" large-context Codex picker variants are the same underlying model
|
||||
# (the suffix is stripped on the wire), so they alias identically.
|
||||
for _base_56 in ("gpt-5.6-sol", "gpt-5.6-terra", "gpt-5.6-luna"):
|
||||
_OFFICIAL_DOCS_PRICING[("openai", f"{_base_56}-pro")] = _OFFICIAL_DOCS_PRICING[
|
||||
("openai", _base_56)
|
||||
]
|
||||
_OFFICIAL_DOCS_PRICING[("openai", f"{_base_56}-900k")] = _OFFICIAL_DOCS_PRICING[
|
||||
("openai", _base_56)
|
||||
]
|
||||
del _base_56
|
||||
|
||||
# The direct Gemini provider currently exposes preview IDs for these two
|
||||
|
||||
Reference in New Issue
Block a user