fix(openai): drop prompt_cache_options from Astra requests — not an SDK kwarg, 30m is the server default

Every direct-API (api.openai.com) Astra request raised
``TypeError: Responses.create() got an unexpected keyword argument 'prompt_cache_options'``
before reaching the network: openai 2.24.0's Responses.create has no such parameter and no
**kwargs, and neither send path relocates it into extra_body. The PR's tests stopped at
build_kwargs/preflight so the SDK boundary was never crossed.

OpenAI's prompt-caching guide states ``prompt_cache_options.ttl`` accepts only ``30m`` and that
``30m`` is the default, so the field carried no information: sending nothing yields the same
cache lifetime. The sanitizer now only removes what the API rejects (none/minimal effort,
sampling/logprob knobs, the pre-5.6 ``prompt_cache_retention``) and never adds a field, which
also keeps the request body byte-stable for the cache prefix.

Also: none/minimal→low no longer needs a bespoke {"", "none", "disabled", "off"} set —
``clamp_effort`` against CODEX_ASTRA_EFFORTS already resolves to the floor (``low``); and the
auxiliary adapter derives ``is_codex_backend`` from ``classify_responses_route`` (the declared
single owner of that predicate) instead of re-implementing the host test inline.

Tests reshaped to contracts: the two proxy/subdomain cases collapse into one parametrised
"exact host only" test asserting effort and temperature pass through untouched.
This commit is contained in:
kshitijk4poor
2026-09-07 21:35:51 +05:30
committed by kshitij
parent fdb76ce643
commit 170a73637d
6 changed files with 31 additions and 58 deletions
+3 -4
View File
@@ -1380,9 +1380,10 @@ class _CodexCompletionsAdapter:
if isinstance(reasoning_cfg, dict) and reasoning_cfg.get("enabled") is not False:
# Truthy-only: Codex 400s on e.g. {"effort": null}, so falsy → default. Shared
# per-model clamp with the main transport ("max" is gpt-5.6-only; "minimal"/"ultra" rejected).
from agent.codex_responses_adapter import classify_responses_route
from agent.reasoning_effort import clamp_effort
from agent.transports.codex import _codex_efforts_for_route
is_codex_backend = base_url_host_matches(host, "chatgpt.com") and "/backend-api/codex" in host.lower()
is_codex_backend = classify_responses_route(SimpleNamespace(base_url=host)).is_codex_backend
effort = clamp_effort(
reasoning_cfg.get("effort") or "medium",
_codex_efforts_for_route(model, host, is_codex_backend=is_codex_backend),
@@ -1439,9 +1440,7 @@ class _CodexCompletionsAdapter:
resp_kwargs["prompt_cache_retention"] = cache_retention
except Exception:
logger.debug("Codex auxiliary: prompt_cache_key derivation skipped", exc_info=True)
# Keep auxiliary Responses calls on the same exact-model contract as main
# turns. This runs last so caller overrides cannot reintroduce unsupported
# Astra reasoning/sampling fields or replace the official cache TTL.
# Last, like the main transport: caller extra_body must not put a rejected Astra field back.
from agent.transports.codex import _sanitize_astra_request_kwargs
_sanitize_astra_request_kwargs(resp_kwargs, model, host)
return resp_kwargs, model, timeout