fix(openai): drop prompt_cache_options from Astra requests — not an SDK kwarg, 30m is the server default
Every direct-API (api.openai.com) Astra request raised
``TypeError: Responses.create() got an unexpected keyword argument 'prompt_cache_options'``
before reaching the network: openai 2.24.0's Responses.create has no such parameter and no
**kwargs, and neither send path relocates it into extra_body. The PR's tests stopped at
build_kwargs/preflight so the SDK boundary was never crossed.
OpenAI's prompt-caching guide states ``prompt_cache_options.ttl`` accepts only ``30m`` and that
``30m`` is the default, so the field carried no information: sending nothing yields the same
cache lifetime. The sanitizer now only removes what the API rejects (none/minimal effort,
sampling/logprob knobs, the pre-5.6 ``prompt_cache_retention``) and never adds a field, which
also keeps the request body byte-stable for the cache prefix.
Also: none/minimal→low no longer needs a bespoke {"", "none", "disabled", "off"} set —
``clamp_effort`` against CODEX_ASTRA_EFFORTS already resolves to the floor (``low``); and the
auxiliary adapter derives ``is_codex_backend`` from ``classify_responses_route`` (the declared
single owner of that predicate) instead of re-implementing the host test inline.
Tests reshaped to contracts: the two proxy/subdomain cases collapse into one parametrised
"exact host only" test asserting effort and temperature pass through untouched.
This commit is contained in:
@@ -1380,9 +1380,10 @@ class _CodexCompletionsAdapter:
|
||||
if isinstance(reasoning_cfg, dict) and reasoning_cfg.get("enabled") is not False:
|
||||
# Truthy-only: Codex 400s on e.g. {"effort": null}, so falsy → default. Shared
|
||||
# per-model clamp with the main transport ("max" is gpt-5.6-only; "minimal"/"ultra" rejected).
|
||||
from agent.codex_responses_adapter import classify_responses_route
|
||||
from agent.reasoning_effort import clamp_effort
|
||||
from agent.transports.codex import _codex_efforts_for_route
|
||||
is_codex_backend = base_url_host_matches(host, "chatgpt.com") and "/backend-api/codex" in host.lower()
|
||||
is_codex_backend = classify_responses_route(SimpleNamespace(base_url=host)).is_codex_backend
|
||||
effort = clamp_effort(
|
||||
reasoning_cfg.get("effort") or "medium",
|
||||
_codex_efforts_for_route(model, host, is_codex_backend=is_codex_backend),
|
||||
@@ -1439,9 +1440,7 @@ class _CodexCompletionsAdapter:
|
||||
resp_kwargs["prompt_cache_retention"] = cache_retention
|
||||
except Exception:
|
||||
logger.debug("Codex auxiliary: prompt_cache_key derivation skipped", exc_info=True)
|
||||
# Keep auxiliary Responses calls on the same exact-model contract as main
|
||||
# turns. This runs last so caller overrides cannot reintroduce unsupported
|
||||
# Astra reasoning/sampling fields or replace the official cache TTL.
|
||||
# Last, like the main transport: caller extra_body must not put a rejected Astra field back.
|
||||
from agent.transports.codex import _sanitize_astra_request_kwargs
|
||||
_sanitize_astra_request_kwargs(resp_kwargs, model, host)
|
||||
return resp_kwargs, model, timeout
|
||||
|
||||
Reference in New Issue
Block a user