Every direct-API (api.openai.com) Astra request raised
``TypeError: Responses.create() got an unexpected keyword argument 'prompt_cache_options'``
before reaching the network: openai 2.24.0's Responses.create has no such parameter and no
**kwargs, and neither send path relocates it into extra_body. The PR's tests stopped at
build_kwargs/preflight so the SDK boundary was never crossed.
OpenAI's prompt-caching guide states ``prompt_cache_options.ttl`` accepts only ``30m`` and that
``30m`` is the default, so the field carried no information: sending nothing yields the same
cache lifetime. The sanitizer now only removes what the API rejects (none/minimal effort,
sampling/logprob knobs, the pre-5.6 ``prompt_cache_retention``) and never adds a field, which
also keeps the request body byte-stable for the cache prefix.
Also: none/minimal→low no longer needs a bespoke {"", "none", "disabled", "off"} set —
``clamp_effort`` against CODEX_ASTRA_EFFORTS already resolves to the floor (``low``); and the
auxiliary adapter derives ``is_codex_backend`` from ``classify_responses_route`` (the declared
single owner of that predicate) instead of re-implementing the host test inline.
Tests reshaped to contracts: the two proxy/subdomain cases collapse into one parametrised
"exact host only" test asserting effort and temperature pass through untouched.
A native Responses compaction checkpoint is opaque ciphertext; the rough
preflight estimator counts it as text (5.17M chars -> ~1.29M tokens against
a 204K trigger) and fires local compression on a request whose real prompt is
~116K. Arm the existing one-response real-usage latch when a replayable
checkpoint is captured (build_assistant_message) or restored into a fresh
agent (_hydrate_from_history), honor it in the post-tool gate and idle
compaction, and require non-empty encrypted_content for a checkpoint.
Squash of the author's source commits from #100642 (0e3c234ea0, 771e1b3365,
bb1505a119) plus the fdf140c81d test refresh, re-based onto current main by
patch application. Source delta is byte-identical to the PR head d6ce3e236d.
Fixes#100611
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Split _preflight_codex_input_items into per-item-type helpers over a
_PreflightCtx; streaming assembly moves into _CodexResponseAssembler with a
per-event dispatch table; run_codex_app_server_turn loses its session/usage/
interrupt regions to _ensure_codex_session/_finish_codex_turn/
_consume_user_interrupt/_queue_token_counts (counts built lazily so stub agents
without a session DB are never touched). Responses input/normalization and
stream callback order verified byte-identical against merge-base.
Follow-up to the #96217 salvage: the codex/xai/github route checks were
re-implemented inline at four sites (codex_responses_adapter helpers,
chat_completion_helpers kwargs build, _is_openai_codex_backend, the
run_agent silent-reject hint). Consolidate them into
classify_responses_route() / ResponsesRouteFlags in
codex_responses_adapter and migrate every site — backend-identity
predicate class (#22548/#70893/#59561/#72468).
Host checks use exact-host-or-subdomain semantics, never substring
matching.
Automatic preflight used the full durable transcript even when the
Codex Responses request would prune around a native compaction
checkpoint. That false-triggered a 600s local summary against history
the main request never sent. Estimate the converted, checkpoint-pruned
payload when native compaction is eligible, and keep the generic
estimate as the conservative fallback.
/simplify-code reuse+quality reviewers both flagged the byte-identical
fc_->call_ synthesis blocks in the assistant and tool-result branches as
a correctness coupling — the two sites MUST stay in lockstep or pairing
breaks. Extract _canonical_call_id_from_fc() and route both through it.
Mutation check: pairing regression test fails when the tool-branch call
is stubbed out, green after restore.
The sweeper review on #49224 flagged that the assistant branch synthesizes
call_<suffix> from an fc_-only id while the tool-result branch kept the raw
fc_... string — so an oversized pair hashed to two DIFFERENT clamped
surrogates and the function_call_output arrived unmatched (HTTP 400).
Canonicalize the tool-result side to the same call_<suffix> before
clamping. Also fixes the pre-existing short-fc_ pairing mismatch
(call_short123 vs fc_short123). Regression test covers both lengths.
A degenerate tool name stored in conversation history (dots, spaces,
unicode from an earlier model degeneration) bricks every subsequent
Codex Responses turn with a non-retryable HTTP 400:
Invalid input[N].name: string does not match pattern '^[a-zA-Z0-9_-]+'
The 400 replays forever until the user manually starts a new session.
Add _sanitize_replayed_fn_name() — replaces invalid chars with '_'
(runs collapsed), degrades all-invalid names to 'fn' instead of empty
(an empty name would trade one 400 for a preflight ValueError). Applied
at both replay sites: the chat-message converter and the preflight
choke-point. Live tool-definition names are left untouched — they must
match the dispatch registry exactly. Pairing is by call_id, so
renaming a replayed function_call is safe.
call_id overflow (the sibling half of #49224) was already fixed on main
by #73492 (_clamp_responses_call_id); this commit covers the remaining
invalid-name defect.
Credit: @Morad37 (#31678 — identified the bug, the replay sites, and
the regex contract), @lubosxyz (#49224 — replace-not-strip semantics
and 'fn' fallback to avoid the empty-name trap).
Fixes#31666
prune_pre_checkpoint_items() had a hardcoded role=='user' filter that
discarded all non-user messages before a checkpoint — including Hermes'
own compression summaries (role='assistant'), causing total context amnesia
about past conversation summaries.
The fix:
- _is_summary_item delegates to the canonical
agent.context_compressor.is_compaction_summary_message provenance check
(not an ad-hoc heuristic)
- Summaries are retained whole (never byte-sliced) within a 32k token budget
- Idempotent across repeated checkpoints (dedup by identical text)
- _chat_messages_to_responses_input threads item_sources (raw chat messages)
through to the pruner, so it can read summary content directly from the
source when the Responses conversion shape is lossy (tool-result carrier
becomes function_call_output, or stale codex_message_items replay shadows
merged content)
Fixes#90975.
Salvage of #90976 by @JoaoMarcos44.
A captured native-compaction checkpoint lives in the persisted
codex_reasoning_items sidecar, but the wire restructure that follows it
(prune_pre_checkpoint_items) ran unconditionally: the native gate only
decided whether context_management went into the request, and no signal
from it ever reached _chat_messages_to_responses_input.
So a single checkpoint kept deleting every pre-checkpoint item from all
later requests — after a mid-session swap out of the gpt-5.6 family,
after compression.enabled: false, after the rejection kill switch, and
after a session resume that reloads the sidecar from state.db. The model
receiving the opaque blob was no longer the one able to decode it, and
nothing was logged.
Thread a single native_compaction_eligible boolean, derived from the same
value that gates the context_management field, into the converter. When
ineligible: do not replay type: "compaction" items and do not prune. Safe
because native compaction never truncates Hermes' local history, so the
fallback still carries the full conversation.
All Responses call sites are covered: build_kwargs and convert_messages
derive the flag via _native_compaction_active, the auxiliary/compression
client is explicitly ineligible, and the converter defaults to False
(pre-feature wire) so future call sites are safe by construction.
Fixes#85914
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Live verification (gpt-5.6 @ api.openai.com) proved the Responses server
renders NOTHING placed before a replayed compaction checkpoint: a fact
stated in a pre-checkpoint input item is invisible to the model, while the
same item after the checkpoint recalls perfectly. Hermes was replaying the
full pre-checkpoint transcript anyway — dead upload weight, and worse, every
plaintext user ask from before the boundary silently vanished from the
model's view, surviving only inside the opaque server summary. That is the
goal-drift failure mode reported against native compaction sessions.
Codex CLI never hits this because it rebuilds history client-side after
compaction, retaining user messages verbatim under a token budget. This
change is the wire-level equivalent: when a replayed checkpoint is present,
_chat_messages_to_responses_input restructures the input as
[newest checkpoint run] + [retained pre-checkpoint user messages,
newest-first within a 64K-token budget] + [post-checkpoint tail]
Histories without a checkpoint are returned unchanged, so non-native
sessions see a byte-identical wire.
Opt-in via compression.codex_responses_native (default: false). When enabled,
gpt-5.6-family models on the direct OpenAI API (api.openai.com) or a ChatGPT
Codex subscription send context_management=[{type: compaction,
compact_threshold: N}] on Responses requests. OpenAI compacts server-side and
returns an encrypted compaction output item; Hermes captures it into the
existing codex_reasoning_items sidecar and replays it on later turns in place
of the pruned history — inheriting persistence, session replay, the
cross-issuer guard, and the encrypted-replay kill switch with zero new state.
Scope is deliberately hard-gated (agent/native_compaction.py, re-checked per
request): gpt-5.6 family only — gpt-5.1/5.2 fail server-side on the field
(HTTP 500 / stream stall, no structured rejection; live-verified) — and
direct OpenAI/Codex routes only; xAI, GitHub/Copilot, OpenRouter, relays,
and local servers never see the field.
Hermes' local compression stays armed as the fallback owner: the native
threshold is clamped ~8K tokens below the local trigger so the server
compacts first, and a structured provider rejection of context_management
disables native compaction for the session and retries without it
(one-shot guard in TurnRetryState).
Live-verified E2E on api.openai.com/gpt-5.6: server compaction fired at a
4K threshold, checkpoints captured and replayed, recall preserved across
3 turns; gpt-5.1 with the flag enabled stays clean (field never sent).
Direction credit: PR #76950 by @laryhorb explored native Responses
compaction; this is a minimal reimplementation on current main.
The codex app-server namespaces MCP tool call ids as
codex_mcp__<server>__<tool>_<codex_call_id>. With an exec-<uuid> component the
built-in hermes-tools server alone overflows the Responses API's 64-char
call_id limit, so the request 400s with a non-retryable "string too long".
The offending item sits near the head of the transcript and replays every
turn, permanently bricking the session — the only recovery is /reset.
Sibling defect to #10788, which clamped input[*].id via
_MAX_RESPONSES_ITEM_ID_LENGTH. Apply the same treatment to call_id at both
Responses emit sites in _chat_messages_to_responses_input: a deterministic
surrogate (call_ + sha256[:32]) for ids over the limit, short ids unchanged.
Because the surrogate is a pure function of the original id, a function_call
and its matching function_call_output — which carry the same original id — map
to the same surrogate and stay paired without correlating the two items.
OpenAI documents GPT-5.5 / GPT-5.5 Pro as extended-cache-only: in-memory
prompt cache retention is not available for them, and only
prompt_cache_retention: "24h" is supported. Responses requests that omit
the field see near-zero cached_tokens even with a stable prompt_cache_key
and identical prefixes (observed on an OpenAI-compatible Responses relay:
0 cached across repeated identical calls before; 97% cache reads after).
Send the field for the gpt-5.5 model family (bare and namespaced ids like
openai.gpt-5.5) on OpenAI-compatible Responses routes, mirrored in the
auxiliary Codex adapter, and pass it through preflight normalization.
Skipped for xAI, GitHub/Copilot, and the chatgpt.com Codex backend, which
reject or ignore body-level cache fields.
Map Codex Responses status=incomplete with incomplete_details.reason=content_filter to finish_reason=content_filter so the existing refusal/fallback path runs instead of burning incomplete continuation attempts.
grok-4.x on the xAI /v1/responses surface sometimes ends a turn with only
reasoning items — no message output item, no tool calls — and those
reasoning items carry no encrypted_content. Two compounding problems:
1. The model occasionally emits its final answer INSIDE the reasoning
channel, delimited by grok's internal "<response>" tag. The answer
exists but is classified reasoning-only → finish_reason=incomplete.
2. An interim assistant message holding only plain-text reasoning replays
as nothing in _chat_messages_to_responses_input, so every continuation
request is byte-identical to the one that just failed. The model
deterministically repeats the reasoning-only response until the retry
budget is exhausted and the turn dies with "Codex response remained
incomplete after 3 continuation attempts".
Fixes:
- _normalize_codex_response (xai_responses only): salvage the
<response>-delimited tail from the reasoning text and promote it to
assistant content; the untagged prefix stays as thinking text.
- Codex-incomplete continuation path: when the interim message has
nothing the input converter will replay (no content, no encrypted
reasoning items, no message items), append a user-role nudge so the
retry actually differs and explicitly asks for the final answer /
pending tool call. Mirrors the existing _get_continuation_prompt
pattern used for length truncation.
Observed live with grok-4.20 on xai-oauth (2026-07-13); sibling of the
grok-composer web_search incomplete-loop fix in transports/codex.py.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Follow-up to the salvaged #64449: the status-trusting branch flipped
github_responses to 'stop' alongside unknown relays. Copilot fronts the
same OpenAI model family as codex_backend and shows the same
reasoning-only 'still thinking' degeneration, so it stays on the
continuation path. Only unrecognized (other:*) backends trust
response.status='completed' as terminal.
Copilot (api.githubcopilot.com/responses) binds replayed assistant
codex_message_items ids to a specific backend "connection". Credential-
pool rotation, a gateway restart, or routine load-balancer churn between
turns all invalidate that binding, and Copilot rejects the stale id with
HTTP 401 "input item ID does not belong to this connection" — even for
short ids well under the #27038 64-char length cap, since this is a
connection-scope problem, not a length problem. Once a session captures
one of these ids it is persisted and replayed forever, permanently
bricking the session.
Thread an is_github_responses flag from build_kwargs/convert_messages
into _chat_messages_to_responses_input and drop the id unconditionally
on that path, mirroring how reasoning items already strip id on replay.
phase/status/content are still replayed so cache-relevant signal isn't
lost — only the connection-scoped id is unsafe to reuse.
Written to apply independently of the #27038 length-cap fix so the two
PRs don't block each other; they touch adjacent conditions in the same
block and merge cleanly in either order.
Fixes#32716
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Codex assigns assistant message items server-side ids that can run
400+ chars (base64 encrypted blobs), but the Responses API caps
input[].id at 64 chars and rejects the whole request with a
non-retryable HTTP 400. Once a session captures one of these long
ids, every subsequent turn replays it and 400s forever, since the
history persists it in codex_message_items.
Add a 64-char length guard at both replay sites — the history-to-
input converter and the final preflight gate — so oversized ids are
dropped while short ids (msg_...) are kept for prefix-cache hits.
Mirrors the existing pattern for reasoning items, which already
strip their id before replay because store=False means the API
can't resolve ids server-side anyway.
Fixes#27038
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
GPT-5.x models on the Codex Responses API emit short pre-tool-call
"preamble" text as message items with phase="commentary". Previously,
_normalize_codex_response() added ALL message items to content_parts
regardless of phase, causing commentary text to leak as visible
assistant content on chat gateways.
Fix: when normalized_phase is "commentary" or "analysis", route the
message text to reasoning_parts instead of content_parts. This keeps
preamble/internal planning in the reasoning channel where it belongs.
FixesNousResearch/hermes-agent#41293
- model_metadata: grok-composer-2.5-fast → 262144 (OAuth slug not in /v1/models)
- codex transport: inject native {"type":"web_search"} for is_xai_responses;
drop client web_search to avoid duplicate-name 400s
- codex adapter: do not treat in-progress server-side *_call items as incomplete
- tests: adapter, transport build_kwargs, model_metadata, oauth recovery