Commit Graph

24 Commits

Author SHA1 Message Date
kshitijk4poor 00f0d92be0 fix(codex): canonicalise persisted issuer stamps via the route-identity owner
`_classify_responses_issuer` reimplemented endpoint canonicalisation with
its own urlsplit/urlunsplit pass. The repo already owns that logic in
`hermes_cli/route_identity.py::normalize_route_base_url` (stdlib-only,
used by agent/backend_identity.py), so delegate to it.

Reasoning items persisted before canonicalisation were stamped with the
raw `other:<agent.base_url>` (trailing slash, host case). Comparing them
verbatim against the now-canonical `current_issuer_kind` marked them
foreign and dropped them on the very endpoint that minted them. Run the
persisted stamp through the same canonicaliser (`_canonical_issuer_kind`,
non-`other:` kinds untouched) before comparing.

Also fixes the `_chat_messages_to_responses_input` docstring, which still
stated the pre-stack rule that legacy endpoint-stamped items drop when the
current model is known; the stack replays them on a matching issuer.
2026-09-15 10:49:19 +05:30
kshitijk4poor 16986c4bff fix(codex): canonicalise the custom-endpoint issuer kind
The openai SDK appends a trailing slash to `client.base_url`, so the aux
adapter stamped `other:https://h/v1/` while the main transport stamped
`other:https://h/v1`. On custom Responses endpoints every aux call
(compression, flush_memories) therefore dropped all main-minted reasoning
items as "foreign".

`_classify_responses_issuer` now strips whitespace and trailing slashes and
lowercases scheme+netloc before stamping. The aux adapter also derives its
route flags from `classify_responses_route` — the single owner of the
codex/xai/github predicates — instead of an inline chatgpt.com host check,
and reuses the same flags for the effort clamp.
2026-09-15 10:49:19 +05:30
kshitijk4poor cb49660bb4 fix(codex): replay legacy endpoint-stamped reasoning without a model stamp
Native compaction checkpoints and reasoning items persisted before model
stamping existed carry only `_issuer_kind`. Treating a missing `_issuer_model`
as foreign dropped every such item once the current model was known, which
wiped existing sessions' native-compaction context on upgrade (four consumer
tests in test_native_compaction / test_native_preflight_estimate /
test_413_compression went red on the stack).

Trust the endpoint stamp when no model stamp is present, as main does today.
Items minted after this change carry the model stamp and still drop on a
same-endpoint model switch; a wrong guess on a legacy item is caught by the
invalid_encrypted_content 400 classifier and the replay kill switch.
2026-09-15 10:49:19 +05:30
Fangliquan 51ebdff570 fix(codex): scope encrypted-reasoning replay to the issuing model
Encrypted reasoning blobs are sealed to the model that minted them, not
only to the endpoint. Switching models on the same custom Responses
endpoint therefore replayed blobs the new model cannot decrypt and the
turn failed with HTTP 400.

Stamp captured reasoning items with `_issuer_model` (the canonical wire
model) alongside `_issuer_kind`, and replay an item only when both the
issuer kind and the model match the current request. Endpoint-stamped
legacy items without model provenance are dropped once the current
model is known (fail closed); ordinary assistant text stays replayable.
The transport threads the effective wire model (request_overrides win)
into conversion and normalization; the auxiliary Codex adapter stamps
and filters against its own model rather than the main agent's. The
400 classifier also recognises the custom-endpoint wording
"encrypted content could not be decrypted or parsed" so recovery strips
the replay state instead of aborting.

Hand-grafted from #95849 (final head d9cf6bcc08) onto current main; the
middleware-model-rewrite half is intentionally left out.

Closes #95834
2026-09-15 10:49:19 +05:30
fangliquanflq f7b6a2b59f fix(codex): drop foreign replay message ids
(cherry picked from commit b58b94e16ccf5c8e01813f25f6a1a94dfb17f79b)
2026-09-15 10:49:19 +05:30
Teknium 11082f603e test: consolidate video rejection variants into one invariant 2026-09-07 06:04:44 -07:00
fangliquanflq 00b4c49939 test(agent): cover all rejected video part types
Co-authored-by: crazyief <8566250+crazyief@users.noreply.github.com>
2026-09-07 06:04:44 -07:00
fangliquanflq 2b7b940046 fix(agent): reject unsupported Codex video input 2026-09-07 06:04:44 -07:00
lEWFkRAD b80b9d8271 fix(codex): preserve assistant image slots in replay 2026-08-28 04:58:14 -07:00
lEWFkRAD 8de45940fb fix(codex): drop assistant images from Responses replay
Fixes #96816
2026-08-28 04:58:14 -07:00
kshitijk4poor 635232ec4e fix(codex): canonicalize fc_-only tool-result ids to match the call side
The sweeper review on #49224 flagged that the assistant branch synthesizes
call_<suffix> from an fc_-only id while the tool-result branch kept the raw
fc_... string — so an oversized pair hashed to two DIFFERENT clamped
surrogates and the function_call_output arrived unmatched (HTTP 400).

Canonicalize the tool-result side to the same call_<suffix> before
clamping. Also fixes the pre-existing short-fc_ pairing mismatch
(call_short123 vs fc_short123). Regression test covers both lengths.
2026-08-26 12:58:35 +05:30
kshitijk4poor 31485d50ea fix(codex): sanitize replayed function_call.name to Responses API pattern (#31666)
A degenerate tool name stored in conversation history (dots, spaces,
unicode from an earlier model degeneration) bricks every subsequent
Codex Responses turn with a non-retryable HTTP 400:
  Invalid input[N].name: string does not match pattern '^[a-zA-Z0-9_-]+'

The 400 replays forever until the user manually starts a new session.

Add _sanitize_replayed_fn_name() — replaces invalid chars with '_'
(runs collapsed), degrades all-invalid names to 'fn' instead of empty
(an empty name would trade one 400 for a preflight ValueError).  Applied
at both replay sites: the chat-message converter and the preflight
choke-point.  Live tool-definition names are left untouched — they must
match the dispatch registry exactly.  Pairing is by call_id, so
renaming a replayed function_call is safe.

call_id overflow (the sibling half of #49224) was already fixed on main
by #73492 (_clamp_responses_call_id); this commit covers the remaining
invalid-name defect.

Credit: @Morad37 (#31678 — identified the bug, the replay sites, and
the regex contract), @lubosxyz (#49224 — replace-not-strip semantics
and 'fn' fallback to avoid the empty-name trap).

Fixes #31666
2026-08-26 12:58:35 +05:30
Xipong 9ceb0858ab fix(codex): defang reserved Harmony tokens in requests 2026-07-31 22:53:20 -07:00
Teknium 6b81590c55 test: prune low-value tests suite-wide (wave 1) — 46,820 → 28,106 test functions
Systematic prune per AGENTS.md test policy, one pass over every major
test tree (gateway, hermes_cli, tools, agent, run_agent, plugins, cli,
cron, tui_gateway, honcho/openviking, root-level):

- DELETE: source-reading tests (read_text/getsource on prod files),
  change-detector tests (exact catalog counts, model-name snapshots,
  config version literals), mock-echo tests (assert a mock returns what
  it was told), assertion-free/trivial tests, near-duplicate
  parametrizations (boundaries + one representative kept), async/sync
  twin duplicates, cosmetic within-file variations.
- KEEP (mandatory): security/redaction/approval guards, message-role
  alternation invariants, prompt-caching/deterministic-call-id
  invariants, issue-number regression tests (deduped), E2E tests.
- 6 test files deleted outright (script-style/no-assert or fully
  redundant); conftest.py, fakes/, fixtures/ untouched.
- tests/acp/conftest.py added: autouse fixture stubs the live
  models.dev/GitHub/Copilot/Anthropic inventory fetches that ACP server
  tests performed on every session create — test_server.py 147s → 3.4s,
  and the tests are now genuinely hermetic.
- Sleep-based slowness shrunk where safe (codex_ttfb_watchdog,
  compression_concurrent_fork, etc.); no wall-clock assertion tightened.

Verification: full hermetic suite via scripts/run_tests.sh —
2439 files, 31,130 tests passed, 0 failed, 0 flaky retries, 315s wall
(baseline: 583s wall, 13,564s subprocess CPU).
2026-07-29 13:10:23 -07:00
PRATHAMESH75 e45f2b39e2 fix(codex): clamp oversized Responses call_id so MCP tools don't brick sessions (#73492)
The codex app-server namespaces MCP tool call ids as
codex_mcp__<server>__<tool>_<codex_call_id>. With an exec-<uuid> component the
built-in hermes-tools server alone overflows the Responses API's 64-char
call_id limit, so the request 400s with a non-retryable "string too long".
The offending item sits near the head of the transcript and replays every
turn, permanently bricking the session — the only recovery is /reset.

Sibling defect to #10788, which clamped input[*].id via
_MAX_RESPONSES_ITEM_ID_LENGTH. Apply the same treatment to call_id at both
Responses emit sites in _chat_messages_to_responses_input: a deterministic
surrogate (call_ + sha256[:32]) for ids over the limit, short ids unchanged.
Because the surrogate is a pure function of the original id, a function_call
and its matching function_call_output — which carry the same original id — map
to the same surrogate and stay paired without correlating the two items.
2026-07-29 00:47:54 +05:30
Sk fe1ab949fd fix(agent): treat Codex incomplete content filter as refusal
Map Codex Responses status=incomplete with incomplete_details.reason=content_filter to finish_reason=content_filter so the existing refusal/fallback path runs instead of burning incomplete continuation attempts.
2026-07-15 09:47:37 -07:00
Teknium 8fa8aabbbb test(codex): pin codex_backend issuer in xai-scoped salvage test (#64844)
test_normalize_codex_response_salvage_is_xai_scoped broke on main when
two same-day merges crossed: #64764 (#64434 — trust response.status for
reasoning-only turns on UNRECOGNIZED Responses backends) changed what a
bare _normalize_codex_response(response) call returns for
status='completed' reasoning-only output (now 'stop'), while #64768
added this test calling with no issuer_kind and expecting 'incomplete'.

The test's intent is that the xAI reasoning-channel salvage does not
leak into other special-cased backends — pin issuer_kind='codex_backend'
so it exercises exactly that (same pattern as
test_normalize_codex_response_treats_summary_only_reasoning_as_incomplete,
which was already pinned for #64434).
2026-07-15 04:27:32 -07:00
Ignacio Pastor 05d1ca549b fix(codex): rescue reasoning-only turns that die with 'remained incomplete after 3 continuation attempts'
grok-4.x on the xAI /v1/responses surface sometimes ends a turn with only
reasoning items — no message output item, no tool calls — and those
reasoning items carry no encrypted_content. Two compounding problems:

1. The model occasionally emits its final answer INSIDE the reasoning
   channel, delimited by grok's internal "<response>" tag. The answer
   exists but is classified reasoning-only → finish_reason=incomplete.

2. An interim assistant message holding only plain-text reasoning replays
   as nothing in _chat_messages_to_responses_input, so every continuation
   request is byte-identical to the one that just failed. The model
   deterministically repeats the reasoning-only response until the retry
   budget is exhausted and the turn dies with "Codex response remained
   incomplete after 3 continuation attempts".

Fixes:
- _normalize_codex_response (xai_responses only): salvage the
  <response>-delimited tail from the reasoning text and promote it to
  assistant content; the untagged prefix stays as thinking text.
- Codex-incomplete continuation path: when the interim message has
  nothing the input converter will replay (no content, no encrypted
  reasoning items, no message items), append a user-role nudge so the
  retry actually differs and explicitly asks for the final answer /
  pending tool call. Mirrors the existing _get_continuation_prompt
  pattern used for length truncation.

Observed live with grok-4.20 on xai-oauth (2026-07-13); sibling of the
grok-composer web_search incomplete-loop fix in transports/codex.py.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 00:13:24 -07:00
Teknium 07443ea21a test(codex): pin codex_backend issuer in summary-only reasoning sibling test
The #64434 change makes unrecognized issuers trust
response.status='completed' for reasoning-only turns, so this sibling
test (which exercised the old default-path behavior) now pins the Codex
backend explicitly — the surface where reasoning-only still means
'still thinking'.
2026-07-15 00:12:40 -07:00
kshitijk4poor bce17bf6a2 fix(codex): enforce Copilot replay policy at dispatch
Reapply the endpoint-aware preflight after request and execution
middleware so no override can reintroduce a connection-scoped ID.
2026-07-11 12:09:27 +05:30
joaomarcos 0b2907f586 fix(codex): drop oversized message ids on Responses input replay
Codex assigns assistant message items server-side ids that can run
400+ chars (base64 encrypted blobs), but the Responses API caps
input[].id at 64 chars and rejects the whole request with a
non-retryable HTTP 400. Once a session captures one of these long
ids, every subsequent turn replays it and 400s forever, since the
history persists it in codex_message_items.

Add a 64-char length guard at both replay sites — the history-to-
input converter and the final preflight gate — so oversized ids are
dropped while short ids (msg_...) are kept for prefix-cache hits.
Mirrors the existing pattern for reasoning items, which already
strip their id before replay because store=False means the API
can't resolve ids server-side anyway.

Fixes #27038

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-10 19:26:08 -07:00
XVVH 6f89e17a33 fix(xai): OAuth Responses native web_search, incomplete guard, grok-composer context
- model_metadata: grok-composer-2.5-fast → 262144 (OAuth slug not in /v1/models)
- codex transport: inject native {"type":"web_search"} for is_xai_responses;
  drop client web_search to avoid duplicate-name 400s
- codex adapter: do not treat in-progress server-side *_call items as incomplete
- tests: adapter, transport build_kwargs, model_metadata, oauth recovery
2026-06-17 17:33:32 -07:00
Teknium a4d8f0f62a feat(prompt): universal task-completion guidance + local Python toolchain probe (#34340)
* fix(codex): surface error code in Responses 'failed' status errors

When a Codex Responses turn ends with status=failed, the response carries
the failure details under `response.error` as
`{code, message, param, ...}`. The previous extractor pulled only
`message`, so users seeing a rate-limit failure got a bare "Slow down"
string indistinguishable from a generic stream truncation; an
internal_error with empty message degraded to a dict dump
("{'code': 'internal_error', 'message': ''}").

Extract a `_format_responses_error()` helper that:
- prefixes `code` when both code and message are present
  (e.g. 'rate_limit_exceeded: Slow down')
- falls back to the bare `code` when message is empty
- accepts both dict and attribute-style payloads (SDK and JSON-RPC paths)
- preserves the prior status-only fallback when no error payload exists

Apply the same helper at the sibling site in
`codex_app_server_session.run_turn()` so codex-CLI subprocess turn
failures get the same treatment.

Tests:
- 8 new unit tests for `_format_responses_error` covering both shapes,
  empty/missing fields, non-string fields, and the status-only fallback.
- 2 regression tests on `_normalize_codex_response` for failed status
  with and without a code, asserting the exact RuntimeError message.
- All 3603 tests in tests/agent/ pass.

Adapted from anomalyco/opencode#28757.

* feat(prompt): universal task-completion guidance + local Python toolchain probe

Two cross-model failure modes get a single-line answer in the cached
system prompt. Both gated by config (default on), both add zero overhead
when not needed, both verified via real AIAgent prompt builds.

## What changed

`TASK_COMPLETION_GUIDANCE` — short prompt block applied to ALL models.
Targets two failure modes observed on a real Sarasota real-estate build
task: (1) Opus stopped after writing an 85-byte stub and gave a prose
response with finish_reason=stop on call #3 of 90; (2) DeepSeek pushed
through a PEP-668 wall, then returned fabricated listings instead of
admitting the blocker. Both behaviors are model-family-agnostic, so the
guidance lives outside the existing tool_use_enforcement gate (~192
tokens, paid once per session via prefix cache).

`tools/env_probe.py` — local Python toolchain probe. Detects
python3/pip/uv/PEP-668 state and emits ONE short line in the system
prompt when something is non-default. Emits NOTHING when the env is
clean (zero token cost for normal users). Skipped entirely for remote
terminal backends (docker/modal/ssh) — they have their own probe.

Example output on a broken environment (the actual case):

    Python toolchain: python3=3.11.15 (no pip module),
    python=missing (use python3), pip→python3.12 (mismatch),
    PEP 668=yes (use venv or uv).

## Config

Both flags live under `agent.` in config.yaml, default True:

    agent:
      task_completion_guidance: true   # universal "finish the job" block
      environment_probe: true          # local Python toolchain hints

Neither addition required a `_config_version` bump — deep-merge fills
defaults in for existing user configs.

## Validation

| Test surface | Result |
|---|---|
| tests/tools/test_env_probe.py | 10/10 pass (probe unit) |
| tests/run_agent/test_run_agent.py — new classes | 8/8 pass (integration) |
| TestToolUseEnforcementConfig | 17/17 pass (no regression) |
| TestBuildSystemPrompt | 9/9 pass (no regression) |
| TestInvalidateSystemPrompt | 2/2 pass (no regression) |
| tests/agent/test_prompt_builder.py | 124/124 pass (no regression) |
| tests/hermes_cli/ | 5662/5662 pass (config defaults) |
| E2E AIAgent build (broken env) | Both blocks present, 2,178 chars |
| E2E AIAgent build (clean env) | 771-char net overhead, env probe silent |
2026-05-28 22:26:09 -07:00
Krishna b1a46b3047 fix(codex): drop transient rs_tmp reasoning replay state 2026-05-27 02:25:59 -07:00