Commit Graph

214 Commits

Author SHA1 Message Date
Teknium eb67765c58 refactor(agent): agent_runtime_helpers — drop dead predicates, dedupe runtime restore/switch/recovery, compact narratives
- Dead: agent_runtime_owns_post_tool_hook, intent_ack_continuation_enabled (only their
  own tests referenced them; tests removed).
- invoke_tool routes inline tools via INLINE_TOOL_EXECUTORS.
- switch_model normalizes provider names once (was 5x); restore_primary_runtime shares
  primary-pool load/match helpers; _apply_primary_runtime_fields and
  _build_anthropic_client_from_runtime shared by transport recovery and turn-start
  restore; recover_with_credential_pool rotate-and-swap helper (4 sites).
- Incident-narrative comments/docstrings compacted; rules, orderings, invariants kept.
5266 -> 3837 LOC.
2026-09-02 13:29:31 -07:00
Alexander Prendota 1131b22856 feat(providers): let a provider profile supply its own client
``create_openai_client`` was a hardcoded if-ladder: copilot-acp builds an ACP
stdio shim, gemini builds a native client, everything else gets an
``openai.OpenAI``. There was no extension point, so a provider whose wire
protocol is not OpenAI-over-HTTP could only be added by editing this function —
which is exactly why an ACP provider cannot ship outside this tree today, even
though ``providers/__init__.py`` has discovered out-of-tree profiles from
``~/.hermes/plugins/model-providers/`` and pip entry points for a while.

``ProviderProfile.create_client(**client_kwargs)`` closes that gap. It returns
``None`` by default, so every provider that wants the standard client is
unaffected and the existing ladder still runs as the fallback. copilot-acp is
migrated onto it — its hardcoded branch is gone and its profile supplies the
client in three lines, which is the same three lines an external package writes.

Resolution goes by provider name first, then by ``base_url`` prefix, so a
runtime configured only by URL still reaches its profile — matching what the
replaced ``startswith("acp://copilot")`` branch did. A profile that raises is
logged and skipped: a third-party plugin can fail to provide a client, but it
cannot take the turn down.

Also replaces the two ``isinstance`` checks in ``agent/auxiliary_client.py``
that mean "this client is complete, do not wrap it" with capability flags the
client class declares — ``HERMES_SKIP_TRANSPORT_WRAP`` and
``HERMES_SKIP_ASYNC_WRAP``, mirroring ``SUPPORTS_HERMES_TOOL_CALLS`` in
``background_review.py``. Two in-tree consumers (the ACP shim and the Gemini
native client), an out-of-tree client is covered by the same declaration, and
the hot path no longer imports those modules just to type-test.

Co-Authored-By: Junie <junie@jetbrains.com>
2026-09-02 09:57:39 -07:00
Teknium bd7cdd7c53 Merge origin/main into core-tool-deferral (resolve show_tip test seam onto the check_tips_enabled gate) 2026-09-01 21:49:14 -07:00
Yong Li f41ed09b51 fix(gemini): strip call ids on insert, name the realignments in the log
Review follow-ups:

- Strip the tool_call id when populating the call-name map so it matches the
  stripped lookup (and pass 1's result_call_ids). A padded id previously
  skipped realignment silently.
- Log which names were rewritten, not just how many.
- Note in the comment that a result whose assistant call frame was pruned is
  already dropped by the orphan pass, so it cannot reach the provider with a
  stale name; cover that with a test.
- Rename test_sanitize_leaves_matching_and_unpaired_tool_result_names_alone,
  which only ever exercised the matching case, and add the padded-id case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 14:56:30 +05:30
Yong Li 2e9435d2f6 fix(gemini): echo bridged tool_call name on the OpenAI-compatible path
Google matches functionResponse.name against functionCall.name and rejects
a mismatch with HTTP 400 INVALID_ARGUMENT. #72089 fixed this for the native
Gemini adapter, where _translate_tool_result_to_gemini() now prefers
tool_name_by_call_id over the result message's internal name.

Requests that reach Gemini through an OpenAI-compatible gateway (OpenRouter,
Vertex/LiteLLM proxies) never run that translation, so they still put the
unwrapped internal tool name on the wire: the model calls the tool_search
bridge tool `tool_call`, make_tool_result_message() labels the result
`mcp__strava__get_recent_activities`, and the next turn 400s with a bare
"Provider returned error". The bad pair stays in the transcript, so every
later request in that session fails too.

Hold the same invariant at the final pre-API chokepoint instead of in the
OpenAI-compat serializer: Gemini arrives under many model strings and base
URLs, so sniffing for "is this really Google?" is unreliable, while every
other provider either ignores the field or already agrees with the call
name. Only a name that is present and disagrees is rewritten, so clean
transcripts still pass through byte-identical for prompt caching, and the
rewrite lands on the per-call copy so the stored trajectory keeps the real
tool name for the session DB and UI. No-op for the native Gemini path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 14:56:30 +05:30
Teknium fb9b2c893f feat(agent): escalate repeated transcript-sanitiser heals with a one-time user notice (#96870)
Builds the escalation layer on top of HexLab98's heal-log windowing
(salvaged from PR #96916):

- Per-session heal counters (heal events + messages healed) tracked by the
  repair path in agent_runtime_helpers.py, session totals preserved across
  10-minute log windows.
- Threshold escalation: after N heals in a session window (default 3,
  configurable via agent.sanitizer_heal_escalation_threshold in
  config.yaml, 0 = off) log ONE ERROR carrying session id + heal pattern
  (events/messages/window/threshold), then stay quiet.
- ONE-TIME out-of-band user notice queued at the threshold and delivered by
  the conversation loop through _emit_warning (status callback -> gateway
  status message / CLI print). Never injected into conversation context or
  the wire copy: prompt caching, role alternation, and durable history are
  untouched. Never re-arms on a new window; scoped per session.
- Counters visible in diagnostics: get_sanitizer_heal_stats() rendered in
  the /debug share // hermes debug report, and the config key surfaced in
  hermes dump overrides. errors.log carries the ERROR line for `hermes logs
  errors`.
2026-08-31 13:11:41 -07:00
HexLab98 20fd5d0b25 fix(agent): stop empty-transcript sanitizer from warning on every send (#96870)
Fill empty non-final user/assistant turns on the wire copy during send-time projection so the sanitizer does not re-heal the same poisoned row every call. When a caller still hits the owner, log WARNING then one ERROR per session window instead of flooding errors.log.
2026-08-31 13:11:41 -07:00
fangliquanflq fd1d8271db fix(cron): isolate lazy imports from stale modules 2026-08-31 09:58:51 -07:00
Teknium 6101f52ba4 Merge remote-tracking branch 'origin/main' into core-tool-deferral 2026-08-30 19:47:30 -07:00
Stephen Chin 48a4201f40 fix(compaction): preserve switch compatibility fixtures
Keep model-switch callers compatible with result objects created before runtime_capabilities was added, and do not roll back minimal agents that lack optional LM Studio helpers. Preserve rollback for real helper failures.
2026-08-30 05:16:10 -07:00
Stephen Chin 5247a6f07f fix(compaction): clarify runtime capability state
Use a distinct runtime_capabilities field on agents, preserve compatibility with earlier snapshots, and resolve the canonical direct OpenAI endpoint when a cross-provider switch omits base_url. Keep ambiguous proxy routes fail-closed.
2026-08-30 05:16:10 -07:00
Stephen Chin 903c36b6d4 fix(compaction): resolve capability from effective switch URL 2026-08-30 05:16:10 -07:00
Stephen Chin 08c7879ca1 fix(compaction): preserve native capability across runtime switches
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
2026-08-30 05:16:10 -07:00
Teknium 3b3ad958d7 fix(runtime): key-scoped fallback extra_body re-resolution + request_overrides in switch_model snapshot
Follow-up hardening on the two cherry-picked contributor commits:

- try_activate_fallback: replace the blanket request_overrides.pop('extra_body')
  with KEY-SCOPED removal — only keys the OLD provider's custom_providers
  entry contributed (value unchanged since the init-time merge) are dropped.
  Caller/profile-provided extra_body keys survive the swap, matching the
  caller-over-provider precedence in agent_init._merge_custom_provider_extra_body.
  The fallback provider's own extra_body is then merged back in.
- switch_model: the live _primary_runtime snapshot it rebuilds now carries
  request_overrides, so a post-switch transport recovery or fallback restore
  reinstates the switched-to identity's overrides instead of dropping them.
- Tests: activation-level stale-key removal + caller-override preservation
  (test_provider_fallback.py), switch-then-recover / switch-then-restore
  (test_primary_runtime_restore.py).

Cache-safety: none of these paths mutate past context or rebuild the system
prompt — only outbound request kwargs change.

Fixes #75091
2026-08-29 19:13:12 -07:00
Adam Fortuna 131501229a fix(runtime): restore request_overrides after transport recovery
Include request_overrides in primary runtime snapshots so transport recovery restores request-level model parameters.
2026-08-29 19:13:12 -07:00
Jack b10b27e6f9 fix(agent): match switched-to custom provider by model+base_url, not name
Addresses the hermes-sweeper review on #53765. The in-place /model switch
helper (_apply_switched_provider_request_overrides) derived a custom
provider's extra_body by provider *name* only, while build-time matching in
agent_init._merge_custom_provider_extra_body matches by provider key, base_url,
AND model. So a different model selected at the same named endpoint could
inherit an extra_body configured for another model.

Reuse the shared agent_init._custom_provider_extra_body_for_agent matcher
(provider key + base_url + model), sourcing custom_providers from the
init-time agent._custom_providers cache (fresh-load fallback if absent). A
stale extra_body is always cleared when no entry matches; non-provider
overrides (service_tier / speed from /fast) are preserved.

Tests: add nonmatching-model and endpoint-mismatch regressions; update the
existing switch tests onto the model/base_url-aware matcher.
2026-08-29 19:13:00 -07:00
Jack d2af990043 fix(agent): carry request_overrides through in-place /model switch (TUI/CLI)
Third in the series. The gateway rebuild path (previous two commits)
carries a custom provider's `request_overrides` (`extra_body`, e.g.
`chat_template_kwargs`) into the agent, but the *in-place* live switch used
by the TUI dashboard and the CLI — `agent.switch_model()` ->
`agent_runtime_helpers.switch_model()` — swapped
model/provider/base_url/api_key without ever updating `request_overrides`.
So a `/model` switch to a thinking-enabled custom provider in the TUI/CLI
kept the previous provider's `extra_body`.

`switch_model()` now re-derives the switched-to provider's
`request_overrides` (via `_get_named_custom_provider`) and applies it in
place, preserving non-provider overrides (`service_tier`/`speed` from
`/fast`). Logic factored into `_apply_switched_provider_request_overrides`
for testability.

Adds tests/agent/test_switch_model_request_overrides.py.
2026-08-29 19:13:00 -07:00
Teknium c1762ff11c test: sweep sibling tests stale on the tool renames + deferral default
The rename sweep in the base commit missed the sibling-test blast radius
(18 red files on CI). Three classes, all fixed:

1. Stale old names in tests (todo/cronjob/process/tour/tip) — updated to
   todo_list/cronjob_manage/process_manage/gui_tour/show_tip at every
   registry.get_entry/dispatch/coerce/preview/allowlist call site, plus
   the coding-brief sentence in agent/coding_context.py now names
   todo_list (and its gating test).
2. Missed rename in production: AGENT_RUNTIME_POST_HOOK_TOOL_NAMES still
   held 'tour' — post-hook ownership would have double-emitted for
   gui_tour via the bridge path.
3. Tests pinning pre-deferral assembly (blank-slate surface, modal
   sandbox resolution, desktop diet, HUD note) now pin their ACTUAL
   contract under the legacy defer:[] override, or assert on granted
   tool names instead of visible schemas.

Also fixes a pre-existing ordering flake surfaced by the sweep:
test_holds_exactly_the_gui_affordances depended on whether an earlier
test had imported apply_layout_tool (registry-registered, not in the
static desktop_ui list) — now forces discovery and pins the full set.

649 tests green locally across all touched files, both orderings.
2026-08-29 18:23:07 -07:00
Teknium e16ad33a9d feat(tool-search): core-tool deferral — curated 19-tool set behind the bridge by default; renames todo_list/cronjob_manage/process_manage/gui_tour/show_tip with legacy aliases (13.4K -> 6.9K desktop schemas, -49%) 2026-08-29 08:26:24 -07:00
Teknium 1d8946b40b fix(prompt-caching): tool-using sessions no longer 400 behind LiteLLM Anthropic proxies (#89886)
LiteLLM OpenAI->Anthropic translation copies tool-message content parts
verbatim, so the envelope-layout part-level cache_control landed at
tool_result.content[0] - a placement the Anthropic Messages schema rejects
with a non-retryable HTTP 400 that killed the whole turn (any tool-using
cron/session on a LiteLLM-fronted Anthropic route).

New envelope_tool_part_cache_markers_supported() predicate (keyed on the
existing _is_litellm_route token matcher) threads a tool_part_markers flag
through build_prompt_cache_plan / apply_anthropic_cache_control and all
four decoration sites (main loop x2, destination replan, MoA). On LiteLLM
routes role:tool messages carry no markers and the breakpoint budget
reallocates to the nearest eligible message; OpenRouter/Nous Portal keep
the part-level form they honor, native Anthropic layout unchanged.
2026-08-29 12:06:48 +05:30
Teknium 217ab2f8df refactor(desktop-tools): consolidate preview + project, diet the desktop_ui suite (3,861 → 2,293 tok/call, −41%) (#97659)
* refactor(desktop-tools): consolidate preview(open/close/read) + project(create/switch/list), diet the desktop_ui suite — 3,861 -> 2,293 tok/call on desktop sessions (-41%)

* rename: preview -> desktop_preview, project -> desktop_project — namespace desktop-app tools against MCP/plugin name collisions

* test: sync remaining old-name pins — per-file registration import, GUI_TOOLS set, post-hook case read_preview -> desktop_preview action=read
2026-08-28 23:10:01 -07:00
Teknium 5b31602c15 docs: reconcile positional pairing with shared _classify_tool_call_orphans (#97167) — classifier docstring reflects its remaining consumer; empty-id filter note updated 2026-08-28 07:51:23 -07:00
fedebyes 93f4dc7561 fix: make positional prune variant-aware; add replayed-call regression tests
Pass 2 of repair_message_sequence matched results only by id/call_id,
pruning calls answered through response_item_id or composite bridge
ids. Use the shared variant helpers (tool_call_id_variants /
tool_result_id_variants) so the unified alias policy applies
(#55626/#63000/#93251).

The positional sanitizer pass changes the crash/resume duplicate shape:
an interrupted first occurrence is now stubbed instead of deduped, so
the replayed call survives with its own immediate result. Update the
#64335 empty-key test to the new semantics and add regression tests for
the #94704 acceptance shape (historical-result + replayed-call +
fresh-call) and the production interrupted-turn shape (session
7d57a602b83d).
2026-08-28 07:51:23 -07:00
Tiberiu Danciu c7761573f5 fix: prune positionally unanswered tool_calls before API send
DeepSeek v4 rejects a payload where an assistant message carries a
tool_call whose tool result does not follow it immediately (HTTP 400
"An assistant message with 'tool_calls' must be followed by tool
messages responding to each 'tool_call_id'"). Context compression can
displace a tool result past a user turn; the result then lands ~100
messages away from its declaring assistant message.

Two gaps let the poisoned shape reach the wire (reproduced from the
production request dump of session 4d8727cbcf04, replayed through both
functions):

1. repair_message_sequence Pass 1 drops the displaced tool RESULT as
   stray but leaves the declaring assistant message carrying the now
   unanswered tool_call (with empty content) in the durable history.
2. sanitize_api_messages stubbed only globally-absent result ids: the
   displaced result still exists in the transcript, so the id survives
   the set-subtraction, no stub is injected, and the payload 400s.

Fix both layers so every path is order-independent:

- repair_message_sequence: new Pass 2 prunes tool_calls that have no
  result in the immediately-following tool run (matching on id or
  call_id, same superset rule as Pass 1). If pruning empties the turn
  (no content/reasoning left), the whole message is dropped rather than
  sending an empty assistant message. Codex interim turns are exempt,
  as in Pass 0.
- sanitize_api_messages: the orphan/stub logic is rewritten as a single
  rolling positional walk that drops results not immediately following
  their declaring assistant (including results appearing BEFORE their
  call) and injects stub results for positionally-uncovered calls even
  when a mispositioned result exists elsewhere.

Adds six regression tests: repair pruning, whole-turn drop when pruned
calls were the only payload, valid-pair negative control, positional
stub injection, result-before-call orphan drop, and a fully-paired
transcript negative control.
2026-08-28 07:51:23 -07:00
Hermes Agent 225fa13bd3 fix(sanitizer): drop duplicated legacy _classify_tool_call_orphans left by cherry-pick auto-merge 2026-08-28 06:32:48 -07:00
isheng c6a426e9ad refactor(sanitizer): extract shared _classify_tool_call_orphans to eliminate drift
sanitize_api_messages (agent_runtime_helpers) and
_sanitize_tool_pairs (context_compressor) both collected
tool-call IDs and classified orphans with near-identical logic
that had already drifted: the canonical sanitizer added dedup
(#58350), but the compressor's copy did not.

Extract the shared orphan-detection logic into
_classify_tool_call_orphans(messages) in agent_runtime_helpers.
Both call sites now delegate to it, preserving their divergent
remediation strategies (insert-stubs vs strip-orphans) while
ensuring id-resolution rules and dedup stay in sync.

Closes #58357
2026-08-28 06:32:48 -07:00
joaomarcos f0ac2c8f12 fix(agent): drop stale api_content sidecar and unpaired tool results
Rebased onto current main to drop the empty-tool_calls fix (already on
main via #86654, cherry-picked from #77944 with @webtecnica's
authorship). This PR now carries only the two fixes unique to it:

1. A pre-existing api_content sidecar left stale on the consecutive-
   assistant merge. The sidecar takes priority over content at
   API-build time, so a merge could silently discard its own freshly
   concatenated content on the next call. Only dropped when the merge
   actually changes the resulting value (wz-heng, #78063 review) --
   content_rewritten compares before/after value, not just whether an
   assignment branch fired, so a falsy new_content (e.g. "") that
   strips to nothing no longer trips a spurious sidecar drop.

2. sanitize_api_messages never flagged a tool result with a missing/
   empty tool_call_id -- its orphan-detection set only ever collected
   truthy ids, so an unpaired result with no id passed the final
   chokepoint untouched.

Addresses teknium1's rebase request and wz-heng's review findings on
2026-08-28 06:32:48 -07:00
Ailirag 5908c577f9 fix(fallback): surface provider transitions and primary recovery 2026-08-25 12:12:08 +05:30
Vignesh Ramesh a0795acc83 fix(codex): identify Hermes requests 2026-08-24 11:25:04 -07:00
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
kshitijk4poor f93b350711 fix: align cache-policy pre-gate identity with the capability matcher
Follow-ups on top of the salvaged #92785 commit:

- Pre-gate now matches base URLs via normalize_route_base_url and
  provider ids via custom_provider_aliases, mirroring the semantics of
  get_custom_provider_model_capability. The raw string comparison
  silently dropped declarations whose config spelling differed only by
  host case or trailing slash (proven empirically: …/v1/ vs …/v1 with a
  non-matching provider name returned (False, False) despite an explicit
  prompt_caching: true).
- get_provider(..., allow_network=False) in the early-init/stub branch:
  the policy runs per request destination (MoA aggregator, auxiliary
  replans via blank_cache_policy_stub, early agent init) and a cold
  models.dev cache triggered a measured ~450 ms foreground registry
  fetch from the send path. A catalog miss degrades to the conservative
  side.
- Debug-log the previously silent provider-lookup exception fallback.
- Tests: _make_agent defaults _custom_providers=[] (post-init reality;
  keeps built-in-route tests off the catalog/config fallback), the two
  early-init tests delete the attr explicitly, and three regression
  tests pin the URL-drift, spaced-legacy-name, and no-network contracts
  (all three fail on the unfixed commit).
2026-08-24 14:58:42 +05:30
Blood Shot 0204e4898e fix(agent): normalize custom provider route identity 2026-08-24 14:58:42 +05:30
Blood Shot 0a3b7efec5 fix(agent): honor prompt_caching for custom providers
Apply explicit per-model prompt_caching capabilities to custom
chat-completions routes, rather than limiting them to recognized providers,
hosts, or model families.

Keep undeclared routes conservative, derive the marker layout from the wire
transport, and leave Responses and Bedrock caching paths unchanged.
2026-08-24 14:58:42 +05:30
fangliquanflq 37411f349a fix(auth): rotate credentials for named custom providers after 401/429
Salvage of #93214 (5 commits squashed onto current main; agent_runtime_helpers.py
diverged since the PR base and was 3-way reapplied). The credential-rotation
guard in recover_with_credential_pool and both restore_primary_runtime paths
only tolerated the custom-naming split when the agent carried the literal label
'custom', so a named custom provider (agent.provider='gemini-no-filter', pool
'custom:gemini-no-filter') tripped the mismatch guard and skipped rotation on
every 401/429. Now all three guard sites use the canonical
credential_pool_matches_provider boundary predicate + resolve_runtime_pool_key,
which recognizes configured named-custom aliases and validates endpoints.

Fixes #93188.
2026-08-23 20:01:18 -07:00
joaomarcos 5496d5995a fix(agent): preserve tool results across ID variants
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
2026-08-23 18:24:43 -07:00
Teknium faa2399e2b fix(agent): make the pre-call dedup pass variant-aware; widen batch regression coverage (#93251)
Follow-up on top of the salvaged cluster: sanitize_api_messages step 3
(duplicate tool_call_id dedup) still tracked only the coalesced
(call_id||id) value in outstanding_call_ids, so after step 2's
variant-aware matching preserved a result keyed on the OTHER id variant,
step 3 deleted it as answering no outstanding call — whole parallel
batches of real results vanished with no stub at all (#93251's total-loss
mode). Track the full variant set per call and consume all siblings when
answered, preserving #58327 duplicate protection and llama.cpp
constant-id re-arm semantics.

Also aligns the #58287 compressor test with the in-flight tool chain
protection (#79278) that landed after that PR was opened: a trailing
user turn keeps the negative-control assistant message out of the
protected trailing window.

New regression tests: divergent-id batch survival through the dedup
pass, sibling-id replay still dropped, constant-id re-arm preserved.
Sabotage-verified: tests fail with the old single-id tracking.
2026-08-23 17:01:20 -07:00
joaomarcos 36b4da5489 fix: repair_message_sequence drops tool results for SDK tool_call objects
The tool_call id-matching pass in repair_message_sequence only read
`.get("id"/"call_id")` on plain dicts, skipping non-dict tool_calls
entirely (`if not isinstance(tc, dict): continue`). Host-fed and
pre-serialization histories can carry unserialized SDK tool_call
objects (e.g. `ChatCompletionMessageToolCall`) instead of dicts, which
left `known_tool_ids` empty for that assistant turn. The following
`tool` message — a legitimate result already produced by executing the
tool — was then misclassified as an orphan and silently dropped,
corrupting the persisted conversation history and leaving the
assistant's tool_calls unanswered (itself a trigger for HTTP 400 on
strict providers).

Fix: extract id/call_id via getattr() for non-dict entries too,
mirroring AIAgent._get_tool_call_id_static's existing dict-or-object
tolerance, instead of skipping them.
2026-08-23 17:01:20 -07:00
Frowtek b9a62f6590 fix(agent): consume every tool_call id variant when pairing tool results
`repair_message_sequence` registers BOTH `id` and `call_id` for each
assistant tool_call, because a matching tool result may be keyed on either
depending on which path built it (#58168). The duplicate guard added for
dropped rather than replayed.

Those two behaviours don't compose: a Codex/Responses tool_call registers
two DIFFERENT ids (`fc_...` and `call_...`), but only the id the first
result referenced is discarded. Its sibling stays in `known_tool_ids`, so a
duplicate result keyed on that sibling still matches and is kept — two tool
messages replayed for one call, which is exactly the HTTP 400 on strict
providers the consume step exists to prevent.

Duplicates of this kind come from the retry / crash / session-resume glitch
the guard was written for; the id-variant split just lets them slip past it.

Track each registered id back to its tool_call's full variant set and
discard all of them on a match. Results keyed on either variant are still
accepted (no false orphaning), and two parallel Codex calls answered via
different variants both survive.

Adds regression tests for the sibling-keyed duplicate and for the
two-calls/mixed-keys case that must NOT be affected.
2026-08-23 17:01:20 -07:00
Bartok9 1a83b1e588 fix(agent): keep tool results keyed on a tool_call's id variant (#55626)
Register every id variant (call_id AND id) of each assistant tool_call in
sanitize_api_messages so a tool result keyed on either variant is treated
as paired. Previously only the coalesced (call_id||id) value was
registered, so Responses-style tool_calls carrying divergent id (fc_...)
and call_id (call_...) had their real results dropped as orphans and
replaced with '[Result unavailable]' stubs.

Cherry-picked from PR #56148 (unrelated busy_ack_templates files dropped
per the author's own follow-up commit).
2026-08-23 17:01:20 -07:00
poisdahl abf87e7248 Merge current main into composite-carrier fix 2026-08-21 15:56:45 +02:00
Teknium ca06b87689 feat: opencode-free is fully keyless — no env var, no account, anonymous wire
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).

On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
  never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
  Zen endpoint routing incl. muse->responses); keyless predicate extended
  with unsuffixed free slugs (big-pickle); free runtime pins EVERY
  opencode-free model keyless; curated catalog replaces the models.dev
  cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
  there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
  wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
  invariant (every curated model must satisfy the keyless predicate)

E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
2026-08-21 00:24:32 -07:00
Rudraksh Chahal 28a9b6c565 feat(providers): add OpenCode Free provider with keyed auth and opencode User-Agent
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.

The free tier requires a real account API key and throttles third-party
clients by User-Agent:

- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
  and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
  Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
  unconditionally (the stale keyless-tier assumption), and credential-pool
  exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
  message.

Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
2026-08-21 00:24:32 -07:00
Teknium a83c3915a3 fix(opencode): family-wide provider predicate + reserved tool-name aliases for custom opencode-* providers
Builds on @Lesnak1's #85619 (issue #85589):

- New opencode_provider_family() single-owner predicate in
  hermes_cli/models.py — resolves built-in AND custom family providers
  (opencode-go-bridge, OpenCode-Zen-Custom, ...) case-insensitively.
  Migrated all 8 inlined family checks (models.py x3, runtime_provider.py
  x4 from the salvaged commits) plus 4 sibling sites the PR missed:
  cli.py api_mode sync, agent_runtime_helpers.py double-/v1 guard,
  model_normalize.py flat-namespace strip, model_switch.py base_url
  normalization.
- Responses transport: alias OpenCode-reserved function names
  (web_search, search_files -> hermes_*) on the wire and map them back on
  dispatch — same pattern as the xAI web_search collision fix. Matches
  family providers and any base_url on opencode.ai. Fixes the HTTP 400
  'custom function name X is reserved' half of #85589.
- Tests: custom-provider routing assertions + 5 new transport alias tests.
2026-08-20 20:21:12 -07:00
Brooklyn Nicholson c57581cd0d feat(tools): drive_preview and annotate_preview — the agent can use the page it opened
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.

Four pieces, and they only make sense together:

  · an in-page engine that inventories what is interactable and performs the
    verb, injected as source because it has to run inside the guest page;
  · the preview.act.request bridge from the gateway into the pane;
  · drive_preview, for acting: elements, click, type, scroll, press, and the
    pane's own back/forward/reload;
  · annotate_preview, for marking without acting.

Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.

Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.

Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
2026-08-20 05:26:37 -05:00
Teknium 449471c334 feat: runtime stall guards — identical-call loop breaker and continue-intent recovery (agent.stall_guards)
Composio eval traces showed Hermes wasting turns re-issuing identical tool
calls (same tool, same args, same result — 3x/4x in one run) and ending
turns by announcing an action it never took. Two conservative, config-gated
guards (agent.stall_guards, default true):

- Identical-call loop breaker: ToolCallGuardrailController.observe_identical_call
  tracks the consecutive streak of (tool, canonical args, result-hash); on
  the 3rd identical call a compact one-line notice is appended to that tool
  RESULT at construction time (cache-safe — tool results are append-only).
  Never blocks the call. Pollers (process, *_get_result, *_poll) are exempt
  via STALL_GUARD_REPEATABLE_TOOLS. Streak resets on any different call,
  changed result, or new turn. Observed on the raw result before the
  tool-loop warning suffix so its changing count can't defeat matching.

- Said-continue-but-stopped recovery: trailing_continue_intent() detects a
  short reply ENDING on an announced next action ('Let me now…', 'I will
  now…', 'Next, I…'); the conversation loop feeds it into the EXISTING
  intent-ack continuation path (same interim-assistant + user-nudge
  mechanism, same codex_ack_continuations cap of 2), preserving message
  alternation — no parallel recovery machinery.

Config: agent.stall_guards in DEFAULT_CONFIG; docs in configuration.md;
unit tests for streak/allowlist/reset/gate and detector pos/neg cases.
2026-08-19 16:34:21 -07:00
kshitij kapoor b2057c1685 refactor: extract duplicated load_config_readonly try/except into helper
The identical 6-line try/except block for reading model.reasoning_echo
from config appeared in both agent_init.py (init) and
agent_runtime_helpers.py (switch_model). Extracted into
AIAgent._read_reasoning_echo_from_config() static method — net -1 LOC.
2026-08-20 00:03:50 +05:30
Yingliang Zhang 663fa68cd4 fix: add reasoning_echo_flag to init snapshot and switch rollback
Address review feedback on PR #76503:

1. Init-time primary snapshot (agent_init.py:2756) was missing
   reasoning_echo_flag — after fallback recovery the flag was
   restored as False even when model.reasoning_echo: true was set.

2. Switch transaction snapshot (agent_runtime_helpers.py:2284) was
   missing _reasoning_echo_flag — a failed client rebuild during
   switch_model would leave the old provider with the new provider
   echo policy.

Both omissions now fixed. No test regressions (56 passed).

Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
2026-08-20 00:03:50 +05:30
Yingliang Zhang 73243b0d2e feat(config): per-provider reasoning_echo opt-in for custom providers
Add model.reasoning_echo (default false) and per-fallback-entry
reasoning_echo to preserve assistant reasoning_content when
replaying history to custom providers and OpenAI-compatible gateways
that proxy thinking-mode models (Kimi K3, GLM-5.2, DeepSeek, etc.)
but are not matched by the built-in host-based _REASONING_ECHO_RULES.

The flag is per-active-provider, not a global toggle:
- Primary: read from model.reasoning_echo at init and switch_model
- Fallback: set by try_activate_fallback from the fallback entry
- Restore: restore_primary_runtime copies the switch_model snapshot

Unlike PR #76019 global agent.reasoning_echo toggle, the
per-provider flag travels with the active provider — falling back to
a strict provider (Mistral, Groq, Cerebras) correctly strips
reasoning_content even when the primary had the flag enabled,
because the flag is False for the strict fallback.

Complements PR #27361 (dynamic detection) which fires after the first
API response; this PR covers turn-1 and history-replay-on-fresh-session
where dynamic detection has not fired yet.

Closes #76018
Refs: #27297, #27361, #76019

Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
2026-08-20 00:03:50 +05:30
Brooklyn Nicholson 23d88c2b0e feat(tools): tour — let the agent walk a user through the UI
One generic tool in the desktop_ui toolset: discover what is on screen,
highlight an element with narration, or hand the user a paged tour. No tour
content lives in the code — the agent authors each one live, which is what
makes 'how does this work?' answerable as a walkthrough instead of a wall
of text.

Rides the existing blocking-prompt bridge (tour.request/.respond) like
read_preview, so it works on every connection topology.
2026-08-19 00:52:54 -05:00
ethernet 879d6a4c78 feat(tui_gateway): batch clarify bridge with per-question locks
One clarify.request carries the question list (qid, question, choices,
multi_select per entry). clarify.respond gains an optional question_id:
each respond locks one answer, a repeat respond overwrites it, and the
batch resolves when every question is locked. A respond without
question_id keeps its existing meaning (cancel the whole prompt).

Locked answers survive the deadline: a timed-out batch returns the
partial answer map with a timed_out flag instead of an empty string.
The reconnect replay snapshot also carries the locked answers, so a
reattached client restores its per-question state.

Both agent-side clarify dispatch sites forward the questions arg.
2026-08-18 21:28:53 -04:00