Commit Graph

28271 Commits

Author SHA1 Message Date
fabiantax d3bfd2e9b1 feat(delegation): forward delegation.request_overrides on direct-endpoint branch
The direct base_url branch of _resolve_delegation_credentials returned no
request_overrides key, so a direct OpenRouter delegation (provider=custom,
base_url=openrouter.ai/api/v1) could not pass routing hints to its children.
The named-provider branch already forwards runtime request_overrides; this
gives the direct branch the same contract, honouring delegation.request_overrides
from config (dict → forwarded, anything else → None).

Primary use: extra_body.provider = {"sort": "throughput"} so delegation
children route to the fastest OpenRouter provider for their model, per the
fab-swarm throughput work (#901).
2026-08-29 19:13:23 -07:00
Teknium ac5186ed82 chore(release): map adamfortuna1324@gmail.com to 0xAdamFortuna 2026-08-29 19:13:12 -07:00
Teknium 3b3ad958d7 fix(runtime): key-scoped fallback extra_body re-resolution + request_overrides in switch_model snapshot
Follow-up hardening on the two cherry-picked contributor commits:

- try_activate_fallback: replace the blanket request_overrides.pop('extra_body')
  with KEY-SCOPED removal — only keys the OLD provider's custom_providers
  entry contributed (value unchanged since the init-time merge) are dropped.
  Caller/profile-provided extra_body keys survive the swap, matching the
  caller-over-provider precedence in agent_init._merge_custom_provider_extra_body.
  The fallback provider's own extra_body is then merged back in.
- switch_model: the live _primary_runtime snapshot it rebuilds now carries
  request_overrides, so a post-switch transport recovery or fallback restore
  reinstates the switched-to identity's overrides instead of dropping them.
- Tests: activation-level stale-key removal + caller-override preservation
  (test_provider_fallback.py), switch-then-recover / switch-then-restore
  (test_primary_runtime_restore.py).

Cache-safety: none of these paths mutate past context or rebuild the system
prompt — only outbound request kwargs change.

Fixes #75091
2026-08-29 19:13:12 -07:00
Adam Fortuna 131501229a fix(runtime): restore request_overrides after transport recovery
Include request_overrides in primary runtime snapshots so transport recovery restores request-level model parameters.
2026-08-29 19:13:12 -07:00
RelaxJonh 91d60d2f9e fix(fallback): re-resolve extra_body when activating fallback provider (#75091)
`try_activate_fallback()` re-resolved `reasoning_config` for the new
fallback provider (fix for #21256), but never re-resolved `extra_body`.
The primary provider's `extra_body` (e.g. `reasoning_effort: "none"`)
rode along onto the fallback provider, which is a different API that
may reject those fields.

Example: primary has `extra_body: {reasoning_effort: "none"}`, fallback
is OpenRouter. After failover, every request to OpenRouter carries both
the stray top-level `reasoning_effort` AND the nested `reasoning` object,
and OpenRouter rejects the pair:
  HTTP 400: "reasoning_effort" and "reasoning.effort" are both provided

The fallback is dead precisely when it is needed.

Fix: after swapping provider/model/base_url, clear the primary's
extra_body from request_overrides, then re-resolve from the fallback
provider's config using the existing _merge_custom_provider_extra_body
helper.  Same pattern as the reasoning_config re-resolution above.
2026-08-29 19:13:12 -07:00
Teknium 5a59ba82dd docs(providers): note extra_body survives gateway turns and /model switches; map gitabtion attribution 2026-08-29 19:13:00 -07:00
Teknium 1859f95799 test(gateway): reused-agent merge-not-overwrite regression via real _run_agent
Port the PR #52432 regression (fast turn then normal turn on a cached
gateway agent) onto the current _run_agent harness: the original test's
host file context no longer exists on main after the TurnRunner
extraction, so the scenario is re-expressed with the existing
_CapturingAgent fixture. Asserts init-time custom-provider extra_body
survives both a /fast turn (service_tier layered on top) and the
following normal turn (only the stale fast-mode key drops).

Salvaged-from: #52432
Co-authored-by: Heng Cai <abtion@outlook.com>
2026-08-29 19:13:00 -07:00
Jack b10b27e6f9 fix(agent): match switched-to custom provider by model+base_url, not name
Addresses the hermes-sweeper review on #53765. The in-place /model switch
helper (_apply_switched_provider_request_overrides) derived a custom
provider's extra_body by provider *name* only, while build-time matching in
agent_init._merge_custom_provider_extra_body matches by provider key, base_url,
AND model. So a different model selected at the same named endpoint could
inherit an extra_body configured for another model.

Reuse the shared agent_init._custom_provider_extra_body_for_agent matcher
(provider key + base_url + model), sourcing custom_providers from the
init-time agent._custom_providers cache (fresh-load fallback if absent). A
stale extra_body is always cleared when no entry matches; non-provider
overrides (service_tier / speed from /fast) are preserved.

Tests: add nonmatching-model and endpoint-mismatch regressions; update the
existing switch tests onto the model/base_url-aware matcher.
2026-08-29 19:13:00 -07:00
Jack 5d238be2ca fix(gateway): carry request_overrides through /model session overrides
Follow-up to the previous commit (which fixed the default/fallback
provider path). A mid-session `/model` switch stores a per-session
override bundle in `_session_model_overrides` that omitted
`request_overrides`, and the two consumers
(`_resolve_session_agent_runtime` fast path and
`_apply_session_model_override`) only copied
provider/api_key/base_url/api_mode. So switching *to* a custom provider
via `/model` did not apply its `extra_body`.

- `ModelSwitchResult` gains a `request_overrides` field, derived for the
  switched provider via `_get_named_custom_provider` /
  `_custom_provider_request_overrides` (the same overrides
  `resolve_runtime_provider` surfaces for the default path).
- Both `/model` override-storage sites in slash_commands.py persist it.
- Both consumers apply it; `_apply_session_model_override` also clears a
  stale value when switching to a provider that has none.

Extends tests/gateway/test_turn_request_overrides.py (3 new cases).
2026-08-29 19:13:00 -07:00
Jack fc00e36c6b fix(gateway): preserve custom-provider request_overrides on agent turns
A `custom_providers` entry can carry an `extra_body` (e.g.
`chat_template_kwargs` to toggle a local vLLM model's thinking).
`resolve_runtime_provider()` correctly surfaces it as `request_overrides`
on the resolved runtime dict, but the gateway never plumbed it through to
the per-turn agent:

- `_resolve_runtime_agent_kwargs()` rebuilt the runtime dict from a fixed
  key whitelist that omitted `request_overrides`.
- `_resolve_turn_agent_config()` rebuilt `runtime` from the same whitelist
  and set `route["request_overrides"]` solely from `/fast` service-tier
  overrides (`{}` otherwise).
- The per-turn `agent.request_overrides = turn_route.get(...)` assignment
  then clobbered the value `_merge_custom_provider_extra_body()` applied at
  agent construction.

Net: on the gateway, a custom provider's configured `extra_body` never
reached the model -- only `/fast` overrides survived. The CLI/TUI path
(which does not go through `_resolve_turn_agent_config`) and the auxiliary
client (which sends `extra_body` directly) were unaffected.

Fix: carry `request_overrides` through the runtime resolvers
(`_resolve_runtime_agent_kwargs`, `_try_resolve_fallback_provider`) and
merge the provider overrides into the per-turn route, layering any `/fast`
service-tier overrides on top (top-level keys, no collision with
`extra_body`).

Adds tests/gateway/test_turn_request_overrides.py.

Known follow-up: the mid-session `/model`-switch override path
(`_session_model_overrides` / `ModelSwitchResult`) does not yet carry
`request_overrides`.
2026-08-29 19:13:00 -07:00
Heng Cai 2f469d7e1e fix(gateway): merge instead of overwrite agent.request_overrides on reused turns
Preserve initialization-time request overrides (custom-provider extra_body
merged at agent construction) while replacing only the previous turn's
routing overrides during the per-turn agent refresh. This keeps
custom-provider extra_body settings without leaving stale fast-mode
service_tier or speed values on cached agents.

Reimplemented from PR #52432 at the code's current location (the per-turn
refresh moved into the TurnRunner path since the original patch), keeping
the original merge semantics: snapshot this turn's route overrides in
agent._gateway_turn_request_overrides, evict only unchanged previous-turn
keys, then layer the new turn overrides on top.

Salvaged-from: #52432
Co-authored-by: Heng Cai <abtion@outlook.com>
2026-08-29 19:13:00 -07:00
Jack d2af990043 fix(agent): carry request_overrides through in-place /model switch (TUI/CLI)
Third in the series. The gateway rebuild path (previous two commits)
carries a custom provider's `request_overrides` (`extra_body`, e.g.
`chat_template_kwargs`) into the agent, but the *in-place* live switch used
by the TUI dashboard and the CLI — `agent.switch_model()` ->
`agent_runtime_helpers.switch_model()` — swapped
model/provider/base_url/api_key without ever updating `request_overrides`.
So a `/model` switch to a thinking-enabled custom provider in the TUI/CLI
kept the previous provider's `extra_body`.

`switch_model()` now re-derives the switched-to provider's
`request_overrides` (via `_get_named_custom_provider`) and applies it in
place, preserving non-provider overrides (`service_tier`/`speed` from
`/fast`). Logic factored into `_apply_switched_provider_request_overrides`
for testability.

Adds tests/agent/test_switch_model_request_overrides.py.
2026-08-29 19:13:00 -07:00
CharZhou a9b696c671 fix(model): initialize switch request overrides 2026-08-29 19:13:00 -07:00
CharZhou 863aac9012 fix: preserve named custom provider request_overrides in gateway and /model switches
Carry provider-derived request_overrides through runtime resolution,
fallback projection, session /model state, restart rehydration, and
turn-route merge so named custom providers keep extra_body and related
overrides.
2026-08-29 19:13:00 -07:00
Teknium 556777ddb1 docs(cron): note request settings carry into scheduled runs 2026-08-29 19:12:51 -07:00
Teknium 1fa3edcb62 chore(attribution): map bsbofmusic noreply email 2026-08-29 19:12:51 -07:00
Patrickk 792dbea777 fix(cron): forward request_overrides into scheduled-job agents
Salvaged from #56876 (cron half only; the delegation half is superseded
by #98237). run_job's ephemeral AIAgent constructor passed api_key /
base_url / provider / api_mode from the resolved runtime but dropped
request_overrides, so cron jobs on custom providers silently lost
extra_body / extra_headers request settings.
2026-08-29 19:12:51 -07:00
Teknium 86a2fdc634 feat(tui): status rule shows cache-hit %, latency, t/s and honors display.status_bar.fields
Extends PR #98250's classic-CLI status-bar upgrades to the Ink TUI:
- tui_gateway/server.py _get_usage() now emits cache_hit_pct,
  avg_latency_s, avg_tps (reads the same per-call deque history from
  agent/conversation_loop.py; keys omitted when no data — Codex
  app-server has no latency, zero cache reads show no %)
- StatusRule renders the three read-outs as width-budgeted tail
  segments (breakpoints 96/104/110 cols, lowest priority — they shed
  first on narrow terminals)
- display.status_bar.fields (the SAME key the classic CLI honors)
  filters TUI segments too: cache_hit, latency, tps, duration,
  compressions, bg_tasks, bg_subagents, voice, battery, title,
  context_pct, context_detail
- values ride the existing usage payload/ticker; constants between
  events so the usage==last dedup keeps suppressing repaints
- 3 new server tests, 5 new TUI tests; full ui-tui suite 1727 green
2026-08-29 19:12:24 -07:00
Teknium 2215fb0e35 fix(providers): mirror new Qwen Cloud models onto alibaba-cn
Follow-up to the #87808 salvage: the domestic alibaba-cn picker list
gets the same five additions (same DashScope catalog, per models.dev).
2026-08-29 19:12:19 -07:00
icocode 04ef14e31f fix(providers): add missing Qwen Cloud (alibaba) models — qwen3.8-max, qwen3.6-flash, glm-5.2, deepseek-v4-pro/flash-0731 2026-08-29 19:12:19 -07:00
Teknium 6cb6aeb168 feat(desktop): real-profile browsing toggle in Capabilities → Tools → Browser
Users reported no GUI switch for browser.use_real_profile — the only
desktop home was the generic Settings → Config editor, which nobody
found. The Browser toolset detail pane now renders a 'Use My Real
Browser Profile' ToggleRow above the backend/provider matrix.

- new BrowserRealProfilePanel: reads the shared profile-scoped config
  record cache, optimistic write-through, rollback on failure
- saveHermesConfigRecord: capability-scoped PUT /api/config counterpart
  of getHermesConfigRecord, so the Capabilities scope selector writes
  the profile it points at (possibly another gateway)
- i18n: en/ja/zh/zh-hant keys (ar inherits en via defineLocale)
- docs: browser.md desktop pointer corrected to the real location

Live E2E on the built app over CDP: clicking the switch flipped
browser.use_real_profile true→false→true in the sandbox HERMES_HOME
config.yaml, GET reflected it, no layout glitches (screenshots in PR).
2026-08-29 19:10:12 -07:00
Brin Shadewater 0ebce1de81 chore: map contributor email
Agent: codex
2026-08-29 19:10:06 -07:00
Brin Shadewater c8bbde7770 feat: allow configured background review tools
Profiles can now grant narrowly scoped tools to the background review runtime whitelist while unrelated tools remain denied. Document the configuration and cover it with a real-config regression test.

Agent: codex
2026-08-29 19:10:06 -07:00
Teknium 5bbb4cd6a9 lint: explicit encoding on harness file opens (ruff unspecified-encoding) 2026-08-29 19:04:29 -07:00
Teknium 7b3c9f86ed Merge origin/main — resolve todo-state seam onto the todo_list rename (accept legacy alias) 2026-08-29 18:55:29 -07:00
hermes-seaeye[bot] 60a4442826 fmt(js): npm run fix on merge (#98265)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-30 01:45:35 +00:00
Teknium 037ce5cf75 eval(tool-search): check in the core-tool-deferral live A/B harness
The harness behind this PR's 288-run maintainer battery, ported from /tmp
into evals/ alongside the readtool / session_search_schema harnesses.
Arm trees parametrized via ABDEFER_BASE_TREE / ABDEFER_PR_TREE (pinned
plain checkouts), results/python roots via ABDEFER_RESULTS / ABDEFER_PYTHON.

- tasks.py: 14 tasks — per-deferred-tool coverage, multistep, long-range,
  clarify ambiguity trap, eager-only control, false-discovery distractor;
  programmatic graders with partial credit
- worker.py: isolated per-cell subprocess, temp HERMES_HOME, hermetic env,
  seeded session DB with decoys, deterministic desktop/computer_use/
  image_generate stubs, interactivity-fairness continuation, exit-3
  infra-abort (misconfig never scores)
- orchestrator.py: resume-safe battery runner, wall timeouts, retry of
  errored records, infra-abort fuse
- report.py: per-task A/B tables + mean-of-task-means
- results/SUMMARY.md: the shipped verdict; rep JSONs gitignored

Ported worker re-verified live post-port (real terra cell, score 1.0).
2026-08-29 18:42:59 -07:00
Teknium c0875ba503 fix(tui): derive resume todo snapshots from already-loaded history
The eager _read_persisted_todo_state(db, target) added a second
get_messages_as_conversation call on every resume, breaking the
one-lineage-SELECT contract pinned by
test_session_resume_uses_parent_lineage_for_display. Derive the
snapshot from the history each resume path already loaded instead;
deferred (defer_history) resumes cache it in the hydration worker once
the transcript arrives.
2026-08-29 18:40:51 -07:00
itsflownium 393af4a310 fix(todo): live task state via revisioned snapshots and a dedicated todo.updated event
Salvaged from PR #97815 by @itsflownium, slimmed to the schema-free core:
- TodoStore gains a monotonic in-memory revision; the todo tool result
  returns it so clients can reject stale updates
- tui_gateway emits a dedicated todo.updated full-snapshot event that
  bypasses optional tool-progress display settings
- session resume/activate responses attach the authoritative todo
  snapshot; renderer restores it with revision arbitration
- desktop store tracks per-session revisions and rejects regressions

The session_todo_state DB table from the original PR is intentionally
dropped: canonical todo tool results already persist in conversation
history, so resume paths derive the snapshot from the stored transcript
instead of a parallel store.
2026-08-29 18:40:51 -07:00
Adolanium cf692532cf fix(desktop): merge todo patches instead of replacing the Tasks list
tool.start for a merge:true todo write used args.todos as a full replace. A one-item status patch became Tasks 1/1, or vanished if content was omitted, so the panel looked stuck at 0/5 until the final complete result. Apply merge by id on start, keep replace for the full result, and show the in-progress spinner on the expanded header too.
2026-08-29 18:40:51 -07:00
Hermes e298dcce49 chore: map contributor email for injaneity 2026-08-29 18:35:17 -07:00
Zane Chee b2e24b986f fix(computer-use): stop launching retired browser-grant runtimes 2026-08-29 18:35:17 -07:00
Teknium 5368598ba1 fix: simplify bad-pin error message (windows-footgun scan tripped on open() inside the string) 2026-08-29 18:35:12 -07:00
Teknium f8546c2eac fix(browser): real-profile follow-ups — reap launched Chrome, headless display-less Linux, register real_profile_pin default + docs
- _terminate_real_profile_chrome(): directly-launched real browsers are ours
  to reap (agent-browser only attaches); wired into the atexit emergency
  cleanup and both launch-failure paths so orphaned Chrome processes can't
  accumulate.
- Display-less Linux gate: append --headless=new (shares the profile's normal
  cookie store, unlike legacy headless) so the direct-launch path doesn't
  regress servers without DISPLAY/WAYLAND_DISPLAY.
- Register browser.real_profile_pin in config_defaults.py and document the
  new launch model + pin in website/docs/user-guide/features/browser.md.
- Drop unused tempfile import from the cherry-picked commit.
2026-08-29 18:35:12 -07:00
Jason Pollak a50b41f843 fix(browser): carry source profile identity into the copy Local State
The Default dir in the snapshot holds the SOURCE profile cookies, but
info_cache['Default'] kept the source user-data-dir own Default entry
(a different person). Chrome saw cookies that belong to profile B while
its profile metadata said profile A, demanded a 'Continue as <name>'
profile-sign-in reconciliation on every launch, and treated the profile
as mid-sign-in. Use the source profile info_cache entry (name + Google
account) for the copy Default.
2026-08-29 18:35:12 -07:00
Jason Pollak 8e746668ba fix(browser): real-profile browsing on macOS - launch real binary, kill sqlite hang, normalize profile copy
Four fixes for real-profile browsing (browser.use_real_profile), found and
verified end-to-end on macOS with a live Chrome:

1. _copy_auth_file: sqlite3.connect('file:...?mode=ro') on a live Chrome
   auth DB can block indefinitely inside lock negotiation - the busy
   timeout never fires, so the 'fail fast' path hangs the launch forever.
   Try immutable=1 first (reads instantly, correct for a committed
   snapshot of a file another process owns); mode=ro stays as fallback.

2. Launch shape: agent-browser's own launch injects --use-mock-keychain /
   --password-store=basic / --headless=new. On macOS the mock keychain
   makes Chrome treat every keychain-encrypted cookie as undecryptable
   and drop it - the copied profile launches signed out (~3 anonymous
   cookies instead of the full jar). Launch the user's real browser
   binary directly on the copy (no mock-keychain switches), wait for
   DevToolsActivePort, then attach agent-browser via --cdp.

3. Snapshot copy: Local State was copied verbatim, still naming the
   SOURCE profile (last_used='Profile 2', info_cache listing several)
   while the copy only contains Default. Chrome opens the missing profile
   dir and starts signed out. Normalize the copy's Local State to
   Default-only.

4. CDP resolution: the agent-browser daemon may report the endpoint of a
   browser IT spawned (throwaway temp profile) instead of the real
   browser we launched on the copy. Trust the port our browser wrote to
   DevToolsActivePort.

Also adds browser.real_profile_pin (optional): pin which source Chromium
profile dir is snapshotted instead of following profile.last_used - on a
machine with a work profile and a personal one, last-used roulette can
silently give the agent the wrong identity. A pin naming a missing dir
fails closed (signed out) rather than falling back to last_used.

Tests: 4 new pin tests + 3 launch tests reshaped to the direct-launch
contract (Popen the real binary, agent-browser attaches). 77 passing.
2026-08-29 18:35:12 -07:00
Teknium 9e017428ba fix: unify status-bar field keys, docs, and tests for salvaged cluster
Follow-up to the cherry-picked #41909/#92696 + #39760 + #97970 cluster:
- single field-key namespace (display.status_bar.fields) instead of the
  second tui_statusbar_fields list; cache_hit/latency/tps/stash/battery/
  title join the existing key set
- cache-hit % prefers the baseline-delta regime (resets on model switch
  and compression) and hides on zero cache reads instead of alarming 0%
- latency/tps segments added to the styled fragment renderer too
- docs updated in website/docs/user-guide/configuration.md
- 7 new tests: rolling latency/t/s, NaN/negative guard, field filtering,
  baseline resets, title badge gating
2026-08-29 18:34:51 -07:00
Turgut Kural 3548fc809b feat(cli): tui status bar per-field toggle + cache/latency/tps
- Add rolling status bar metrics:
  - cache hit ratio (◈) delta since model/compression reset
    (hit = cache_read / prompt, verified against live logs)
  - avg latency (◷) and throughput (↑ t/s) over last 10 API calls
    (deque in agent, displayed in wide bar only)
- Add display.tui_statusbar_fields list to filter segments:
  model, ctx, ctx_bar, cache_hit, latency, tps, compressions,
  bg_tasks, bg_processes, bg_subagents, goal, duration, prompt,
  idle, focus, yolo, stash, battery, title
  Missing/null -> all enabled (backward compat). Unknown keys ignored.
  Title gated via right-align; stash/battery also gated.

- Wide bar (≥76 cols) respects fields, narrow/medium filtered,
  overflow trim preserved. Battery also respects display.battery.

No private data; mock data in tests.

Test: pytest tests/cli/test_cli_status_bar.py etc. 68 passed,
check-windows-footguns clean.
2026-08-29 18:34:51 -07:00
Cheri Wen 4bc7e624d6 feat(cli): show prompt cache hit rate in status bar
Add a ◎ XX% indicator to the CLI status bar showing the prompt cache
hit rate (cache_read / prompt_tokens). This helps users monitor how
effectively their provider's prompt caching is working.

Features:
- Color-coded: green (≥70%), yellow (40-70%), red (<40%)
- Adaptive precision: integer on narrow terminals, one decimal on wide
- Only shown when cache data is available (provider supports it)
- Compatible with OpenAI, Anthropic, DeepSeek, xiaomi, and other
  providers that return prompt_tokens_details.cached_tokens

Tests: 6 new test cases, 43/43 passing
2026-08-29 18:34:51 -07:00
liuhao1024 fb786d2f5b feat(cli): add display.status_bar.fields config for customizing status bar
Allow users to control which fields appear in the interactive CLI status
bar via display.status_bar.fields in config.yaml.

Available fields: model, context_pct, context_detail, compressions,
bg_tasks, bg_processes, duration, prompt_elapsed, yolo, total_tokens.

When the list is empty (default), all fields are shown as before.
The field order is fixed (model always first); the config controls
visibility only. Narrow terminals (<76 cols) automatically drop
context_detail regardless of config.

total_tokens is opt-in only (not shown by default) to avoid width
overflow in the prompt_toolkit fragment renderer.

Closes #41909
2026-08-29 18:34:51 -07:00
Teknium b4403a942a fix(auth): carry the spent-rotation verdict across processes via a durable sidecar registry
The consumed-but-uncommitted rotation verdict was process-local
(_SPENT_ROTATION_FINGERPRINTS), while the credential it protects is
explicitly cross-process: ~/.claude/.credentials.json is shared by every
Hermes profile and process. A fresh interpreter could lease the stale
access token or re-POST the already-spent single-use refresh token and
burn the credential family into invalid_grant.

- Persist non-secret one-way fingerprints to a sidecar registry next to
  the shared singleton source (claude_code / hermes_pkce), written under
  the same path-keyed cross-process lock that serializes refreshes.
- Consult the sidecar in the pool resolver, the pool refresh path, and
  the direct claude_code resolver/refresh before leasing or POSTing.
- Two-process regression: A rotates and loses the commit; B (fresh
  interpreter, empty local registry) must neither lease the stale pair
  nor POST the spent refresh token. Plus a no-verdict control.

Closes the remaining P1 from the exact-head review of f228439b on
PR #87891.
2026-08-29 18:34:35 -07:00
Teknium b095c3d958 fix(tests): carry hermetic guards across the anthropic adapter module split
The adapter godfile split moved credential resolution into
agent/anthropic_credentials.py, which silently disarmed two repository
guards still pointed at the old seam:

- tests/conftest.py::_neutralize_macos_keychain_creds patched only the
  adapter re-export, so the default suite lost its protection against
  reading the operator's real macOS Keychain. Patch the implementation
  owner AND the adapter alias.
- test_oauth_setup_token_keeps_inherited_stdin read only the old source
  file; it now scans both seams and fails loudly if the call moves again.
- test_hermetic_side_effect_guards isolates the owner module directly.
2026-08-29 18:34:35 -07:00
Teknium 1c6b835b3b chore: drop triage-notes markdown files from salvage of #87891 2026-08-29 18:34:35 -07:00
joaomarcos b7a9db8b9a fix(auth): keep the borrowed claude_code row out of token authority and carry the spent-rotation verdict through resolution
Two runtime blockers from the exact-head review of c057ef5.

1. A sanitized `claude_code` pool row was treated as token authority.

`claude_code` is a borrowed source: it is absent from the owned-source
allowlist, so `sanitize_borrowed_credential_payload` strips `access_token`
and `refresh_token` before the row reaches `auth.json`. `load_pool()`
re-hydrates the live pair from the singleton on every load, which is what
makes `~/.claude/.credentials.json` — not the pool store — authoritative
for this source.

`_sync_anthropic_entry_from_pool_store()` re-read that persisted row during
refresh. Being token-less, it "differed" from the live entry, so it was
adopted as a rotation performed by another process: `_refresh_entry()`
replaced a usable credential with an empty one and returned it before
`_claude_code_credentials_lock()` and the authoritative re-read were ever
entered. The empty OAuth entry then stayed selectable, because the
empty-runtime-key guard in `_available_entries()` covered API-key rows only.

Repairs: the pool-store sync refuses borrowed sources outright (plus a
defensive refusal of any token-less row, for future sources that sanitize on
write); the `claude_code` branch of `_refresh_entry()` now runs before the
generic adopt-and-return shortcut, so the path-keyed lock and the
authoritative re-read are always entered before deciding to POST or adopt;
and an OAuth entry with no access token is never leased.

2. A failed commit still fell through to the same spent credential.

`_refresh_oauth_token()` correctly returns None when the refresh POST
rotated the single-use token but the replacement could not be committed.
That verdict did not survive the caller: `resolve_anthropic_token()`
continued to `_resolve_anthropic_pool_token()`, which enumerates read-only
(`clear_expired=False, refresh=False`) over a pool that `load_pool()` had
just re-seeded from the unchanged singleton — so the pair whose refresh half
was already spent came back as a healthy token, and
`_refresh_provider_credentials("anthropic")` reported success and evicted
its cached clients.

Repair: every commit-failure path records the consumed pre-rotation pair as
non-reversible fingerprints (bounded, process-local), and both the Claude
Code file resolver and the pool resolver refuse a credential whose
fingerprint is on that list. `_refresh_provider_credentials("anthropic")`
consequently returns False when the spent family is the only credential,
while genuinely independent pool credentials stay eligible.

Coverage: `test_anthropic_borrowed_row_authority.py` starts from
`load_pool()` reading an actually persisted, actually sanitized row, forces
a refresh, and asserts the full pair survives with exactly one POST and one
commit, that the shared-file lock is entered, and that no empty OAuth entry
can be leased. `test_anthropic_spent_rotation_verdict.py` takes the full
resolver path: successful POST plus failed commit must make
`resolve_anthropic_token()` return None, make
`_refresh_provider_credentials("anthropic")` return False, and keep the
spent fingerprint out of every lease — with a control proving a successful
commit quarantines nothing and an independent credential still resolving.
Five of the seven new borrowed-row tests fail on the previous head, and the
three resolution tests fail with the verdict disabled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gcoy6nLTg5R6FHHhjcLZEC
2026-08-29 18:34:35 -07:00
joaomarcos 7cbffdd125 refactor(anthropic): split the adapter godfile into four modules
`agent/anthropic_adapter.py` was 3,423 lines and this PR adds another auth
boundary to it. Split along the seams that were already there, so the
credential surface this PR changes has a single owner instead of being
interleaved with request building:

- `agent/anthropic_endpoints.py` (258) — base-URL/endpoint-family predicates.
  Pure functions over a URL string, which is what lets both of the modules
  below depend on it without a cycle.
- `agent/anthropic_message_convert.py` (1,225) — OpenAI-style to Anthropic
  Messages payload conversion: model ids, tool schemas, content/thinking
  blocks, tool_use pairing, cache_control, screenshot eviction, blank-block
  scrubbing.
- `agent/anthropic_credentials.py` (910) — credential sources, the OAuth
  flows, and the refresh commit (`CredentialPersistError` and both singleton
  writers).
- `agent/anthropic_adapter.py` (1,215) — client construction and the Messages
  API call, re-exporting every name from the three modules above so existing
  `from agent.anthropic_adapter import ...` imports keep resolving. The
  re-export surface was diffed against the pre-split module: nothing dropped.

Call sites that read a moved name through the adapter's namespace at runtime
(`credential_pool._refresh_entry_impl`, `auxiliary_client`) now import it from
the defining module, so there is one patchable seam rather than two bindings
that can disagree. The tests that monkeypatched those seams were retargeted to
match; no assertion was changed.

No behavior change.
2026-08-29 18:34:35 -07:00
joaomarcos 07faed33bb fix(auth): make the Anthropic refresh commit part of the transaction
Anthropic OAuth refresh tokens are single-use: the POST that returns a new
pair invalidates the one that was sent. The replacement therefore only
becomes real once it reaches its authoritative store -
~/.claude/.credentials.json for claude_code entries,
~/.hermes/.anthropic_oauth.json for hermes_pkce ones. Both writers caught
OSError/IOError, logged at debug level and returned nothing, so no caller
could tell a durable commit from a failed one.

That let a refresh spend the only refresh token, report success, and leave
the consumed pre-rotation pair on disk. _seed_from_singletons() re-reads
those files on every load_pool(), so the next process seeded the spent pair
back over the fresh pool row and the following refresh replayed a consumed
token (invalid_grant / refresh_token_reused) - exactly the failure this PR
set out to remove.

- _write_claude_code_credentials() and _write_hermes_oauth_credentials()
  now raise CredentialPersistError instead of swallowing the write error.
- _refresh_oauth_token() treats a failed commit as a failed refresh and
  returns None rather than handing back an access token whose refresh half
  was lost.
- _refresh_entry_impl() fails closed on both the primary and the recovery
  path: the rotated pair is never marked, persisted or returned, and the
  entry is quarantined DEAD with a credential_persist_failed reason so it
  leaves rotation and surfaces as an explicit re-auth instead of a silent
  fallback to another provider. The retry path now commits to the singleton
  before persisting the pool row.
- _upsert_entry() no longer treats re-seeding a borrowed source as a
  rotation. Borrowed rows (claude_code, env-backed) are written to auth.json
  without their secret, so comparing the re-seeded token against the empty
  stored value reported a rotation on every load and cleared the DEAD state
  the previous process had just written - resurrecting the quarantined,
  already-consumed credential on restart. It now compares the incoming
  token against the row's secret_fingerprint.

Adds failure-injection coverage for both writers, the direct resolver, the
claude_code and hermes_pkce pool paths and the retry path, each asserting
that a reload cannot bring the pre-refresh pair back as a usable credential.
2026-08-29 18:34:35 -07:00
joaomarcos 0099f250c2 fix(auth): close Anthropic OAuth review gaps 2026-08-29 18:34:35 -07:00
joaomarcos e1a210652a fix(auth): harden claude_code refresh lock and remove dashboard Anthropic OAuth
Add a cross-process lock over the shared ~/.claude/.credentials.json file
so concurrent Hermes processes racing a claude_code-sourced Anthropic
refresh resync instead of losing the update (mirrors the existing
per-profile auth-store lock, kept as the outer lock per the documented
lock-ordering invariant).

Remove the dashboard-triggered Anthropic PKCE OAuth flow entirely rather
than continue patching it: an unattended HTTP endpoint minting Claude
Pro/Max subscription tokens outside Anthropic's own client sits on the
wrong side of Anthropic's OAuth usage policy. The provider catalog entry
is now flow == "external", pointing at `hermes auth add anthropic`
(terminal PKCE, unaffected, out of scope). Drop the now-dead PKCE
functions/constants and the tests that exercised only that removed code.
2026-08-29 18:34:35 -07:00
joaomarcos 5490029a71 docs(auth): record manual A/B validation of OAuth API-key shadowing fix
Confirms the fix from 41b7aba875 with a real before/after test: without it,
resolve_anthropic_token() keeps returning a stale API key after an OAuth
dashboard login; with it, the key is auto-cleared and OAuth wins. Also notes
the installed app still needs to be updated past main@8c8d55b to pick this up.
2026-08-29 18:34:35 -07:00
joaomarcos 739dc6d198 fix(auth): close Anthropic OAuth CSRF gap, cross-process refresh race, and API-key shadowing
Dashboard PKCE login reused the code_verifier as the OAuth state (leaking
it and disabling CSRF validation) and never checked state on callback --
the same class of bug already fixed for the CLI flow. Credential-pool
refresh excluded "anthropic" from the cross-process lock Codex/xAI already
get, so concurrent Hermes processes racing a single-use refresh token could
leave the loser stuck exhausted with no recovery for hermes_pkce/dashboard
sources. The dashboard OAuth save also never cleared a stale
ANTHROPIC_API_KEY, which resolve_anthropic_token() prioritizes over the
OAuth pool entry by design -- so a leftover key silently kept billing
pay-per-token after a Claude Pro/Max login.

A concurrency stress test written to validate the refresh-race fix under
load surfaced a fifth, unrelated bug: _auth_store_lock()'s Windows
lock-file "ensure content" write was unguarded and could raise an uncaught
PermissionError under real contention -- affecting every single-use-token
provider sharing that lock, not just Anthropic.

Fixes #87887, #87888, #87889.
2026-08-29 18:34:35 -07:00