Commit Graph

3725 Commits

Author SHA1 Message Date
joaomarcos b7a9db8b9a fix(auth): keep the borrowed claude_code row out of token authority and carry the spent-rotation verdict through resolution
Two runtime blockers from the exact-head review of c057ef5.

1. A sanitized `claude_code` pool row was treated as token authority.

`claude_code` is a borrowed source: it is absent from the owned-source
allowlist, so `sanitize_borrowed_credential_payload` strips `access_token`
and `refresh_token` before the row reaches `auth.json`. `load_pool()`
re-hydrates the live pair from the singleton on every load, which is what
makes `~/.claude/.credentials.json` — not the pool store — authoritative
for this source.

`_sync_anthropic_entry_from_pool_store()` re-read that persisted row during
refresh. Being token-less, it "differed" from the live entry, so it was
adopted as a rotation performed by another process: `_refresh_entry()`
replaced a usable credential with an empty one and returned it before
`_claude_code_credentials_lock()` and the authoritative re-read were ever
entered. The empty OAuth entry then stayed selectable, because the
empty-runtime-key guard in `_available_entries()` covered API-key rows only.

Repairs: the pool-store sync refuses borrowed sources outright (plus a
defensive refusal of any token-less row, for future sources that sanitize on
write); the `claude_code` branch of `_refresh_entry()` now runs before the
generic adopt-and-return shortcut, so the path-keyed lock and the
authoritative re-read are always entered before deciding to POST or adopt;
and an OAuth entry with no access token is never leased.

2. A failed commit still fell through to the same spent credential.

`_refresh_oauth_token()` correctly returns None when the refresh POST
rotated the single-use token but the replacement could not be committed.
That verdict did not survive the caller: `resolve_anthropic_token()`
continued to `_resolve_anthropic_pool_token()`, which enumerates read-only
(`clear_expired=False, refresh=False`) over a pool that `load_pool()` had
just re-seeded from the unchanged singleton — so the pair whose refresh half
was already spent came back as a healthy token, and
`_refresh_provider_credentials("anthropic")` reported success and evicted
its cached clients.

Repair: every commit-failure path records the consumed pre-rotation pair as
non-reversible fingerprints (bounded, process-local), and both the Claude
Code file resolver and the pool resolver refuse a credential whose
fingerprint is on that list. `_refresh_provider_credentials("anthropic")`
consequently returns False when the spent family is the only credential,
while genuinely independent pool credentials stay eligible.

Coverage: `test_anthropic_borrowed_row_authority.py` starts from
`load_pool()` reading an actually persisted, actually sanitized row, forces
a refresh, and asserts the full pair survives with exactly one POST and one
commit, that the shared-file lock is entered, and that no empty OAuth entry
can be leased. `test_anthropic_spent_rotation_verdict.py` takes the full
resolver path: successful POST plus failed commit must make
`resolve_anthropic_token()` return None, make
`_refresh_provider_credentials("anthropic")` return False, and keep the
spent fingerprint out of every lease — with a control proving a successful
commit quarantines nothing and an independent credential still resolving.
Five of the seven new borrowed-row tests fail on the previous head, and the
three resolution tests fail with the verdict disabled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gcoy6nLTg5R6FHHhjcLZEC
2026-08-29 18:34:35 -07:00
joaomarcos 7cbffdd125 refactor(anthropic): split the adapter godfile into four modules
`agent/anthropic_adapter.py` was 3,423 lines and this PR adds another auth
boundary to it. Split along the seams that were already there, so the
credential surface this PR changes has a single owner instead of being
interleaved with request building:

- `agent/anthropic_endpoints.py` (258) — base-URL/endpoint-family predicates.
  Pure functions over a URL string, which is what lets both of the modules
  below depend on it without a cycle.
- `agent/anthropic_message_convert.py` (1,225) — OpenAI-style to Anthropic
  Messages payload conversion: model ids, tool schemas, content/thinking
  blocks, tool_use pairing, cache_control, screenshot eviction, blank-block
  scrubbing.
- `agent/anthropic_credentials.py` (910) — credential sources, the OAuth
  flows, and the refresh commit (`CredentialPersistError` and both singleton
  writers).
- `agent/anthropic_adapter.py` (1,215) — client construction and the Messages
  API call, re-exporting every name from the three modules above so existing
  `from agent.anthropic_adapter import ...` imports keep resolving. The
  re-export surface was diffed against the pre-split module: nothing dropped.

Call sites that read a moved name through the adapter's namespace at runtime
(`credential_pool._refresh_entry_impl`, `auxiliary_client`) now import it from
the defining module, so there is one patchable seam rather than two bindings
that can disagree. The tests that monkeypatched those seams were retargeted to
match; no assertion was changed.

No behavior change.
2026-08-29 18:34:35 -07:00
joaomarcos 07faed33bb fix(auth): make the Anthropic refresh commit part of the transaction
Anthropic OAuth refresh tokens are single-use: the POST that returns a new
pair invalidates the one that was sent. The replacement therefore only
becomes real once it reaches its authoritative store -
~/.claude/.credentials.json for claude_code entries,
~/.hermes/.anthropic_oauth.json for hermes_pkce ones. Both writers caught
OSError/IOError, logged at debug level and returned nothing, so no caller
could tell a durable commit from a failed one.

That let a refresh spend the only refresh token, report success, and leave
the consumed pre-rotation pair on disk. _seed_from_singletons() re-reads
those files on every load_pool(), so the next process seeded the spent pair
back over the fresh pool row and the following refresh replayed a consumed
token (invalid_grant / refresh_token_reused) - exactly the failure this PR
set out to remove.

- _write_claude_code_credentials() and _write_hermes_oauth_credentials()
  now raise CredentialPersistError instead of swallowing the write error.
- _refresh_oauth_token() treats a failed commit as a failed refresh and
  returns None rather than handing back an access token whose refresh half
  was lost.
- _refresh_entry_impl() fails closed on both the primary and the recovery
  path: the rotated pair is never marked, persisted or returned, and the
  entry is quarantined DEAD with a credential_persist_failed reason so it
  leaves rotation and surfaces as an explicit re-auth instead of a silent
  fallback to another provider. The retry path now commits to the singleton
  before persisting the pool row.
- _upsert_entry() no longer treats re-seeding a borrowed source as a
  rotation. Borrowed rows (claude_code, env-backed) are written to auth.json
  without their secret, so comparing the re-seeded token against the empty
  stored value reported a rotation on every load and cleared the DEAD state
  the previous process had just written - resurrecting the quarantined,
  already-consumed credential on restart. It now compares the incoming
  token against the row's secret_fingerprint.

Adds failure-injection coverage for both writers, the direct resolver, the
claude_code and hermes_pkce pool paths and the retry path, each asserting
that a reload cannot bring the pre-refresh pair back as a usable credential.
2026-08-29 18:34:35 -07:00
joaomarcos 0099f250c2 fix(auth): close Anthropic OAuth review gaps 2026-08-29 18:34:35 -07:00
joaomarcos e1a210652a fix(auth): harden claude_code refresh lock and remove dashboard Anthropic OAuth
Add a cross-process lock over the shared ~/.claude/.credentials.json file
so concurrent Hermes processes racing a claude_code-sourced Anthropic
refresh resync instead of losing the update (mirrors the existing
per-profile auth-store lock, kept as the outer lock per the documented
lock-ordering invariant).

Remove the dashboard-triggered Anthropic PKCE OAuth flow entirely rather
than continue patching it: an unattended HTTP endpoint minting Claude
Pro/Max subscription tokens outside Anthropic's own client sits on the
wrong side of Anthropic's OAuth usage policy. The provider catalog entry
is now flow == "external", pointing at `hermes auth add anthropic`
(terminal PKCE, unaffected, out of scope). Drop the now-dead PKCE
functions/constants and the tests that exercised only that removed code.
2026-08-29 18:34:35 -07:00
joaomarcos 739dc6d198 fix(auth): close Anthropic OAuth CSRF gap, cross-process refresh race, and API-key shadowing
Dashboard PKCE login reused the code_verifier as the OAuth state (leaking
it and disabling CSRF validation) and never checked state on callback --
the same class of bug already fixed for the CLI flow. Credential-pool
refresh excluded "anthropic" from the cross-process lock Codex/xAI already
get, so concurrent Hermes processes racing a single-use refresh token could
leave the loser stuck exhausted with no recovery for hermes_pkce/dashboard
sources. The dashboard OAuth save also never cleared a stale
ANTHROPIC_API_KEY, which resolve_anthropic_token() prioritizes over the
OAuth pool entry by design -- so a leftover key silently kept billing
pay-per-token after a Claude Pro/Max login.

A concurrency stress test written to validate the refresh-race fix under
load surfaced a fifth, unrelated bug: _auth_store_lock()'s Windows
lock-file "ensure content" write was unguarded and could raise an uncaught
PermissionError under real contention -- affecting every single-use-token
provider sharing that lock, not just Anthropic.

Fixes #87887, #87888, #87889.
2026-08-29 18:34:35 -07:00
Teknium 1ee30352ca fix: background review can now read skills before patching — denial storm ended, cache parity intact (#61521, #39996)
The self-improvement review fork advertises the parent's full tool schema
(deliberate — tools[] must stay byte-identical for prompt-cache parity)
but denied everything except memory/skill tools at dispatch. Models
naturally reach for read_file to inspect a SKILL.md before patching, got
denied, then attempted a blind skill_manage patch which the
read-before-write guard correctly refused. One deployment logged ~142
denials + ~204 refusals over 2 days: the self-improvement loop ran
continuously but almost never landed a skill patch.

Fix is dispatch-side ONLY — zero request-body change, cache untouched:

- Whitelist read_file + search_files on the review fork (reads are
  side-effect-free). Write tools (write_file/patch/terminal) stay denied:
  autonomous maintenance must go through skill_manage's validation.
- read_file now registers full reads with the review fork's
  read-before-write guard (same as skill_view), so the natural
  read_file -> skill_manage(patch) sequence lands. Partial reads
  (offset>1 / truncated) don't count. No-op outside review forks.
- Self-correcting deny message: names skill_view/skill_manage/memory as
  substitutes so one denial redirects the model instead of a storm
  (the actionable half of #61521's proposal 2).

Rejects #39997's alternative (narrow the advertised schema on local
endpoints): local backends have KV/prefix caches too, and re-prefilling
a large snapshot is most expensive exactly there.

Live A/B (real dispatch path, isolated HERMES_HOME): on main,
read_file DENIED -> patch REFUSED (read-before-write); on this branch,
read_file OK -> patch LANDED. tools[] identical in both.
2026-08-29 18:33:41 -07:00
Teknium c1762ff11c test: sweep sibling tests stale on the tool renames + deferral default
The rename sweep in the base commit missed the sibling-test blast radius
(18 red files on CI). Three classes, all fixed:

1. Stale old names in tests (todo/cronjob/process/tour/tip) — updated to
   todo_list/cronjob_manage/process_manage/gui_tour/show_tip at every
   registry.get_entry/dispatch/coerce/preview/allowlist call site, plus
   the coding-brief sentence in agent/coding_context.py now names
   todo_list (and its gating test).
2. Missed rename in production: AGENT_RUNTIME_POST_HOOK_TOOL_NAMES still
   held 'tour' — post-hook ownership would have double-emitted for
   gui_tour via the bridge path.
3. Tests pinning pre-deferral assembly (blank-slate surface, modal
   sandbox resolution, desktop diet, HUD note) now pin their ACTUAL
   contract under the legacy defer:[] override, or assert on granted
   tool names instead of visible schemas.

Also fixes a pre-existing ordering flake surfaced by the sweep:
test_holds_exactly_the_gui_affordances depended on whether an earlier
test had imported apply_layout_tool (registry-registered, not in the
static desktop_ui list) — now forces discovery and pins the full set.

649 tests green locally across all touched files, both orderings.
2026-08-29 18:23:07 -07:00
Hyusein Leshov 835a913ffd fix(compression): arm the failure cooldown when codex compaction fails
Closes #75364.

`_compress_context_via_codex_app_server` returns the transcript unchanged
when the codex thread reports `interrupted` or `error`. The session is
therefore still above threshold, and nothing records that the attempt
failed — so the next turn retries immediately, and keeps retrying for as
long as the condition persists.

Every other compression path arms the shared failure cooldown, records an
ineffective-compression strike, or both. This path records neither:

* `_hygiene_compression_failure_cooldowns` is set only on
  `asyncio.TimeoutError`, or behind `_last_compress_aborted`, which is
  assigned exclusively in `context_compressor.py` on the Hermes summarizer
  path.
* `compression_ineffective_count` lives in `ContextCompressor`, and this
  path returns before any compressor bookkeeping runs.

`compress_context` already documents the rule this path was missing —
"Every automatic entrypoint must honor compressor-owned cooldown and
breaker state" — but the codex branch dispatches above that block and
returns from inside it.

`result.interrupted` needs no unusual configuration to occur: an ordinary
user message arriving mid-compaction sets it (see
`codex_app_server_session.py`, which produces the "compact turn
interrupted" string). Observed in production on a Discord gateway session
at ~315k tokens against a 258k window, where compaction was attempted on
essentially every turn for ~70 minutes; the session's
`compression_ineffective_count` was still 0 afterwards.

This reuses the existing cooldown rather than adding a new mechanism:

* arm `_record_compression_failure_cooldown` with the existing
  `_SUMMARY_FAILURE_COOLDOWN_SECONDS` when compaction returns
  interrupted/error;
* honor an active cooldown on entry, matching the Hermes path.

`force=True` bypasses both, so an explicit /compress is never braked by a
failure it did not cause, and a successful compaction arms nothing.
2026-08-29 22:29:28 +05:30
Teknium e16ad33a9d feat(tool-search): core-tool deferral — curated 19-tool set behind the bridge by default; renames todo_list/cronjob_manage/process_manage/gui_tour/show_tip with legacy aliases (13.4K -> 6.9K desktop schemas, -49%) 2026-08-29 08:26:24 -07:00
Teknium 578f85cfb0 feat: /btw rides the background-review cache-parity fork for full-context answers
The initial /btw implementation (#97937) answered from a rendered
plain-text transcript digest — truncated context, cold-written tokens on
every question. Teknium's call: reuse the self-improvement review fork
instead, which keeps the entire prompt cache stable for the fork and
gives it the complete conversation for very cheap.

- agent/background_review.py: extract the review-fork construction into
  build_cache_parity_fork() — same runtime/credentials as the parent,
  byte-identical system prompt / tools[] / reasoning config on the
  same-model path, shared session_id for prefix warmth, full persistence
  detachment (no state.db writes, no rotation, no external memory,
  in-place-only compaction). The review thread now calls the helper;
  behavior unchanged (full review test suite green).
- agent/side_question.py: /btw prefers the fork when a live parent
  AIAgent exists — replays the untruncated snapshot as warm cache reads,
  denies every tool at dispatch via an empty thread whitelist (tools[]
  stays byte-identical for cache parity), attributes usage to the parent,
  and trims a mid-turn snapshot tail so role alternation holds. The
  one-shot digest remains as fallback (no live agent = cold cache anyway,
  and any fork failure degrades gracefully).
- CLI passes self.agent, TUI passes the session agent, gateway looks up
  the chat's cached agent (parity with how turns reuse it).

Live-verified: /btw on the worktree runs the fork path (agent.log shows
the side question as a forked conversation turn on the parent session_id
with the full history replayed), answers correctly from context.
2026-08-29 08:23:47 -07:00
kshitijk4poor ac5c8f58db fix: drop duplicate hy4-preview context entry — main's 1_048_576 wins
Follow-up for salvaged PR #96939: main already added hy4-preview at
1_048_576 (f7c79efbac); the cherry-picked duplicate key later in the
dict silently overrode it with 1024000.
2026-08-29 20:51:17 +05:30
simonweng 0fb5cab0d4 feat:add hy4-preview model and tokenplan provider 2026-08-29 20:51:17 +05:30
kshitijk4poor b954547e72 fix(nebius): route effort through canonical clamp_effort — hand-rolled map inverted the ladder
Review finding on salvaged #28253: the hand-rolled mapping sent
ultra -> medium while xhigh -> high (stronger request, weaker wire
value). Declare NEBIUS_EFFORTS in agent/reasoning_effort.py and use
clamp_effort like the zai/kimi/tokenhub call sites; disable detection
stays ahead of the clamp since clamp_effort('none', ...) returns the
floor, not off. Adds a monotonicity regression test.
2026-08-29 20:39:44 +05:30
kshitijk4poor 1c5ee5815f fix(router): pytest guard on caps warmer + debug log in fail-open efforts lookup
Review findings on salvaged #93548: _warm_efforts_async now returns
early under PYTEST_CURRENT_TEST (matching the canonical OpenRouter caps
warmer) so a test that forgets to monkeypatch it can't fire live HTTP
when RAMP_ROUTER_API_KEY is set; the codex transport's fail-open
except in _profile_declared_efforts logs at debug instead of silently
swallowing profile-hook bugs.
2026-08-29 20:29:04 +05:30
Neel Patel eeb7916000 review: host-resolved efforts, ladder-validated ingest, deduped catalog
Addresses the automated review on this PR:

- _profile_declared_efforts falls back from provider name to the
  endpoint's host (via model_metadata's URL->provider map), so a named
  custom provider pointed at api.router.com — which the host mandate
  already routes onto this transport — gets the catalog clamp instead
  of the default vocabulary and a Router 400.
- _parse_efforts validates catalog levels against EFFORT_LADDER at
  ingest, logging and dropping unrecognized tiers; a model whose whole
  vocabulary is unrecognized stays out of the map (transport defaults)
  instead of passing the requested effort through unclamped.
- fetch_models dedupes ids while preserving Router's deliberate listing
  order.
- plugin.yaml credits the human contributor per repo convention.
2026-08-29 20:29:04 +05:30
Neel Patel 804f8b4732 feat(providers): add Ramp Router (router.com) provider plugin
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:

- plugins/model-providers/router/: RouterProfile plugin —
  api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
  RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
  GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
  Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
  runtime_provider._detect_api_mode_for_url: api.router.com ->
  codex_responses. The host is Responses-only — POST /v1/chat/completions
  does not exist and 404s — so this is a genuine host mandate (exact
  hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
  hook (tri-state: None=defer, ()=model takes no reasoning params,
  tuple=clamp set). Router validates reasoning.effort per model and
  returns HTTP 400 invalid-argument on levels outside the model's
  published vocabulary, and 400 unsupported_parameter when a
  non-reasoning model receives any reasoning field (both verified live).
  The profile answers from a cached copy of the catalog's
  router.capabilities.reasoning block: cache-only on the hot path,
  seeded for free by fetch_models(), disk-mirrored across processes
  (/cache/router_catalog.json), background-warmed when cold
  — same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
  the generic effort-clamp branch (xai/actual/github branches untouched;
  profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
  document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
  rejection, profile registration + auth auto-registry wiring, catalog
  parsing, and transport clamp/suppression/fallback paths.

Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.
2026-08-29 20:29:04 +05:30
Teknium 74a95a3ddf feat: /btw now answers side questions with conversation context; /background renamed to /bg
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.

/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.

Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
2026-08-29 07:25:17 -07:00
Teknium 11b98a1429 feat(system_prompt): two-line conversation clock — anchored start + rebuild-day line (salvages #96224) (#97930)
* test(system_prompt): cover session-start anchoring + fix Windows-portable expectation

- Add TestSessionStartLike unit tests for _session_start_like(): session-id
  embedded timestamp, session_start fallback, now fallback, non-matching id.
- Add a build-level regression: a session started Jan 1 must still render
  'Conversation started: Thursday, January 01' when the prompt is rebuilt
  on Jan 2 (the rebuild-drift bug).
- test_coding_prompt_preserves_legacy_workspace_order hardcoded '/hermes'
  while production renders str(Path('/hermes')) — backslash on Windows made
  the suite fail on Windows (CI runs Linux, so it was never caught). Build
  the expectation via str(Path()) to match production on every platform.

* fix(system_prompt): anchor 'Conversation started' to the real session start

The timestamp line stamped hermes_time.now() at system-prompt build time.
The prompt is rebuilt on compression, fresh-agent gateway turns, and
resume-without-stored-prompt, so the date silently advanced to whatever
day the prompt was last rebuilt — a chat that started on Wednesday read
as 'Conversation started: Thursday' after a Thursday-morning resume,
contradicting the fresh per-turn time hint.

Resolve the true start via _session_start_like(): the timestamp embedded
in the session id (YYYYMMDD_HHMMSS_..., immutable for the session life)
-> agent.session_start -> now() only as last resort. Box-local stamps are
attached to the box's local zone then converted to the rendered zone so
the date is consistent with the per-turn clock. The line stays date-only
and is now byte-stable for the whole session (never moves on rebuild),
preserving prefix-cache KV. The zone suffix and _bot_chat_timeless_prompt
behaviour are untouched.

* feat(system_prompt): two-line conversation clock — anchored start (salvaged #96224, credit @bobaba76) + as-of-last-rebuild date for multi-day sessions

---------

Co-authored-by: bobaba76 <79245850+bobaba76@users.noreply.github.com>
2026-08-29 07:06:54 -07:00
Teknium f89f0a2eaa feat(prompt): default identity rewritten as a behavior spec — sizing rule, named prohibitions, anti-sycophancy, earned depth; exploration-thrift line deliberately removed (models under-explore) (#97926) 2026-08-29 07:00:53 -07:00
Teknium 5241df3d4a fix(prompt): skills-section cleanup — drop '(mandatory)' header, delete hermes-agent paragraph duplicating the help guidance, gate the skill pointer on the skill actually being installed, cut the 'when the two differ' dead clause (#97918) 2026-08-29 06:53:02 -07:00
Teknium e387cbc0aa refactor(prompt): platform-hint diet — 7 heavies compressed, −657 tok across the map (facts probe-pinned) (#97899)
* refactor(prompt): platform-hint diet — shared _MEDIA_NATIVE spine; seven heavies compressed with every verified fact intact (3,175 -> ~2,520 map total, -657)

* refactor(prompt): steer-channel note diet 225 -> 155 — marker is self-describing since its own provenance+replay clauses; prompt keeps only anti-lookalike + authority + latest-results scope (#40240/#76805 archaeology in comment)
2026-08-29 06:37:06 -07:00
Teknium ccc367dce0 fix(prompt)+feat(gateway): platform-hint truth pass + universal voice-bubble transcode (all 22 hints source-verified) (#97873)
* fix(prompt): platform-hint truth pass — CLI/TUI file-delivery reality (paths/URLs only, MEDIA: prints literally), CLI no-markdown verified live, Slack/Discord markdown+tables truth, shared local-cron constant

* feat(gateway): universal voice-bubble delivery — shared transcode_to_ogg_opus; telegram [[audio_as_voice]] any-format; feishu native voice; hints to new truth

* chore: delete the webui ghost hint (tombstone comment, audit-verified); sync send_voice signature pin in tts routing test
2026-08-29 05:57:13 -07:00
Teknium 9d9f44d638 refactor(prompt): desktop hint diet — recipe-first widget teaching verified against the renderer source; setup_mcp sentence delegated to its schema (442 -> 307 tok) (#97850) 2026-08-29 04:35:08 -07:00
Teknium fae063fc74 fix(desktop): MEDIA: non-media files get the preview file card, not a degraded 'Open' anchor (#97812)
* fix(desktop): MEDIA:-delivered non-media files route to the preview pipeline — PDFs/data files get the file card instead of a dead 'Open' anchor (extends #84951 to every extension)

* docs(prompt): desktop guidance aligned with any-file MEDIA: delivery — preview card truth, local-markdown-image block warning
2026-08-29 02:53:42 -07:00
Teknium a2e19d484c refactor(prompt): memory/skills guidance — one builder, positive posture, session-accurate wording (537 → 255 served, −282/call) (#97760)
* refactor(prompt): diet the memory/skills guidance block — schema-taught curricula removed, form rule + pruning contract kept (537 -> 223 tok in the combined block)

* polish: literal check-glyphs in source; memory capacity posture — save proactively, replace/consolidate when full

* refactor: single spine for memory/profile guidance — form rule + capacity posture written once, variants differ only in opening frame

* refactor: ONE memory-guidance builder — frame adapts to enabled stores, body written once, positive posture leads (maintainer direction)

* fix wording: memory is loaded per SESSION, not injected per turn (maintainer correction)
2026-08-29 02:30:44 -07:00
Machan-Army 31b974dfb7 fix(gateway): split turn-hold expiry from idle-timeout failure path
Introduce HygieneTurnHoldExceeded exception so turn-hold budget expiry
no longer collapses into the generic asyncio.TimeoutError handler.

- Add HygieneTurnHoldExceeded exception (availability boundary, not a failure)
- Add dedicated handler: stamps AGENT_COMPRESSION_TURNHOLD provenance,
  sends deferral notice, does NOT increment failure cooldown
- Preserve #87011 contract: idle timeout still sends 'no output' message
  and takes failure path
- Add behavior witnesses: turn-hold ≠ idle timeout, cooldown untouched

Fixes semantic boundary collapse flagged in PR #90845 review.
2026-08-29 13:02:13 +05:30
Jakub Wolniewicz 23bae43cfa fix(agent): normalize list-shaped streaming content deltas 2026-08-29 12:45:43 +05:30
kshitijk4poor c03c72a19e refactor(vertex): fold _creds_cache_key into _sa_snapshot (simplify pass)
Final /simplify-code pass on the full diff: after the content-digest
rework, _creds_cache_key had become a wrapper with zero production
callers — get_vertex_credentials inlined the same read/except dance —
kept alive only by its own tests. One helper now owns the
(bytes, cache_key) resolution for all three cases (ADC sentinel,
readable file, unreadable fallback); prod calls it, tests target it.
No behavior change: 10/10 green, same mutation-check results.
2026-08-29 12:33:44 +05:30
kshitijk4poor 91608eb20e docs: note the MEASURED_1H_PROVIDERS exception in effective_cache_ttl's docstring
Review finding: the docstring still claimed all Qwen/Alibaba routes clamp
to 5m, contradicting the new allow-list one paragraph of code below.
2026-08-29 12:23:41 +05:30
Jack Powrie d51d66e869 fix(cache): preserve the configured 1h TTL on the OpenCode Go route
`effective_cache_ttl` evaluates the generic `is_qwen_model` clamp before any
route-level allowance, so a configured `prompt_caching.cache_ttl: 1h` is
silently rewritten to `5m` for every Qwen model on opencode-go. Reproduced on
this commit's parent, no provider traffic:

    effective_cache_ttl('1h', 'opencode-go', 'qwen3.7-plus')  -> '5m'
    _build_marker('5m')                                       -> {'type': 'ephemeral'}

A repair for this exists historically -- payload d6b33faae1, merged as
a43fe4918d -- but neither commit is reachable from current main
(`git merge-base --is-ancestor` returns non-zero for both), so the regression
is live on this lineage.

Restores the precedence fix: MEASURED_1H_PROVIDERS (an allow-list holding only
opencode-go, the one route measured with a delayed read past five minutes) is
consulted ahead of the generic Qwen clamp, with NO_1H_TIER_MODELS nested inside
it for models measured to ignore the tier even on a capable route.

opencode-go stays in ALIBABA_FAMILY_PROVIDERS. That set is also the
cache-marker-layout opt-in read by
agent_runtime_helpers.anthropic_prompt_cache_policy, so narrowing it would turn
working five-minute caching into *no* caching rather than extending the window.
The two sets are kept separate on purpose and a test pins the separation.

One deliberate divergence from the historical payload, found by independent
review: that patch checked NO_1H_TIER_MODELS globally, ahead of the provider
gate. MiniMax on its own Anthropic-compatible endpoint is a separate and
genuinely cache-eligible route, and the global check regressed its configured
1h to 5m off the back of an opencode-go observation -- an unrelated-provider
change this repair must not make. Measured:

    base      effective_cache_ttl('1h', 'minimax', 'MiniMax-M2.5') -> '1h'
    payload                                                        -> '5m'
    here                                                           -> '1h'

Scope note on the evidence: the delayed-read run covered qwen3.8-max and
glm-5.2. The rule is keyed on the route, not the model, so the deployed
qwen3.7-plus is covered by it but has never itself been measured. This change
restores *sending* the requested 1h marker; it does not establish that the
provider honours 1h retention. Those stay separate claims, and the provider
labels every write `ephemeral_5m_input_tokens` whatever ttl was requested, so
only a delayed read past five minutes with no intervening call can settle it.

Tests: 11 new cases covering the deployed model, marker shape, precedence
against the generic clamp, cache eligibility surviving the repair, negative
controls for unrelated routes, provider case normalization, and a closed-set
audit of every marker-emitting call site. Ten mutations were applied to scratch
copies and each turned the suite red, including hoisting the generic clamp back
above the allowance, dropping opencode-go from ALIBABA_FAMILY_PROVIDERS,
un-nesting the model denial, and adding a new unclamped sender.

Known risk, not discharged here: this marker shape has never been sent on the
Go relay. #77217 records the sibling Zen relay returning HTTP 400 on an
unexpected marker shape, so field validation must check HTTP status before it
looks at any cache counter.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QV9Ld4f5ZTqdZWncpfZcQh
2026-08-29 12:23:41 +05:30
kshitijk4poor 913b910434 fix(vertex): content-digest cache key — stat signature is not credential identity
Security review on #97701 (unsupportedpastels) found and reproduced a
metadata-signature collision: atomic replacement that preserves size
and mtime (deployment tools that restore metadata; equal-length JSON)
rotates the private key under an IDENTICAL (path, mtime_ns, size)
signature, so the cache kept serving the old identity — exactly the
failure the PR set out to close.

The stat idiom is right for config caches (guards a parse, mtime
collisions are harmless). It is wrong for a credential cache (guards
an identity). The key is now (path, sha256(content)):

- _read_sa_file() reads the file ONCE per call, returning both the
  bytes and the digest key; on a miss the credentials are built from
  that same snapshot via from_service_account_info — closing the
  stat->read TOCTOU the reviewer also flagged (key and credentials can
  never describe different bytes).
- Read failure degrades to the bare-path key and the SDK's own file
  read: byte-for-byte the pre-signature behavior.
- Cost: one read + sha256 of a ~2KB JSON per cache probe, noise next
  to the OAuth token mint the cache exists to avoid.

New regression test per the review: equal-length content swap with
mtime restored and atomic replace — asserts a NEW cache key and a NEW
Credentials object (fails on the stat-keyed version, reproducing the
reviewer's cache_keys_equal=True). Fake google-auth harness gains
from_service_account_info.

Suite 10/10; ruff green.
2026-08-29 12:21:49 +05:30
kshitijk4poor ef4cd77ffc fix(vertex): restore ADC->SA retry killed by tuple cache keys (review finding)
The /simplify-code reviewer caught a real regression in the signature-
keyed cache commit: the ADC-failure fallback still compared
`cache_key == "__adc__"` (string), but keys are tuples now — the
comparison is always False, silently disabling the retry that picks up
a service-account file added after startup. The lead's pre-verification
had checked sentinel collision, failure-path pop, and TOCTOU, but
missed this consumer of the OLD key shape.

Guard now tests the actual condition (`not resolved_path` — this
attempt was ADC) instead of a key literal, so it can't rot again if
the key shape changes. New regression test drives the full path: ADC
raises, SA file appears on re-resolution, retry succeeds — fails on
the tuple-comparison version AND on any future key-shape change that
breaks the guard.
2026-08-29 12:21:49 +05:30
kshitijk4poor 56a2623316 fix(vertex): pick up rotated service-account files — signature-keyed creds cache
Pattern-D fix (stale cache after out-of-band change): the Vertex
credentials cache was keyed on the service-account file PATH alone, so
rotating the file on disk (key revoked and re-issued, new identity)
kept serving tokens minted from the OLD Credentials object for the
life of the process. Operators rotate compromised keys precisely when
they most need the new identity to take effect.

The cache key is now the file's (path, mtime_ns, size) signature — the
established idiom (agent/skill_utils.py:414, hermes_cli/config.py:3343,
and the shape #89792 applies to model overrides). Rotation bumps the
signature, forcing one re-read; the superseded entry for the same path
is evicted on insert so the cache stays bounded at one Credentials per
file. ADC keeps a stable sentinel key ("__adc__",) and its existing
expiry/refresh handling; a stat failure degrades to the bare-path key,
i.e. exactly the pre-signature behavior.

Tests (tests/agent/test_vertex_adapter.py):
- rotation invalidates: rewrite + mtime bump -> new key, new
  Credentials object, old entry evicted (fails on main: main serves
  the first identity's object after rotation)
- stat failure falls back to bare-path key, never raises
- ADC sentinel stable across None/empty resolved paths

Suite 8/8 green; ruff green.
2026-08-29 12:21:49 +05:30
Teknium 1d8946b40b fix(prompt-caching): tool-using sessions no longer 400 behind LiteLLM Anthropic proxies (#89886)
LiteLLM OpenAI->Anthropic translation copies tool-message content parts
verbatim, so the envelope-layout part-level cache_control landed at
tool_result.content[0] - a placement the Anthropic Messages schema rejects
with a non-retryable HTTP 400 that killed the whole turn (any tool-using
cron/session on a LiteLLM-fronted Anthropic route).

New envelope_tool_part_cache_markers_supported() predicate (keyed on the
existing _is_litellm_route token matcher) threads a tool_part_markers flag
through build_prompt_cache_plan / apply_anthropic_cache_control and all
four decoration sites (main loop x2, destination replan, MoA). On LiteLLM
routes role:tool messages carry no markers and the breakpoint budget
reallocates to the nearest eligible message; OpenRouter/Nous Portal keep
the part-level form they honor, native Anthropic layout unchanged.
2026-08-29 12:06:48 +05:30
Teknium 217ab2f8df refactor(desktop-tools): consolidate preview + project, diet the desktop_ui suite (3,861 → 2,293 tok/call, −41%) (#97659)
* refactor(desktop-tools): consolidate preview(open/close/read) + project(create/switch/list), diet the desktop_ui suite — 3,861 -> 2,293 tok/call on desktop sessions (-41%)

* rename: preview -> desktop_preview, project -> desktop_project — namespace desktop-app tools against MCP/plugin name collisions

* test: sync remaining old-name pins — per-file registration import, GUI_TOOLS set, post-hook case read_preview -> desktop_preview action=read
2026-08-28 23:10:01 -07:00
StanleyStetson 24e54b55f5 fix(desktop): preserve streamed assistant text and unify atomic persistence (#95514)
- Preserve streamed assistant text in Desktop UI when message.complete delivers empty text.

- Prevent destructive hydration in Desktop useMessageStream over rendered text on empty completion.

- Recover stream buffer in finalize_turn when final_response is empty on healthy turns.

- Unify in-place blank assistant repair, watermark clone resolution, non-blank concurrent winner adoption, and batch row appends into a single atomic guarded SessionDB transaction.

- Synchronize canonical committed content to live in-memory messages dicts and preserve all-or-nothing rollback semantics on persistence failure.
2026-08-29 11:33:05 +05:30
joaomarcos 88a78ecc96 fix(bedrock): recover from server-side cachePoint rejections per placement
Bedrock's cachePoint rules are per-model-family AND per-field. Amazon Nova
accepts a cachePoint block in `system` and `messages` but rejects it inside
`toolConfig.tools`, failing the whole request with

    ValidationException: Malformed input request: #/toolConfig/tools/18:
    extraneous key [cachePoint] is not permitted

so every tool-enabled Nova turn fails, with no retry path and no way for the
user to turn cache markers off (#97281).

The adapter decided placement from one static allowlist that answers only
"does this model cache at all", never "in which section". Any family whose
placement rules differ breaks 100% of turns until someone edits the table and
ships a release — the same maintenance trap the `_NON_TOOL_CALLING_PATTERNS`
comment already admits to ("if a model fails with a tool-related
ValidationException, add it here").

Make Bedrock's own verdict authoritative alongside the table: classify the
rejection by the JSON pointer AWS returns, drop the marker for that one
section, retry the request once, and remember the verdict for the rest of the
process so later turns are built clean. The other sections keep their cache
markers, so Nova still gets system/messages caching instead of losing prompt
caching wholesale. This mirrors the module's existing self-heal idiom
(`is_streaming_access_denied_error` → non-streaming `converse()`).

Applied at all four boto3 call sites: `call_converse`, `call_converse_stream`,
and both Bedrock dispatch sites in `chat_completion_helpers` (the streaming
one is the path in the report). A rejection with no marker to strip returns
None so the caller re-raises instead of looping.

Fixes #97281

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gcoy6nLTg5R6FHHhjcLZEC
2026-08-29 11:33:05 +05:30
Teknium 3951ead838 fix(cache): over-length caller prompt_cache_key no longer 400s Chat Completions requests
Port from anomalyco/opencode#44571: OpenAI caps prompt_cache_key at 64
chars (DeepSeek and Zai inherit the same limit via their OpenAI-compatible
APIs) and rejects longer values with HTTP 400. The Responses transport
already bounds keys via _bounded_prompt_cache_key, but the Chat
Completions transport passed caller-supplied keys (request_overrides,
top-level or extra_body) through unmodified on both the profile and
legacy kwargs paths.

Bound caller keys with the same pck_<sha256[:24]> hash shape codex.py
uses so both transports behave identically; blank keys are dropped
instead of sent empty. Hermes-generated keys were already safe
(content-addressed pck_ hashes).
2026-08-29 11:33:05 +05:30
ericmaddox f0d5f1298b fix(caching): prevent whitespace-only text blocks in prompt cache prefix splits 2026-08-29 11:33:05 +05:30
kshitijk4poor 7eee066c30 refactor: fold simplify-code review findings into #96667 salvage
- Anthropic classifier counts signature_delta and citations_delta payloads
  (content-bearing delta types the transport emits — relay_llm.py handles
  both) so signed-thinking/cited-text generation keeps ticking the fence.
- Fix the stale 'per streamed event' comment above the Anthropic
  on_stream_event lambda (left over from the conflict resolution).
- Rename test_completed_response_without_stream_payload_does_not_tick to
  test_completed_response_ticks_only_terminal_signals — the old name
  contradicted its own assertion (dispatch + shim ticks are expected).
2026-08-29 11:08:11 +05:30
StanleyStetson 7ff70f1709 fix(agent): count only substantive auxiliary progress
Salvages the stream-progress fix from #80122 on current main while incorporating review feedback for empty provider deltas and tool-call scaffolding.
2026-08-29 11:08:11 +05:30
Brooklyn Nicholson d889758521 feat(pet): gate Unicode placeholders to kitty and Ghostty
WezTerm speaks kitty APC but not U+10EEEE, so detect_terminal_graphics()
== "kitty" is the wrong gate for the placeholder path.
2026-08-28 23:38:59 -05:00
rob-maron f7c79efbac add tencent/hy4-preview to model pickers 2026-08-28 19:53:06 -07:00
Teknium 93de1d3430 vision_analyze diet + image routing: explicit aux vision backend becomes the de-facto route (reverses #29135) (#97339)
* refactor(vision_analyze): schema diet — routing mechanics removed (automatic; native path's own result teaches), region flow kept (~271 -> 181 tok/call, -33%)

* feat(image-routing): explicit auxiliary.vision backend is the de-facto image route — reverses #29135 (maintainer decision); native stays default when unset, image_input_mode:native stays absolute
2026-08-28 12:15:24 -07:00
Mariano Nicolini 4d482ed344 refactor(nous): trim comments and drop an unused field 2026-08-28 15:39:23 -03:00
Teknium d3a1c46510 feat(agent): context size anchors on provider-reported usage — estimation shrinks to the last turn
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.

- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
  with a structural base-message identity check that fails closed on any
  transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
  site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
  pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
  compaction (codex_runtime), session reset/switch (run_agent), plus the
  fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
  estimation fallback (first request of a session).
2026-08-28 07:51:31 -07:00
Teknium 5b31602c15 docs: reconcile positional pairing with shared _classify_tool_call_orphans (#97167) — classifier docstring reflects its remaining consumer; empty-id filter note updated 2026-08-28 07:51:23 -07:00
fedebyes 93f4dc7561 fix: make positional prune variant-aware; add replayed-call regression tests
Pass 2 of repair_message_sequence matched results only by id/call_id,
pruning calls answered through response_item_id or composite bridge
ids. Use the shared variant helpers (tool_call_id_variants /
tool_result_id_variants) so the unified alias policy applies
(#55626/#63000/#93251).

The positional sanitizer pass changes the crash/resume duplicate shape:
an interrupted first occurrence is now stubbed instead of deduped, so
the replayed call survives with its own immediate result. Update the
#64335 empty-key test to the new semantics and add regression tests for
the #94704 acceptance shape (historical-result + replayed-call +
fresh-call) and the production interrupted-turn shape (session
7d57a602b83d).
2026-08-28 07:51:23 -07:00
Tiberiu Danciu c7761573f5 fix: prune positionally unanswered tool_calls before API send
DeepSeek v4 rejects a payload where an assistant message carries a
tool_call whose tool result does not follow it immediately (HTTP 400
"An assistant message with 'tool_calls' must be followed by tool
messages responding to each 'tool_call_id'"). Context compression can
displace a tool result past a user turn; the result then lands ~100
messages away from its declaring assistant message.

Two gaps let the poisoned shape reach the wire (reproduced from the
production request dump of session 4d8727cbcf04, replayed through both
functions):

1. repair_message_sequence Pass 1 drops the displaced tool RESULT as
   stray but leaves the declaring assistant message carrying the now
   unanswered tool_call (with empty content) in the durable history.
2. sanitize_api_messages stubbed only globally-absent result ids: the
   displaced result still exists in the transcript, so the id survives
   the set-subtraction, no stub is injected, and the payload 400s.

Fix both layers so every path is order-independent:

- repair_message_sequence: new Pass 2 prunes tool_calls that have no
  result in the immediately-following tool run (matching on id or
  call_id, same superset rule as Pass 1). If pruning empties the turn
  (no content/reasoning left), the whole message is dropped rather than
  sending an empty assistant message. Codex interim turns are exempt,
  as in Pass 0.
- sanitize_api_messages: the orphan/stub logic is rewritten as a single
  rolling positional walk that drops results not immediately following
  their declaring assistant (including results appearing BEFORE their
  call) and injects stub results for positionally-uncovered calls even
  when a mispositioned result exists elsewhere.

Adds six regression tests: repair pruning, whole-turn drop when pruned
calls were the only payload, valid-pair negative control, positional
stub injection, result-before-call orphan drop, and a fully-paired
transcript negative control.
2026-08-28 07:51:23 -07:00