Commit Graph

963 Commits

Author SHA1 Message Date
Teknium 42dd219f46 test: trim salvage of #65076 to a lean regression set
Drop the bulk test additions from the original PR; keep only mandatory
picker-assertion adaptations (Mantle IDs join the discovery lists), one
allowlist routing test covering all four Mantle model IDs, the 272K
context check, and the two review-mandated auxiliary regressions
(config-region-beats-env for the Mantle path, aux Responses client).
2026-08-21 15:02:29 -07:00
Nathaniel Branscum e57d55fc7a fix(moa): keep Bedrock slots on provider runtime
Preserve the Bedrock provider identity for MoA reference and aggregator slots so Bedrock OpenAI Responses models use the aws_sdk/SigV4 runtime instead of being downgraded to a generic custom endpoint. Add regression coverage for Bedrock GPT-5.5 MoA slots.
2026-08-21 15:02:29 -07:00
JoaoMarcos44 fb27614add fix(native_compaction): preserve compression summary messages during pre-checkpoint pruning
prune_pre_checkpoint_items() had a hardcoded role=='user' filter that
discarded all non-user messages before a checkpoint — including Hermes'
own compression summaries (role='assistant'), causing total context amnesia
about past conversation summaries.

The fix:
- _is_summary_item delegates to the canonical
  agent.context_compressor.is_compaction_summary_message provenance check
  (not an ad-hoc heuristic)
- Summaries are retained whole (never byte-sliced) within a 32k token budget
- Idempotent across repeated checkpoints (dedup by identical text)
- _chat_messages_to_responses_input threads item_sources (raw chat messages)
  through to the pruner, so it can read summary content directly from the
  source when the Responses conversion shape is lossy (tool-result carrier
  becomes function_call_output, or stale codex_message_items replay shadows
  merged content)

Fixes #90975.

Salvage of #90976 by @JoaoMarcos44.
2026-08-21 17:24:22 +05:30
kshitijk4poor b883756b79 fix: foreground priority for background review cancel timeout
Change fail-closed behavior to proceed-with-warning when a background
review does not acknowledge cancellation within the bounded deadline.
The review is non-critical self-improvement work and must never block
a user-facing turn (#84423). Keep the off-thread interrupt to ensure
a broken abort path cannot stall the bounded wait.
2026-08-21 16:12:57 +05:30
qixuancao 1b92a94962 refactor(agent): simplify background review run state 2026-08-21 16:12:57 +05:30
qixuancao 37da0d4d50 fix(agent): synchronize background review cancellation 2026-08-21 16:12:57 +05:30
Teknium ca06b87689 feat: opencode-free is fully keyless — no env var, no account, anonymous wire
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).

On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
  never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
  Zen endpoint routing incl. muse->responses); keyless predicate extended
  with unsuffixed free slugs (big-pickle); free runtime pins EVERY
  opencode-free model keyless; curated catalog replaces the models.dev
  cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
  there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
  wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
  invariant (every curated model must satisfy the keyless predicate)

E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
2026-08-21 00:24:32 -07:00
Rudraksh Chahal 28a9b6c565 feat(providers): add OpenCode Free provider with keyed auth and opencode User-Agent
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.

The free tier requires a real account API key and throttles third-party
clients by User-Agent:

- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
  and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
  Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
  unconditionally (the stale keyless-tier assumption), and credential-pool
  exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
  message.

Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
2026-08-21 00:24:32 -07:00
kshitijk4poor c26357ad6a refactor(prompt_caching): shallow strip copy, exact-count guards, dedupe idempotency tests
Follow-up to the #90972 salvage:

- strip loop: copy.deepcopy(msg) -> dict(msg). strip_anthropic_cache_control
  is copy-on-write on content parts by contract (pops the top-level key,
  rebuilds content lists/part dicts fresh), so a shallow top-level copy
  preserves the caller-non-mutation guarantee — verified for all four
  marker shapes — and removes the redundant second deepcopy the re-mark
  path paid on already-decorated input. Docstring updated to match.
- tests: moved the surviving idempotency tests into
  tests/agent/test_prompt_caching.py (where this module's tests live) as
  TestApplyIdempotency; dropped the three tests that duplicated existing
  coverage (dynamic_tool_accounting ~= TestPromptCachePlan::
  test_copies_sections_and_keeps_canonical_tools_plain which already
  asserts == 4; can_carry_marker_envelope_vs_native ~= TestCanCarryMarker;
  never_exceeds_four_markers subsumed by the idempotency test).
- exact-count assertions per review: idempotency fixture pins == 4,
  no-tools fallback pins == 3 (marker loss can no longer masquerade as
  safety); added the one new _can_carry_marker assertion (native=True
  empty assistant) to TestCanCarryMarker.
- new part-level stale-marker mutation guard (the other detection branch,
  where part-dict aliasing is the risk); fails on pre-fix base with
  marker accumulation (9 > 4), passes with the fix.
2026-08-21 12:23:47 +05:30
joaomarcos 0fc52b055f fix(prompt_caching): make apply_anthropic_cache_control idempotent on pre-decorated input
apply_anthropic_cache_control never stripped pre-existing cache_control
markers before placing new ones, so calling it twice (or handing it
messages a prior call already marked) accumulated markers past
Anthropic's 4-breakpoint limit and produced HTTP 400
'cache_control can only be specified up to 4 times'.

Strip any pre-existing markers from per-message copies before marking,
mirroring the strip-then-mark pattern build_prompt_cache_plan already
uses. Only messages that already carry a marker pay the copy cost; the
copy-on-write contract (caller-owned messages are never mutated) is
preserved. Repeated calls now converge to byte-identical output.

Salvaged from #90972 by @JoaoMarcos44 (net diff of the PR's commit
stack, intermediate reverts collapsed).

Related: #90971
2026-08-21 12:23:47 +05:30
kshitijk4poor 2cf7b36e11 fix(memory): enforce independent built-in store permissions
Normalize malformed memory config during initialization and bind per-target write permissions to the session MemoryStore so direct and staged writes cannot update a disabled built-in store.
2026-08-20 20:20:23 -07:00
Brooklyn Nicholson c57581cd0d feat(tools): drive_preview and annotate_preview — the agent can use the page it opened
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.

Four pieces, and they only make sense together:

  · an in-page engine that inventories what is interactable and performs the
    verb, injected as source because it has to run inside the guest page;
  · the preview.act.request bridge from the gateway into the pane;
  · drive_preview, for acting: elements, click, type, scroll, press, and the
    pane's own back/forward/reload;
  · annotate_preview, for marking without acting.

Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.

Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.

Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
2026-08-20 05:26:37 -05:00
Teknium c2f5d2da21 test: vary marathon-turn fixture args — identical calls now legitimately dedupe to stubs 2026-08-20 00:16:22 -07:00
kshitijk4poor b7e12decc6 fix(agent): route relay-wrapped output-cap 429s into the output-cap handler
Salvage follow-up for #72283: instead of a second pre-retry clamp block
(which bypassed the #55546 clamp+compress path and broke its three
regression tests), parse the output cap ONCE at classification time and:
- exempt parseable wrapped output-cap 429s from the eager rate-limit
  provider fallback (a deterministic request-shape failure that failover
  cannot fix but the clamp fixes in one retry), and
- widen is_context_length_error so they reach the SAME #55546
  clamp+compress recovery as plain output-cap 400s.

Adds both #72283 regression scenarios plus an ordering guard proving a
NON-EMPTY fallback chain does not consume the wrapped 429 (fallback
slot unspent, model unchanged). 119 fallback/rate-limit tests green.
2026-08-20 11:37:01 +05:30
kshitijk4poor 0596ccdeb3 fix(compression): salvage follow-up — todo snapshot last-resort, reuse prune helpers
Review follow-up on the salvaged #90353:
- Todo snapshot (+ coupled pruned-skill reload notice, 7a16840add) is now
  reduced only as a LAST resort after reasoning/tool/summary shrink ops,
  and the reload notice survives even then.
- Reuse existing helpers/constants instead of re-hardcoding:
  _PRUNED_TOOL_PLACEHOLDER, _PRUNE_MIN_CHARS, _NEWEST_TURN_ONLY_BUDGET_KEYS,
  and _prune_stale_reasoning_replay (codex sidecar shrink, #71058 boundary).
- Assistant-role messages without the summary metadata key are no longer
  truncatable by the summary-cap heuristic.
- Caller passes budget so the estimator runs 3x, not 5x, per would-grow pass.
2026-08-20 11:36:54 +05:30
MindDragonLabs fb96247eaf fix(compression): salvage grown candidates before refusal 2026-08-20 11:36:54 +05:30
Teknium 481bc9391e fix(memory): profile-only config gets narrow USER_PROFILE_GUIDANCE instead of the full memory block
With memory_enabled: false but user_profile_enabled: true, the memory tool
stays (it backs USER.md) but the full MEMORY_GUIDANCE told the model to save
notes to a MEMORY.md store that does not exist. Split the guidance: a
profile-only block is injected for that configuration, directing writes to
target='user' only.
2026-08-19 22:59:17 -07:00
HexLab98 a969c5a93d test(memory): cover the disabled built-in memory surface
Walks the real resolution chain -- config.yaml on a temp HERMES_HOME ->
check_memory_requirements -> get_tool_definitions -- rather than mocking
the availability check, since the bug was in how the flags reach the
schema. Covers both flags off, either one alone, no config file at all,
and a config read that raises (must fail open).

Also asserts the external provider's tools survive with the built-in tool
gone, so the fix cannot regress into taking Hindsight/Mem0 down with it,
while disabled_toolsets keeps its documented "hide everything" meaning.

The existing MEMORY_GUIDANCE test built a skip_memory agent whose flags
were both false, so it was asserting the old tool-presence-only behavior;
it now states its precondition and gains the false-case mirror.
2026-08-19 22:59:17 -07:00
Teknium f4a866b484 fix(cli): /config displays the live agent credential, not the env-var constructor seed 2026-08-19 19:42:07 -07:00
Teknium d762ed9b3c feat: execution-discipline guidance now reaches all tool-capable models (config model.execution_guidance)
Un-fences OPENAI_MODEL_EXECUTION_GUIDANCE from the gpt/codex/grok substring
check and gives it its own injection gate, independent of
tool_use_enforcement, controlled by config.yaml `agent.execution_guidance`
(auto/true/false/list — same semantics as tool_use_enforcement). The "auto"
list (EXECUTION_GUIDANCE_MODELS) now also covers deepseek, kimi, qwen, glm,
minimax, mimo, and mistral.

Composio agentic-eval traces showed Hermes+DeepSeek/Kimi failing where
competitors passed: financial math done in prose, no read-back after
external writes, malformed identifiers "repaired", completeness claimed
despite count mismatches. The discipline block existed but those models
never received it.

The block is extended with compact clauses distilled from that analysis:
- external-write read-back (tool-call success is not task success; internal
  file edits already confirmed by the tool are not re-verified)
- count reconciliation (declared totals/has_more are hard assertions)
- literal preservation (never normalize identifiers that fail a stated
  format; lookup success does not validate a malformed token)
- retry-differently (empty/partial/suspiciously narrow results get a
  broader retry before concluding)
- completion gated on verification (done = every named acceptance
  criterion verified, never a plausible subset)

The todo tool description now encourages enumeration-as-checklist for
"all N items" tasks and gates completed status on verified work, never
intent.

Guidance is chosen once at session start keyed on model name, so the
system prompt stays byte-stable for the life of a conversation.

Supersedes/absorbs prior contributor proposals: #20588, #35087, #41874
(MiMo), #53847 (GLM tool-calls-as-text stall).

Co-authored-by: Mat-London <56627804+Mat-London@users.noreply.github.com>
Co-authored-by: intelac <8803887+intelac@users.noreply.github.com>
Co-authored-by: 6ylqq <51219463+6ylqq@users.noreply.github.com>
Co-authored-by: tauros1983 <267660491+tauros1983@users.noreply.github.com>
2026-08-19 16:16:37 -07:00
Falko 8430c1b4da test(codex): cover nested retention entry paths 2026-08-20 03:15:14 +05:30
kshitijk4poor f6d1d774a1 refactor(codex): name the dropped shape in the wire-guard warning
Review polish from the 3-angle pass on the final stack:

- The warning now says WHICH shape leaked (top-level, extra_body, or both).
  Relay injects top-level while request_overrides typically inject via
  extra_body, so the shape identifies the offending middleware when
  debugging.
- Fold the 'always returns a fresh mapping' assertion into the parametrized
  real-endpoint test (the caller mutates the result with stream=True, so the
  copy contract is load-bearing on no-drop paths too) and drop the
  SimpleNamespace stub test it strictly subsumes. The nested-preserve stub
  stays: the parametrized test only exercises top-level retention.
2026-08-20 03:15:14 +05:30
kshitijk4poor 26530e7df5 fix(codex): strip nested extra_body retention at the consumer Codex wire
The wire guard only removed the top-level prompt_cache_retention kwarg, but
the OpenAI SDK merges extra_body into the outgoing JSON body, so a nested
extra_body.prompt_cache_retention reaches chatgpt.com/backend-api/codex just
the same and still triggers the non-retryable HTTP 400. Both injection
vectors are real and probe-verified: the Relay overlay's 'key not in
baseline' arm admits an interceptor-added extra_body, and
request_overrides={'extra_body': {...}} lands verbatim in build_kwargs
output.

Close the gap in the same helper: strip the nested field too (copy-on-write,
never mutating the caller's mapping), drop extra_body entirely when it
empties, and log the same warning. Compatible endpoints keep nested
retention untouched.

Mutation-verified: removing the extra_body leg fails both new nested tests.

Reported by egilewski's review on #89969.
2026-08-20 03:15:14 +05:30
kshitijk4poor ba4bc39afd test(codex): pin the retention drop to real endpoints
The salvaged compatibility test stubs `_is_codex_backend=lambda: False` on a
SimpleNamespace, so it proves the helper honors its own boolean but not that
the boolean is right for any real endpoint. A predicate change that widened
the drop onto retention-supporting hosts would keep it green.

Adds a parametrized test that builds a real AIAgent per base URL and asserts
the drop only fires for chatgpt.com/backend-api/codex, while api.meta.ai,
bedrock-mantle.*.api.aws, api.openai.com and a same-host/different-path
backend keep their supported 24h value. Also asserts prompt_cache_key
survives untouched on every endpoint, since retention and cache-key routing
are independent and the guard must not disturb caching.

Verified non-vacuous: relaxing the guard's condition to drop on every
endpoint fails 4 of the 6 cases (Meta, Bedrock, OpenAI, non-codex path).

Drive-by on the guard itself: drop the dead `None` default on the `pop` that
is already gated by an `in` check, and record why the predicate is resolved
via getattr -- run_codex_stream is driven with lightweight stand-in agents
that lack `_is_codex_backend`, so a bare call would raise AttributeError.
2026-08-20 03:15:14 +05:30
Falko 8e2949495a fix(codex): strip unsupported cache retention at wire 2026-08-20 03:15:14 +05:30
Caleb DeLeeuw 4ee8a9a169 test: production-shape resolver + init-read coverage for reasoning_echo
Salvaged from #73811 per the consolidation triage on #76503, adapted to this
PR's per-active-provider model.reasoning_echo design.

Existing reasoning_echo tests hand-set _reasoning_echo_flag; none drives the
real config path init_agent uses:
    load_config_readonly().get("model").get("reasoning_echo")
which is wrapped in `except Exception: False`, so a broken read would silently
disable the feature untested. This adds a temp-HERMES_HOME test that resolves a
named custom provider via the real resolve_runtime_provider (asserting
provider == "custom"), materializes the flag from a real config file, and
checks reasoning_content is preserved with the flag on and stripped with it off.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ab7R4NizLNbNyQZShtciKg
2026-08-20 00:03:50 +05:30
Brooklyn Nicholson 93b50ea0bb test+docs(tour): cover the engine and document the API
Ten behavior tests for target discovery, stable-selector ordering, step
paging, recovery hints, and the self-containment contract the preview
injection depends on. Docs cover data-tour markup and the curated-tour
entry point alongside the tool itself.
2026-08-19 00:52:54 -05:00
Teknium 210cdb0ed3 fix(agent): legacy hidden redirect placeholders get the neutral wire payload at projection time (#88955)
The salvaged writer-side fix stamps api_content on NEW hidden redirect
placeholders, but rows persisted before it (content="" + display_kind=hidden,
no sidecar) would keep re-triggering repair_empty_non_final_messages on every
call forever. Substitute [response interrupted] on the wire copy at the
api_content/display_kind projection stage so legacy sessions converge too.
Never the interrupt scaffold (#81841). Durable transcript untouched.

Regression tests drive run_conversation end-to-end with a spied sanitizer:
the projection must leave the sanitizer nothing to heal (its per-turn warning
spam is the bug), verified failing via sabotage run against the writer-only
fix.

Projection-side approach credit: @JoaoMarcos44 (PR #88996).
2026-08-18 16:27:32 -07:00
Axl Ibiza, MBA 0ee9bc8d1e fix(agent): give interrupted-turn hidden placeholder a neutral provider-replay sidecar
Bot-mode interrupted member turns with no visible assistant text persisted an
empty assistant row (content="" + display_kind="hidden"). The pre-call
sanitizer repair_empty_non_final_messages() re-healed that row on every later
call (wire copy only), so the loop never converged (#88955).

Stamp api_content="[response interrupted]" (the canonical
_INTERRUPTED_PLACEHOLDER) on the hidden placeholder instead. display_kind is
stripped before sanitization, but api_content is projected back into content
for historical assistant rows, so the provider sees a non-empty neutral turn
and the sanitizer stops touching the row — while the durable transcript stays
hidden and empty. Uses the neutral interruption text, never the
_INTERRUPTED_SCAFFOLD_MARKER, which replaying as assistant text caused #81841.

Adds regression coverage proving (A) the placeholder carries the replay
sidecar, (B) two consecutive projections converge without sanitizer healing,
(C) the sanitizer still repairs genuinely-empty unmarked assistants.

Refs #88955
2026-08-18 16:27:32 -07:00
Jeffrey Quesnelle 9664e386f6 Merge pull request #85581 from bbednarski9/codex/fix-openai-sparse-response-objects
fix(openai): tolerate sparse response objects
2026-08-18 11:42:24 -04:00
Teknium 83d451e7b5 test: budget-replacement assertions accept persisted-output blocks
The steer-survives-budget tests pinned the inline 'Truncated:' fallback
shape, which only occurred because env=None persistence was broken.
Now that host-side spillover succeeds, budget enforcement produces a
<persisted-output> block instead. Assert the actual contract — the
oversized payload was replaced (persisted OR truncated) — via a shared
helper, not which replacement shape was used.
2026-08-18 02:12:13 -07:00
Bryan Bednarski a51337ecee Merge main into fix-openai-sparse-response-objects
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-17 18:58:48 -07:00
kshitij 979ca57a50 Merge pull request #88244 from kshitijk4poor/fix/handoff-cleanup-race
fix: prevent handoff leg data loss + surface state.db corruption to users
2026-08-17 17:41:44 +05:30
Teknium df68cc1c1a test: adapt stale dedup test to outstanding-call semantics, add cross-turn coverage
Follow-up to the salvaged #70734 fix:

- test_sanitize_dedup_drops_tool_calls_key_when_all_removed encoded the old
  global-uniqueness assumption (its second assistant call reused the id AFTER
  the first call was answered, which is now a legitimate new call). The
  replayed call now precedes the result, making it a true duplicate of a
  still-outstanding call, preserving the intended empty-tool_calls key-drop
  coverage from #64335.
- New test: Hermes' own deterministic local counter ids repeating across
  turns (the #76632 scenario) survive sanitization.
- New test: the 50-step constant-id field repro from #70724 (Kimi K3 /
  llama.cpp) — stock main kept 1/50 tool results, now 50/50.
2026-08-17 03:25:07 -07:00
flaviovargasbrandao 0b8fd04bea fix(agent): keep tool results when a server reuses one tool_call_id
The #58327 dedup passes treat a repeated tool_call_id as garbage from a
retry/crash/resume glitch and drop it. That assumes tool_call_id is
globally unique, which it is not: llama.cpp emits a single constant id
for every tool call it ever returns (verified — three separate
completions from one server all carried the same id).

Under a seen-once-drop-forever rule, the SECOND legitimate tool result
of such a session looks like a duplicate and is deleted. From the second
tool call onward the model never sees any result: it announces its next
action, the turn ends, and the task is left unfinished. Bisected to
dba585c17 over a 2258-commit range; reproduced live on v0.19.0 (1/6 runs
completed a 4-step file task, vs 20/20 on the last release before that
commit, same model and server).

Key off OUTSTANDING calls instead of every id ever seen. Both original
protections are preserved: a replayed result still answers no pending
call and is still dropped, and duplicate tool_calls sharing an id within
one assistant message are still collapsed. A genuine new call that
reuses the id re-arms it first.

repair_message_sequence needs no change — it already resets its id set
per assistant message, so only the final pre-API pass mis-fires.

Live result after the fix: 8/8 runs complete, 17-26s each (was 1/6 with
runs hitting a 150s ceiling).
2026-08-17 03:25:07 -07:00
kshitij 59f302fef9 fix: prevent handoff leg data loss + surface state.db corruption to users
Two data-loss bugs reported by users:

1. /handoff CLI→gateway race (#88234): After /handoff completed, CLI
   cleanup called finalize_session on the session the gateway just
   reopened. This set end_reason on a row the gateway was actively
   writing to, causing the handoff leg to vanish from session history
   and breaking session_search recall. Fix: add _handed_off_session_ids
   module-level set (mirrors _single_query_finalize_attempted_session_ids
   pattern). _handle_handoff_command registers the session_id on
   completion; _should_emit_cleanup_session_finalize and
   _emit_interrupted_session_end check it before firing.

2. state.db corruption silent failure (#88235): When SessionDB init
   failed at gateway startup, the error stayed in logs — messages
   flowed but nothing was persisted, with no user-visible indication.
   Fix: store _session_db_init_error on GatewayRunner, broadcast a
   recovery-guidance message to all home channels via
   _send_session_db_warning_notifications() after the gateway connects.
   Also improved the 'corrupt' persistence cause wording in
   _format_turn_completion_explanation to include the full recovery
   path (hermes doctor --fix, sqlite3 .recover, backups).

Tests: 6 new tests for handoff cleanup race, 3 for corruption wording.
All existing CLI/turn-completion tests pass.
2026-08-17 13:19:58 +05:30
Shannon Sands d10f87245e refactor(agent): move empty-response guard settings from env vars to config.yaml
Per project policy, .env / HERMES_* env vars are reserved for
credentials; behavioural settings belong in config.yaml. Replaces
HERMES_DETERMINISTIC_EMPTY_GUARD and
HERMES_EMPTY_RETRY_COST_THRESHOLD_USD with an additive
agent.empty_response_guard section:

  agent:
    empty_response_guard:
      enabled: true            # false = legacy fixed 3-retry behaviour
      cost_threshold_usd: 0.25 # per-attempt cost that halves the budget

- hermes_cli/config_defaults.py: new documented subsection under agent
  (additive key, no config-version bump needed).
- agent/empty_response_guard.py: resolve_guard_settings() maps the
  section to (enabled, threshold) with fail-open tolerance for
  malformed values; guard_enabled()/_cost_threshold_usd() now read the
  init-resolved agent attributes instead of os.environ.
- agent/agent_init.py: resolves the section once at init into
  agent._empty_guard_enabled / agent._empty_guard_cost_threshold_usd,
  following the existing tool_use_enforcement extraction pattern.
- Tests updated to config-attr injection; new TestResolveGuardSettings
  covering malformed sections, YAML string booleans, bad thresholds,
  and a DEFAULT_CONFIG sync check; new integration test proving
  enabled:false restores the legacy 1+3-call behaviour.

Requested by isak-ialogics on PR #75115.
2026-08-16 22:06:08 -07:00
Shannon Sands ac06c2ff8b fix(agent): stop re-billing deterministic empty responses (NS-503)
Every empty-response retry re-sends the full conversation input at full
price. On large contexts a single turn that produces no visible output
could bill the user several dollars across the 3-retry + fallback-chain
walk (reported: ~$2.33 for one empty answer on a ~26K-token session).

Signaled refusals (finish_reason=content_filter, Anthropic refusal
stop_reason, guardrail interventions) are already terminal today and
never reach this loop. The uncovered class is *unsignaled* refusals:
the provider returns 200 with zero output tokens and a generic finish
reason. Those are deterministic — resending the identical prompt
reproduces the same empty — so burning the remaining retry budget only
multiplies the charge.

New agent/empty_response_guard.py, two independent guards, both failing
OPEN to today's behaviour:

- Deterministic-empty detection: two consecutive empty attempts with
  usage present, output_tokens == 0 (reasoning tokens count as output),
  and identical (model, provider, finish_reason) skip the remaining
  retries and go straight to the fallback chain — a different model may
  well answer. Missing usage, nonzero output, or any signature change
  keeps the full budget.
- Cost-aware retry budget: when one attempt's estimated input cost
  exceeds HERMES_EMPTY_RETRY_COST_THRESHOLD_USD (default $0.25), the
  empty-retry budget drops 3 -> 1 for that streak. Unknown pricing or
  included/subscription routes are untouched.

At exhaustion the status trace now includes the estimated cost of the
empty attempts so the charge is at least explained in-session.

Streak state lives on the agent and self-clears whenever
_empty_content_retries resets to 0, transparently honouring every
existing reset site (turn start, tool success, compaction, fallback
activation) without touching them.

Set HERMES_DETERMINISTIC_EMPTY_GUARD=0 to disable both guards.

Tests: tests/agent/test_empty_response_guard.py (26 unit tests) plus
two loop-level integration tests in tests/run_agent/test_run_agent.py
proving the api_call reduction and the fail-open path.

Refs NS-503.
2026-08-16 22:06:08 -07:00
Teknium 250232ff91 Revert "fix(agent): harden canonical tool call deduplication"
This reverts commit 8fc4189edd.
2026-08-16 10:51:11 -07:00
Teknium 06b9141109 fix(state): classify structural DB corruption as its own persistence cause
'database disk image is malformed' contains the word 'disk', so
classify_persistence_error bucketed SQLITE_CORRUPT / SQLITE_NOTADB
failures as 'disk' and the turn-completion explainer told users to
free disk space for a structurally damaged state.db (the #77386-family
misdiagnosis, reproduced in the v0.20.0 malformed-DB incident report).

- hermes_state: new 'corrupt' bucket in PERSISTENCE_ERROR_CAUSES,
  matched via _DB_CORRUPTION_MARKERS BEFORE the locked/disk buckets
- run_agent: explainer text for 'corrupt' points at hermes doctor and
  explicitly says freeing space will not help
- cron explainer-variant suppression picks the new variant up
  automatically (it iterates PERSISTENCE_ERROR_CAUSES)
2026-08-16 10:34:23 -07:00
Teknium d709d29f19 fix(agent): trim background_review to the enabled switch
Follow-up to #87400: drop the max_iterations and prompt_file knobs from
auxiliary.background_review. The aux model routing (provider/model/
base_url/...) predates #87400 and stays; the enabled switch and the
usage telemetry stay. The fork's iteration budget returns to the
historical hardcoded 16.
2026-08-16 10:27:52 -07:00
Ojas Sharma 7095e23eb2 fix(agent): attribute background-review usage and add cost controls
Persist fork token usage under session_model_usage task=background_review,
emit a per-fork completion log line, and expose enabled/max_iterations/
prompt_file so operators can see and bound the automatic review cost.

Address review feedback: load auxiliary.background_review once per spawn,
classify completion logs by summarize action prefixes, treat explicit
api_call_count=None as the documented default of 1, and WARNING on the
fail-open enabled-gate path.
2026-08-16 06:38:38 -07:00
kshitij 4ce0d64be4 test(caching): pin signal precedence and the openrouter-host opt-out
Review follow-up. Three coverage gaps in the LiteLLM matrix:

- The operator opt-out on a litellm-named provider pointed at an OpenRouter
  host. That route previously took the OpenRouter branch and ignored an
  explicit per-model `prompt_caching: false`; it is the only cell in the
  differential matrix where the salvage REMOVES caching, so pin it as
  intended rather than leaving it to be read as a regression.
- Signal precedence: an explicitly litellm-named provider grants even on a
  lookalike host, because the provider id is an independent signal and only
  the host-derived signal is token-gated. Intentional, now documented.
- A hyphen-delimited host label (`my-litellm-gw.internal.example.com`),
  which the token matcher handles but nothing exercised.

Traded the redundant `claude-3-7-sonnet` parametrize cell for the new host
case, so the matrix covers more shapes with the same cell count.

Tests: 83 passed. All three production fixes re-mutation-checked against
the final stack.
2026-08-16 14:34:30 +05:30
kshitij 435d6f30b5 fix(caching): match the litellm provider id token-wise too
Self-review follow-up. The previous commit fixed substring matching on the
HOST but left the provider-id side as a bare substring, so a user-named
provider like `custom:notlitellm` or `mylitellmthing` still matched and was
handed Anthropic markers — the same bug class, half-fixed.

Both signals now match `litellm` as a whole delimited token via a shared
helper. Real spellings (`litellm`, `custom:litellm`, `litellm-router`, and
the already-lowercased `LiteLLM`) still match; lookalikes no longer do.

Tests: 71 passed. Adds lookalike-provider and real-spelling guards; both
new guards mutation-checked. Differential matrix over 2688 configs vs
origin/main: 60 changes, every one a Claude model on a genuine LiteLLM
route getting the envelope layout, zero pre-existing routes altered.
2026-08-16 14:34:30 +05:30
kshitij 1b0e953b47 fix(caching): use the envelope layout for LiteLLM Claude on the OpenAI wire
Follow-up to the salvaged LiteLLM cache grant. The grant itself is right;
four things about how it was scoped were not.

1. Layout. The branch returned the native inner-block layout
   (use_native_layout=True) on api_mode == "chat_completions". That layout
   writes a TOP-LEVEL msg["cache_control"] on role:tool and empty-content
   messages and depends on the Anthropic adapter to relocate it into the
   block — but that adapter only runs for api_mode == "anthropic_messages"
   (agent/transports/anthropic.py registers there), and the
   chat_completions transport does no relocation. Measured on a 3-tool-turn
   transcript: 2 of the 4 available breakpoints landed on markers the
   provider never sees. Worse, when LiteLLM itself relocates a top-level
   marker for an OpenRouter-backed Claude route
   (OpenrouterConfig._move_cache_control_to_content), the marker lands on
   an empty assistant turn and produces a cache_control-marked empty text
   block — the HTTP 400 "text content blocks must contain" shape already
   guarded in agent/anthropic_adapter.py (#69512). Switched to the envelope
   layout, matching every other OpenAI-wire grant in this function:
   4 of 4 breakpoints honored, zero empty blocks.

2. Host matching. `"litellm" in base_url_hostname(...)` is the substring
   false-positive class base_url_hostname's own docstring warns against; it
   granted Anthropic markers to notlitellm.example.com,
   foolitellmbar.example and friends. Replaced with a label-token match in
   a named helper, so "litellm" must be a whole dot- or hyphen-delimited
   token. All three of the original test hosts still match; a "litellm"
   path segment on an unrelated host still does not.

3. Transport gate. `not is_anthropic_wire` also swept in codex_responses,
   bedrock_converse and codex_app_server. Gated on
   api_mode == "chat_completions" explicitly.

4. Operator override. The grant is inferred from a provider/host name, but
   the custom-provider capability lookup was gated on is_anthropic_wire, so
   an explicit `prompt_caching: false` for the route+model was honored on
   /v1/messages and silently ignored on /v1/chat/completions. The lookup
   now also runs for a LiteLLM route, and its layout follows the transport
   rather than the declaration (an explicit `true` must not promote a
   chat_completions request to the native layout).

Tests: 64 passed. Adds the wire-shape contract the original matrix was
missing (asserts no breakpoint sits on the message envelope, rather than
only checking the returned tuple), plus lookalike-host, other-transport,
and both operator-override directions. All five guards mutation-checked —
reverting each fix turns the corresponding test red.
2026-08-16 14:34:30 +05:30
Justin Bowes ff4df5e54d fix(caching): engage prompt caching for LiteLLM Claude on the OpenAI wire
anthropic_prompt_cache_policy() only granted Anthropic cache_control
markers to LiteLLM over the native Anthropic wire
(api_mode == "anthropic_messages"). A LiteLLM deployment exposing the
OpenAI-compatible surface instead (/v1/chat/completions, /v1/messages
-> 404) matched no grant branch and fell through to (False, False): no
cache_control injected, the system prompt sent as a plain string, and
the provider serving zero cache hits -- the entire prompt re-billed at
full price on every turn. Silent: no error, no warning, usage simply
shows 100% uncached input forever.

Add one branch after the is_anthropic_wire/is_claude case that grants
caching to Claude-family models on a LiteLLM endpoint regardless of
wire, with the native inner-block layout. Same failure class already
documented in-function for Qwen/DashScope.

Design:
- Gated on the Claude family only (is_claude); a Gemini/GPT/Qwen route
  through the same proxy must not receive markers (they may reject the
  cache_control block format -- cf. the DeepSeek/OpenCode exclusion).
- Matches on provider string OR base_url host, since provider naming
  varies per install (litellm, custom:litellm, or a bare custom alias
  pointed at a LiteLLM host).
- prompt_caching.cache_ttl: false still wins (the _cache_disabled early
  return is untouched).
- Generic strict OpenAI-wire custom providers (e.g. Fireworks) remain
  excluded -- verified by the existing over-reach regression test.

Tests: adds TestLiteLLMOpenAIWire covering the grant (several model
spellings x provider/host signals), no-over-reach (non-Claude on the
same proxy get nothing; operator disable wins), and adjacent behavior
(LiteLLM in Anthropic proxy mode still native layout). Full module:
43 passed.

Closes #84506. Original diagnosis, patch design, and measurements by
@ottosulin.
2026-08-16 14:34:30 +05:30
fangliquanflq 8fc4189edd fix(agent): harden canonical tool call deduplication 2026-08-16 01:59:08 -07:00
fangliquanflq a55d29d9fb fix(agent): canonicalize duplicate tool call arguments 2026-08-16 01:59:08 -07:00
Martin Koistinen da392043a9 fix(agent): include timezone and UTC offset in system prompt timestamp
The "Conversation started:" line carried a bare date (%A, %B %d, %Y). Tools
that accept instants -- nutrition, calendar and similar MCP servers -- reject
naive datetimes and require an explicit UTC offset, so the model had to infer
EST vs EDT from the date alone. Near a DST boundary that is a coin flip, and a
wrong guess does not error: it silently writes the record onto the wrong day.

Append the IANA zone (when configured), the zone abbreviation and the UTC
offset, e.g.:

  Conversation started: Saturday, August 15, 2026 (America/New_York, EDT, UTC-04:00)

get_timezone() returns None when no timezone is configured; in that case the
line falls back to the abbreviation and offset of the server-local (still
tz-aware) time, so behaviour is unchanged for users who never set one:

  Conversation started: Saturday, August 15, 2026 (EDT, UTC-04:00)

Daily byte-stability is preserved -- the property the date-only format exists
to protect (PR #20451). Zone name, abbreviation and offset are all constant for
the whole day; they shift only at a DST transition, where a change is correct.
The static-prefix reconstruction guard in _restore_plugin_sections matches on
"\n\nConversation started:" and is unaffected by a suffix after the date.

test_datetime_is_date_only_not_minute_precision used `re.search(r":\d{2}")`
over the whole line as a proxy for "no time-of-day". A UTC offset also matches
that pattern, so the check now applies to the date portion (everything before
the zone parenthetical) and the invariant is tightened rather than relaxed:

- test_datetime_includes_utc_offset asserts the offset is present
- test_datetime_line_is_stable_across_rebuilds asserts two rebuilds in the
  same day produce a byte-identical line

Fixes #87403

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 01:57:59 -07:00
webtecnica b48ab1b4ad fix(agent+discord): guard truncated-response continuation loops and cap Discord split delivery (#86581) 2026-08-16 01:55:51 -07:00