Commit Graph

326 Commits

Author SHA1 Message Date
Alexander Prendota 613164dadc fix(acp): key the ACP runtime exclusions on the scheme, not on one vendor
An ACP client talks to a CLI over subprocess stdio: it returns a plain
completion object rather than an iterable stream, and it does not implement
the Responses API surface. Both exclusions spelled out `acp://copilot`, so the
next ACP client silently inherited the wrong defaults — a Responses upgrade
its shim cannot serve, and a streaming call that tries to iterate a
`SimpleNamespace`.

Match on the `acp://` scheme instead. `acp+tcp://` was already handled this
way; copilot-acp's behaviour is unchanged, and the new tests pin that a
non-ACP URL still upgrades, so this is not a blanket opt-out.
2026-08-26 10:10:11 -07:00
Alexander Prendota 07200e9cd6 feat(agent): fold an agent-as-provider's own tool work back into the turn
Most providers are models: they ask Hermes to run a tool and Hermes runs it,
so the transcript and the loop's counters see every tool iteration. Some
providers are agents — an ACP CLI behind a client shim, or the codex
app-server, which already takes an analogous path in `agent/codex_runtime.py`.
They execute their own read/edit/execute tools inside their own session, and
by the time Hermes sees the response that work is done.

Those calls must never come back as pending `tool_calls` — Hermes would
re-run finished work. But summarising them into `reasoning` blinds two
subsystems:

- the self-improvement loop, which distils memories and skills by replaying
  `messages`; a one-line activity feed teaches it nothing;
- the skill-review nudge, whose `_iters_since_skill` counter only moves on
  Hermes tool iterations, of which there are none.

So a client may hand both back on the completion object —
`hermes_projected_messages` (completed assistant(tool_calls) + tool(result)
rows) and `hermes_provider_tool_iterations` — and
`splice_provider_projection` applies them. Rows go through `append_message`
like every other live-transcript append, so they carry a timestamp and
persist the same way the codex projection path's rows do.

The splice is append-only, sits before this turn's assistant message so the
order reads call -> result -> answer, and is a no-op for every client that
sets neither attribute, i.e. every ordinary OpenAI-compatible provider.
Garbage attribute values are tolerated rather than allowed to break the turn.
2026-08-26 10:10:11 -07:00
Gille 21b92d2687 fix(agent): bypass response cache for empty retries 2026-08-23 23:31:54 -07:00
Finn763 74e6885f0d fix(review): fail-closed compressor detachment + warm-cache first request (#93057 review)
Adversarial-review fixes for the #93057 snapshot-compaction PR:

- Fail-closed detachment: only re-enable compression after
  bind_session_state successfully severs the engine's parent binding.
  A failed rebind keeps the historical compression_enabled=False
  behavior and warns, instead of running compaction against a
  compressor still bound to the parent's SessionDB (#38727 re-open).
- Warm-cache parity: defer both compression gates (turn-prologue
  preflight + pre-API pressure check) until the fork's first provider
  response, so the first request replays the full snapshot as the
  intended cached read and compaction applies from the second request
  on — matching the documented budget mental model.
- Tests: regression for the rebind-failure fail-closed path (red on
  pre-fix code) and the existing threshold-crossing test reworked to a
  two-request review asserting the warm first request + compacted
  second request. 116 tests green across all touched suites; ruff
  clean.
2026-08-23 18:25:19 -07:00
Finn763 4202a508fd fix(review): bound same-model background review replay
Detach the review fork's compressor from the parent SessionDB/session_id
and re-enable in-memory-only compaction for oversized snapshots, instead
of the historical compression_enabled=False guard that left the fork's
replayed transcript unbounded (350k-384k input tokens per request, 1.49M
total across one 8-request review). Add an aggregate input-token budget
(auxiliary.background_review.max_input_tokens, default 600k) so repeated
tool calls cannot recreate an unbounded transcript; the tool loop stops
before the provider call that would cross it.

Closes #93057
2026-08-23 18:25:19 -07:00
joaomarcos 5496d5995a fix(agent): preserve tool results across ID variants
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
2026-08-23 18:24:43 -07:00
kshitijk4poor b8cd00f968 fix: set failed=True for repeated_outer_errors exit + drop append_message
Follow-up to @BrunoBza's #93062:

1. Set failed=True only for the new repeated_outer_errors exit reason.
   Previously the error exit left failed=False, so finalize_turn reported
   completed=True for a turn that actually failed — incorrect.

2. Don't append_message the assistant response at the break. A thinking-
   prefill or interim assistant may already be the tail, and appending
   would create assistant→assistant role-alternation violation.
   finalize_turn (lines 341-353) handles this safely by checking
   _tail_role != 'assistant' before appending.

3. Update test to assert failed=True and completed=False for the
   repeated_outer_errors exit.
2026-08-24 04:47:07 +05:30
BrunoBza 56e7fd2adf fix(loop): bound outer-loop error retries per turn instead of relying on max_iterations (#92450)
The outer conversation-loop except handler only left the loop on a
local-processing error or when api_call_count >= max_iterations - 1.
With the turn budget now unlimited by default (sys.maxsize), a
permanent failure that escaped the inner retry/fallback machinery
retried forever: ~64 retries/s, one core pegged, and the rotated
agent.log history overwritten within minutes.

Bound the loop with a small per-turn cap on total escaping exceptions
(_MAX_OUTER_LOOP_ERRORS = 8, scaled down by a tiny explicit
max_iterations so a manually bounded budget still governs). The legacy
local-processing and near-limit exits are byte-identical; a new
'repeated_outer_errors' exit reason gets a user-facing explanation.

The inner retry/fallback layer owns transient API recovery and
terminates on its own, so only exceptions that escape it reach this
cap - a successful turn is unaffected.

Fixes #92450
2026-08-24 04:47:07 +05:30
Teknium 26cf456598 refactor: reconcile with #93269 — one shared shutdown predicate, keep both guard sites
PR #93269 (kshitijk4poor) landed the outer-handler break for the same
symptom while this branch was in flight. Keep his guard (it covers
shutdown errors from local post-processing and does the resume-hint +
best-effort persist) and keep this branch's inner-retry-handler return
(it fires BEFORE the ⚠️ retry trace, credential rotation, and fallback
attempts that the outer handler never sees). Point his
_is_interpreter_shutdown_error at tools/interpreter_shutdown.py so the
class has exactly one text-matching site, preserving his RuntimeError
type gate and all 7 of his tests.
2026-08-23 16:04:41 -07:00
Teknium 9ea7fe9938 fix: quitting the CLI no longer spams shutdown-race API errors onto the shell
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.

Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
  (matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
  predicates now delegate to the shared home (tool_executor previously
  matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
  shutdown signal and abandons the turn — one log warning, no print,
  no traceback, no debug dump, no retry; outer handler gets the same
  guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
  (set by the background-review fork) instead of bypassing it.

Refs #55924 #58720 (same class in cron delivery), adjacent to #90683.
2026-08-23 16:04:41 -07:00
kshitijk4poor e63786fc23 fix(agent): break immediately on interpreter shutdown in conversation loop
When the Python interpreter begins teardown (user closes hermes, SIGTERM,
OOM-kill), every executor-backed operation raises 'cannot schedule new
futures after interpreter shutdown'. The outer except handler in
run_conversation caught this error but did not recognize it as fatal —
it kept retrying (API calls #4, #5, #6) until max_iterations, each time
hitting the same dead executor and printing another traceback.

The fix adds an early check: if sys.is_finalizing() or the error matches
the 'cannot schedule new futures' pattern, break immediately with a clean
interpreter_shutdown exit reason instead of retrying. The codebase already
had this pattern in cron/scheduler.py and agent/tool_executor.py — the
conversation loop just wasn't using it.
2026-08-24 04:14:26 +05:30
Teknium 334bcbac93 fix(desktop): error card honors the classifier's retry verdict + failing-session identity (review feedback)
Addresses @helix4u's review on #91493:
- conversation_loop now stamps failure_retryable (the real ClassifiedError
  verdict) next to failure_reason; error_surface prefers it and only falls
  back to the reason set for older results. Fallback set corrected to match
  classify_api_error (auth, format_error, billing_unverified now
  non-retryable).
- The descriptor carries the failing session's provider/model captured at
  classification time; Copy error details prefers them over the foreground
  composer atoms.
- Open logs is labeled 'Open Desktop logs' on remote/cloud connections —
  the local folder holds transport logs, not the remote runtime's.
- API-exception module allowlist widened to botocore/boto3/google/grpc/
  requests/aiohttp so other adapter SDKs don't misclassify as gateway.
2026-08-21 15:24:03 -07:00
qixuancao 37da0d4d50 fix(agent): synchronize background review cancellation 2026-08-21 16:12:57 +05:30
kshitijk4poor cef999c5cf refactor(agent): consolidate uncompressed-overflow guard to one warn site + re-arm
Review follow-up on the salvaged #89444:
- Warn fires only from the conversation-loop pre-API site, reusing the
  unconditionally computed request_pressure_tokens (zero marginal cost,
  covers turn-start AND mid-turn growth) — drops the duplicate every-turn
  estimate the turn-context block paid.
- Turn-context block now only RE-ARMS the dedup once the session is back
  under the window, so warn -> /compress -> regrow warns again (the dedup
  was previously never cleared with compression disabled).
- Char pre-check treats non-string (multimodal) content as over-gate —
  len() of a part list defeated the 20k char floor (probe: 10 'chars' vs
  ~70k real tokens) — and compares against the window, not a flat 20k.
- Deletes the unreachable get_model_context_length fallback from both
  sites (context_compressor always exists; its context_length property
  hard-floors positive; the fallback would have been a synchronous
  network probe mid-turn that also bypassed config overrides) and the
  undeduped inline _emit_warning fallback (third copy of the message).
- Tests bind the PRODUCTION warn/clear methods (previously a verbatim
  fake reimplementation left them uncovered) and add dedup, re-arm,
  no-rearm-while-over, and multimodal-gate coverage.
2026-08-20 12:35:20 +05:30
joaomarcos 4d1fc6ca0a fix(agent): add mid-turn uncompressed context overflow guardrail (#89297) 2026-08-20 12:35:20 +05:30
kshitijk4poor b7e12decc6 fix(agent): route relay-wrapped output-cap 429s into the output-cap handler
Salvage follow-up for #72283: instead of a second pre-retry clamp block
(which bypassed the #55546 clamp+compress path and broke its three
regression tests), parse the output cap ONCE at classification time and:
- exempt parseable wrapped output-cap 429s from the eager rate-limit
  provider fallback (a deterministic request-shape failure that failover
  cannot fix but the clamp fixes in one retry), and
- widen is_context_length_error so they reach the SAME #55546
  clamp+compress recovery as plain output-cap 400s.

Adds both #72283 regression scenarios plus an ordering guard proving a
NON-EMPTY fallback chain does not consume the wrapped 429 (fallback
slot unspent, model unchanged). 119 fallback/rate-limit tests green.
2026-08-20 11:37:01 +05:30
Teknium 449471c334 feat: runtime stall guards — identical-call loop breaker and continue-intent recovery (agent.stall_guards)
Composio eval traces showed Hermes wasting turns re-issuing identical tool
calls (same tool, same args, same result — 3x/4x in one run) and ending
turns by announcing an action it never took. Two conservative, config-gated
guards (agent.stall_guards, default true):

- Identical-call loop breaker: ToolCallGuardrailController.observe_identical_call
  tracks the consecutive streak of (tool, canonical args, result-hash); on
  the 3rd identical call a compact one-line notice is appended to that tool
  RESULT at construction time (cache-safe — tool results are append-only).
  Never blocks the call. Pollers (process, *_get_result, *_poll) are exempt
  via STALL_GUARD_REPEATABLE_TOOLS. Streak resets on any different call,
  changed result, or new turn. Observed on the raw result before the
  tool-loop warning suffix so its changing count can't defeat matching.

- Said-continue-but-stopped recovery: trailing_continue_intent() detects a
  short reply ENDING on an announced next action ('Let me now…', 'I will
  now…', 'Next, I…'); the conversation loop feeds it into the EXISTING
  intent-ack continuation path (same interim-assistant + user-nudge
  mechanism, same codex_ack_continuations cap of 2), preserving message
  alternation — no parallel recovery machinery.

Config: agent.stall_guards in DEFAULT_CONFIG; docs in configuration.md;
unit tests for streak/allowlist/reset/gate and detector pos/neg cases.
2026-08-19 16:34:21 -07:00
Teknium 803397ecc3 feat: wall-clock run budget — wrap-up injection at 80% and deadline-scaled stale timeouts (agent.run_budget_seconds / --run-budget) 2026-08-19 16:32:17 -07:00
Teknium 210cdb0ed3 fix(agent): legacy hidden redirect placeholders get the neutral wire payload at projection time (#88955)
The salvaged writer-side fix stamps api_content on NEW hidden redirect
placeholders, but rows persisted before it (content="" + display_kind=hidden,
no sidecar) would keep re-triggering repair_empty_non_final_messages on every
call forever. Substitute [response interrupted] on the wire copy at the
api_content/display_kind projection stage so legacy sessions converge too.
Never the interrupt scaffold (#81841). Durable transcript untouched.

Regression tests drive run_conversation end-to-end with a spied sanitizer:
the projection must leave the sanitizer nothing to heal (its per-turn warning
spam is the bug), verified failing via sabotage run against the writer-only
fix.

Projection-side approach credit: @JoaoMarcos44 (PR #88996).
2026-08-18 16:27:32 -07:00
Axl Ibiza, MBA 0ee9bc8d1e fix(agent): give interrupted-turn hidden placeholder a neutral provider-replay sidecar
Bot-mode interrupted member turns with no visible assistant text persisted an
empty assistant row (content="" + display_kind="hidden"). The pre-call
sanitizer repair_empty_non_final_messages() re-healed that row on every later
call (wire copy only), so the loop never converged (#88955).

Stamp api_content="[response interrupted]" (the canonical
_INTERRUPTED_PLACEHOLDER) on the hidden placeholder instead. display_kind is
stripped before sanitization, but api_content is projected back into content
for historical assistant rows, so the provider sees a non-empty neutral turn
and the sanitizer stops touching the row — while the durable transcript stays
hidden and empty. Uses the neutral interruption text, never the
_INTERRUPTED_SCAFFOLD_MARKER, which replaying as assistant text caused #81841.

Adds regression coverage proving (A) the placeholder carries the replay
sidecar, (B) two consecutive projections converge without sanitizer healing,
(C) the sanitizer still repairs genuinely-empty unmarked assistants.

Refs #88955
2026-08-18 16:27:32 -07:00
Shannon Sands ac06c2ff8b fix(agent): stop re-billing deterministic empty responses (NS-503)
Every empty-response retry re-sends the full conversation input at full
price. On large contexts a single turn that produces no visible output
could bill the user several dollars across the 3-retry + fallback-chain
walk (reported: ~$2.33 for one empty answer on a ~26K-token session).

Signaled refusals (finish_reason=content_filter, Anthropic refusal
stop_reason, guardrail interventions) are already terminal today and
never reach this loop. The uncovered class is *unsignaled* refusals:
the provider returns 200 with zero output tokens and a generic finish
reason. Those are deterministic — resending the identical prompt
reproduces the same empty — so burning the remaining retry budget only
multiplies the charge.

New agent/empty_response_guard.py, two independent guards, both failing
OPEN to today's behaviour:

- Deterministic-empty detection: two consecutive empty attempts with
  usage present, output_tokens == 0 (reasoning tokens count as output),
  and identical (model, provider, finish_reason) skip the remaining
  retries and go straight to the fallback chain — a different model may
  well answer. Missing usage, nonzero output, or any signature change
  keeps the full budget.
- Cost-aware retry budget: when one attempt's estimated input cost
  exceeds HERMES_EMPTY_RETRY_COST_THRESHOLD_USD (default $0.25), the
  empty-retry budget drops 3 -> 1 for that streak. Unknown pricing or
  included/subscription routes are untouched.

At exhaustion the status trace now includes the estimated cost of the
empty attempts so the charge is at least explained in-session.

Streak state lives on the agent and self-clears whenever
_empty_content_retries resets to 0, transparently honouring every
existing reset site (turn start, tool success, compaction, fallback
activation) without touching them.

Set HERMES_DETERMINISTIC_EMPTY_GUARD=0 to disable both guards.

Tests: tests/agent/test_empty_response_guard.py (26 unit tests) plus
two loop-level integration tests in tests/run_agent/test_run_agent.py
proving the api_call reduction and the fail-open path.

Refs NS-503.
2026-08-16 22:06:08 -07:00
Teknium 4e22d070f4 feat(agent): one-time protocol upgrade for legacy Bot Chat sessions
Bot Chats created before the epoch mechanism persisted prompts with no
protocol section and no stamp — the staleness check only fires on
stamped prompts, so pre-existing bots would never learn to message
teammates. stored_bot_chat_prompt_needs_upgrade() migrates them: one
rebuild, title-gated to Bot Chat, only when the probe would actually
emit a section (SOUL-append legacies and unmanaged installs are left
alone — rebuilding those would loop). The rebuilt prompt carries the
stamp, so the upgrade can never re-fire.

E2E v3b through the real restore path: legacy Bot Chat upgraded once
then verbatim-reused; legacy regular sessions byte-untouched.
tests/agent/ 4648/4648.
2026-08-16 18:30:53 -07:00
Teknium ea4310e76c feat(agent): capability-refresh + timeless prompts for eternal Bot Chat sessions
Bot Chats break the "new sessions come often" assumption behind
build-once system prompts: capability edits used to sit invisible until
/new or compression, and the frozen birth date became misinformation.

- tools/bot_mode_probe.py: capability_fingerprint() hashes the profile's
  capability surface (disabled skills, toolset pins, MCP config, SOUL.md,
  installed skills, Bot-Mode roster); Bot Chat prompts embed the 12-hex
  epoch stamp
- agent/conversation_loop.py restore path: stored Bot Chat prompt whose
  epoch mismatches disk → ONE rebuild (through a cleared skills-prompt
  cache so new installs appear), persisted so the next turn reuses the
  new bytes verbatim. Prompts without a stamp — every non-Bot-Chat
  session — never take the branch; probe failure fails closed to reuse
- agent/system_prompt.py: Bot Chat prompts are timeless — the
  "Conversation started:" date is dropped (timezone kept); no ticking
  fields in an eternal session
- tui_gateway: _sync_bot_capabilities at turn start rebuilds the live
  agent (tool definitions are construction-baked) when the fingerprint
  moves, same session id/history, with a user-visible notice

Cache stance: this is the /model exception applied to capabilities — a
loud, user-initiated, once-per-change prefix break. Unchanged state
hashes identically and stored bytes are reused verbatim (E2E-proven).

Validation: 9 probe unit tests incl. per-axis fingerprint changes;
E2E v3 against the real restore path (fresh build → verbatim reuse →
skill install → single refresh w/ new skill in index → verbatim reuse;
regular sessions dated, unstamped, never refreshed); tests/agent/
4647/4647.
2026-08-16 18:30:53 -07:00
webtecnica b48ab1b4ad fix(agent+discord): guard truncated-response continuation loops and cap Discord split delivery (#86581) 2026-08-16 01:55:51 -07:00
Tuck 79a41c1d32 fix: convert remaining messages.append() calls to append_message() 2026-08-15 01:04:19 -07:00
Teknium 4b7b2b0049 fix: widen base-URL hostname identity class to remaining substring sites
Follow-up to #85737, which migrated five provider-identity sites onto
utils.base_url_host_matches()/base_url_hostname(). This completes the class
sweep (never-patch-predicates: one owner, every site) and folds in the two
open contributor PRs attacking individual sites:

- agent/auxiliary_client.py ZAI/Kimi OpenAI-wire rewrite (PR #85715,
  pierrenode): 'bigmodel'/'api.z.ai'/'api.kimi.com' substring checks
  rewrote proxy paths containing those markers.
- hermes_cli/runtime_provider.py Azure endpoint detection (PR #74721,
  RelaxJonh, issue #74312): 'azure.com' substring picked the Azure key
  for non-Azure hosts whose path contained the text.
- run_agent.py: _is_azure_openai_url, _is_copilot_url, Anthropic
  credential-refresh azure guard, _anthropic_preserve_dots host
  allowlist, OpenRouter/mistral reasoning gates.
- agent/chat_completion_helpers.py: nousresearch / nvidia detection.
- agent/conversation_loop.py: GitHub Models 413 hint.
- agent/usage_pricing.py: localhost billing-route detection.
- hermes_cli/model_switch.py: api.openai.com catalog fallback and
  localhost custom-provider detection.
- cli.py: local-model autodetect and Ollama/LM Studio context-length
  hints (port-anchored instead of '11434' in URL).
- tools/mcp_oauth.py: Figma remote-MCP detection.
- tools/skills_hub.py: raw.githubusercontent.com source-URL check.

Regression tests extend tests/hermes_cli/test_base_url_host_identity.py
(azure/copilot/dotted-model/figma proxy-path + lookalike cases) and
tests/agent/test_minimax_auxiliary_url.py (ZAI/Kimi path false positives).

Closes #74312. Salvages #85715 and #74721 with authorship preserved.
2026-08-14 22:04:16 -07:00
Jack Lau f57209bc9f fix(agent): carry the ambiguity of Anthropic's 'out of extra usage' 400 through classification, cooldown, and terminal surfaces
Review follow-up (egilewski): the previous commit only hedged the guidance
text; the exact Anthropic 400 was still classified, persisted, and surfaced
as confirmed billing exhaustion. Carry the ambiguity all the way through:

- agent/error_classifier.py: 'out of extra usage' matches on the 400 and
  status-less paths now attach error_context {billing_unverified,
  possible_content_filter}. Reason stays FailoverReason.billing (rotation +
  fallback remain the right recovery either way); ClassifiedError grows a
  billing_unverified property.

- agent/credential_pool.py: new FAILURE_REASON_BILLING_UNVERIFIED. An
  unverified billing exhaustion gets the short transient cooldown instead of
  the one-hour bench, regardless of pool size: a content-filter rejection
  leaves the credential healthy and fails identically on every key, and the
  hour-long sole-credential latch is what replayed the stored error and made
  real fixes look ineffective. A true 402 keeps the full bench. The marker
  persists with the entry so a restart cannot upgrade it back to a bench.

- agent/agent_runtime_helpers.py + run_agent.py: recover_with_credential_pool
  threads billing_unverified and hands the pool 'billing_unverified' as the
  persisted failure_reason.

- agent/conversation_loop.py: the fallback-switch status, max-retries status,
  terminal label, and both structured terminal results hedge when the verdict
  is unverified. New _billing_terminal_label + _billing_failure_result build
  the returned terminal response in one place; the result dict now carries
  billing_unverified and the billing_block gains 'unverified': true. The
  confirmed-billing path (a real 402 or an API-key credit depletion) keeps
  the original assertive wording, so the caveat no longer dilutes it.

Regression tests: classifier marking (400 + status-less + unambiguous-body
negative), pool cooldown TTLs + persistence round-trip, pool failure_reason
plumbing, and the returned terminal response for both unverified and
confirmed verdicts.

Note: tests/agent/test_credential_pool_routing.py::TestFailureAttribution::
test_unmatched_key_does_not_retry_only_pool_entry fails identically on
current main without this change (pre-existing, unrelated).
2026-08-14 21:54:56 -07:00
Jack Lau 6fbbe18be8 fix(agent): reword SKILLS_GUIDANCE trigger and stop mislabelling its 400 as billing
On an Anthropic subscription OAuth credential, every request failed with
HTTP 400 "You're out of extra usage. Add more at claude.ai/settings/usage".
That is not a billing condition: Anthropic's server-side content filter rejects
the first sentence of Hermes' own built-in SKILLS_GUIDANCE prompt, and the
rejection is surfaced with a billing-shaped message. Because the message points
at the usage settings page, it reliably sends people to buy quota they do not
need — the reporter lost three debugging sessions to it.

Bisected against the live API with the real 71,721-char assembled prompt: the
first SKILLS_GUIDANCE sentence alone reproduces the 400 and removing it alone
clears it. Size was ruled out (20 KB of unrelated filler returns 200) and so was
the system[0] identity gate (that returns 429, a different failure).

Three changes, all serving the same outcome — a subscription user can no longer
be misdirected by this 400:

- agent/prompt_builder.py: reword the triggering sentence to the phrasing the
  reporter verified returns 200. Meaning, the skill_manage reference, and the
  ## Skill Safety Rule block are all preserved. The reword is empirically
  validated rather than understood, so a comment records the bisect and warns
  that any rewrite must be re-verified against an OAuth token, not an API key.

- agent/conversation_loop.py: the Anthropic branch of the billing guidance no
  longer asserts exhaustion as fact. It hedges the opening line, names the
  content-filter alternative, and gives the operator a way to tell the two apart
  (if the usage page still shows quota, suspect a content rejection). It also
  points at `hermes auth reset anthropic`, because the credential exhaustion
  latch replays the stored error for ~60 min without issuing a request — which
  makes a real fix look like it did not work.

- hermes_cli/auth.py: document that CLAUDE_CODE_OAUTH_TOKEN is an OAuth token,
  not an API key, despite auth_type="api_key". It stays in api_key_env_vars
  because that tuple doubles as the credential-discovery list; removing it would
  stop Hermes finding a `claude setup-token` credential at all.

Docs updated to match the reworded prompt.

Fixes #82154
2026-08-14 21:54:56 -07:00
Teknium 8dc5608a78 fix(compression): adopt live continuation tip at flush across multi-hop chains
A turn writing against a session already closed by compression died with
session_persistence_failed and a misleading "this is often a full disk"
dialog, even though the store was healthy and a live continuation existed
(#82001). Depth-1 recovery (find_live_compression_child) could not resolve
lineages with >=2 compression hops (root -> mid -> tip), reproduced
independently on two- and three-hop chains.

- run_agent.py flush chokepoint: on CompressionSessionClosedError, resolve
  tip = db.get_compression_tip(old_id) (canonical bounded transitive walk),
  adopt only when tip != old_id AND the tip row is live, retry the flush
  exactly once (adoption budget); otherwise fail closed.
- gateway/session.py append_to_transcript: replace the depth-1 live-child
  lookup with the same tip + liveness contract, so gateway transcript
  reroutes follow full chains.
- agent/conversation_compression.py _adopt_live_compression_child: turn-start
  recovery preflight now resolves via get_compression_tip with the same
  liveness check, closing the last depth-1 consumer in this family.
- classify_persistence_error: new "compression_closed" bucket; the turn-end
  explanation names compression rotation and tells the client to refresh the
  session id instead of blaming a full disk.

Tests: depth-1 adoption, multi-hop chain adoption (agent + gateway), fail
closed with no continuation / stale-closed (ws_orphan_reap) tip, exactly-once
adoption budget, and error-wording guards (compression-closed never mentions
disk; real disk failures keep disk guidance).

Closes #82001

Co-authored-by: Al3xand3r1987 <125030427+Al3xand3r1987@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
2026-08-14 21:39:44 -07:00
Drexuxux ab879a1e22 fix(agent): stop the MoA prepared request reaching a swapped-in native client
`_moa_prepared_request` is a private handshake between the conversation
loop and MoAChatCompletions.create. It is attached whenever
agent.provider == "moa", on the assumption that agent.client is still the
in-process MoA facade.

Credential rotation, provider fallback and dead-connection cleanup all
rebuild agent.client from _client_kwargs between attempts, and
pending_moa_prepared_request deliberately carries a prepared request
across exactly that boundary. The rebuilt client is a native OpenAI
client while provider stays "moa", so the key reaches an SDK that has
never heard of it:

    TypeError: Completions.create() got an unexpected keyword argument
    '_moa_prepared_request'

That error is non-retryable, so every remaining turn on the session
fails. Both dispatch paths are affected: the non-streaming one calls
agent.client directly, and _create_request_openai_client returns
agent.client unchanged for provider "moa".

Re-check the live client at the point the key is attached, which covers
both paths at once. When the facade is gone, send the prepared prompt
without the handshake and log the downgrade.
2026-08-14 21:36:41 -07:00
Teknium fbec3c78fd fix(compression): clear preflight block on provider-confirmed rearm
A provider-confirmed rearm (#85846) resets the shared attempt budget, but
an earlier insufficient-progress verdict left _preflight_compression_blocked
armed, keeping the pre-API gate dark for the rest of the turn — a later
pressure spike could still grow unchecked until the provider overflow
handler fired. Clear the blocker and the stale pressure reading inside the
provider-confirmed rearm branch: the prompt is proven back below the
threshold, so the old verdict describes a request shape that no longer
exists.

Builds on @h-mascot's #84995, whose commit is preserved on this branch;
his rearm condition was superseded by #85846's latch-verified variant, but
the blocker-clear half was correct and is kept.
2026-08-13 23:43:54 -07:00
Henry Mascot 2b9da1f252 fix(compression): re-arm same-turn attempt budget 2026-08-13 23:43:54 -07:00
Gille 380e4da3d1 fix(compression): rearm budget from verified usage 2026-08-13 23:08:19 -07:00
ai-ag2026 72c828ca2c fix(agent): refund per-turn compression budget after real progress
A marathon tool turn burned all compression_attempts on *successful*
pre-API compactions; the gate then went permanently dark and the context
grew unchecked until the provider rejected the request terminally
("Context length exceeded: max compression attempts (3) reached", session
f087963205f9, 2026-08-01). The budget now refunds at loop top when the
assembled request is back under threshold * 0.8 AND the compressor's own
should_compress() agrees the pressure is gone.

Anti-thrash intent of the cap is preserved (#11529): no-progress passes
never reach the refund margin, divergent-signal cases (should_compress
still True) keep the budget burnt, and the insufficient-progress blocker
is untouched.

Validation: new behavioral suite (7) + all 130 compression/context tests
via scripts/run_tests_hermetic.py.

Rückbau: Commit revertieren; kein Zustand, keine Migration.

(cherry picked from commit 041b489d566bfb1d6816d53d9acd5ac50b8d5af0)
2026-08-13 23:08:19 -07:00
kshitij 1df33dccfe fix: restore stale-base revert hunks in conversation_loop.py and moa_loop.py
The diff-apply salvage introduced stale-base revert hunks — the PR was 1246
commits behind main, and its diff for conversation_loop.py and moa_loop.py
silently dropped symbols added after the PR's base (e.g.
_CODEX_ACK_CONTINUATION_NUDGE, _INTERRUPT_SCAFFOLD_MARKER, cache_ttl plumbing,
finalize_turn import, _restore_user_after_reference_handoff).

Restored both files to origin/main and re-applied only the PR's additive
changes: _moa_reference_metrics_for_hook, _system_prompt_for_hooks, the
system_prompt= and moa_references= hook kwargs, _last_reference_metrics
attribute and accessors, and the slot_metrics population in the fan-out path.

Fixes CI ImportError: cannot import name '_CODEX_ACK_CONTINUATION_NUDGE' from
'agent.conversation_loop'.
2026-08-13 23:10:16 +05:30
kshitij e665300d6b feat(langfuse): widen tracing to errors, sessions, subagents, and MoA fan-out
Salvaged from PR #83437 by @erosika, with adopted fixes from @bgodlin (#81054),
@aldoeliacim (#82332), @nftpoetrist (#42326), @rodboev (#39653), @FnExpress
(#64292, supersedes #32175 by @db-aeon), @Per0-1 (#61166), @NaMinhyeok (#64797),
and @liuhao1024 (#43130).

Widens the bundled Langfuse plugin from 6 to 11 hooks and fixes two
attribution bugs. Also adopts shutdown/atexit lifecycle fixes and composes
8 prior community PRs with interaction-fix follow-ups.

Model attribution: on_pre_llm_request and on_post_llm_call now prefer the
wire value (request body model, response model) over the agent attribute,
which goes stale after /model switch or provider fallback.

Cost total: both cost paths now send a summed total alongside the per-type
breakdown, since Langfuse does not derive calculatedTotalCost from
cost_details keys. Subscription-included routes send no cost keys at all.

New coverage: api_request_error closes failed generations with ERROR level;
on_session_finalize/on_session_end close dangling traces for tool-only and
interrupted turns; subagent_start/subagent_stop trace delegated children as
spans; MoA advisor fan-out emits one generation per advisor priced at the
advisor's own model.

Capture modes: HERMES_LANGFUSE_CAPTURE=metadata|sanitized|full (default
sanitized). Sanitized mode redacts secret patterns before truncation.

Adopted lifecycle fixes: shutdown client at session finalize when
reason=shutdown (not on session rotation); atexit finalizer ends open root
spans for short-lived processes; root context manager exited to prevent
interpreter-teardown TypeError; TOCTOU on _get_langfuse() fixed with lock;
reasoning_content surfaced in traces; system prompt included in generation
input for Anthropic/Codex/Bedrock; SDK v3 update_trace replaces set_trace_io.

Closes #29482, #43129, #72661.
Supersedes #81054, #82332, #42326, #39653, #64292, #32175, #61166, #64797, #43130.
Partially addresses #67544 (capture modes + secret redaction; user_id remains open).
2026-08-13 23:10:16 +05:30
Teknium e029e300ca fix(compression): harden native compaction rejection matcher + config coercion (#82777)
Two reliability gaps from #82777:

1. Rejection matcher required only a field-name mention, so a transient
   5xx/timeout whose body echoed the request (which contains
   context_management) permanently downgraded native compaction for the
   session. Now requires rejection language (unknown/unsupported/invalid/...)
   alongside the field name, and when a parsed HTTP status is available,
   400 specifically — non-400 statuses never match. Message-only transports
   (no status attribute) keep working unchanged.

2. compression.codex_responses_native was coerced with bool(), so the
   strings "false"/"off" enabled the feature. Now uses the shared
   utils.is_truthy_value helper.

Conversation-loop call site passes api_error.status_code through.
Sabotage-verified: reverting the matcher to field-name-only fails the new
echo and non-400 tests.
2026-08-13 03:04:31 -07:00
kshitij a723351a92 refactor: hoist preflight clear into restart handler, single-source qwen predicate
Simplify-pass follow-ups on the salvage stack (all guard tests re-run,
mutation-checked):

1. conversation_loop.py: moved `_preflight_compression_blocked = False`
   from 9 per-site copies into the restart_with_rebuilt_messages handler
   (its single consumer). Besides removing the 9 duplicated blocks, this
   fixes a 10th pre-existing retry-loop site (content-filter stall
   failover, #32421) that set the flag and broke WITHOUT clearing the
   preflight block — a content-filter failover previously restarted with
   preflight compression still blocked against the fallback's smaller
   window, the same #84733 bug class. The outer-loop empty-response site
   keeps its own clear (it never passes through the handler). New AST
   guard test_restart_handler_clears_preflight_block pins the hoisted
   clear (mutation-checked).

2. agent_runtime_helpers.py: extracted _raw_cache_ttl_from_config() —
   prompt_caching_disabled_from_config and configured_cache_ttl were
   verbatim copies of the same config read. Added VALID_CACHE_TTLS.

3. prompt_caching.py: added is_qwen_model() next to
   ALIBABA_FAMILY_PROVIDERS; effective_cache_ttl and
   anthropic_prompt_cache_policy now share both the family set and the
   qwen predicate — neither can desync.

4. Guard-test hardening: assert every _try_activate_fallback reference
   is a direct `if agent._try_activate_fallback(...):` site, so a future
   `activated = ...` form can't silently escape the restart-discipline
   guard.
2026-08-13 15:24:47 +05:30
kshitij d7517d6e73 fix: restore empty-response fallback retry, dedupe alibaba set, thread TTL into aux replan
Follow-ups on the salvaged #84782 (webtecnica):

1. conversation_loop.py: the empty-response fallback site sits directly
   in the OUTER iteration loop, not the retry loop. The salvaged commit's
   `break` there exited the conversation loop and ended the turn without
   ever calling the just-activated fallback (caught by CI:
   test_empty_response_triggers_fallback_provider). Restored `continue`
   (which already re-runs the pre-API preflight at the top of the next
   outer iteration) while keeping the `_preflight_compression_blocked`
   reset. The other 9 sites are inside the retry loop, where `break` to
   the restart_with_rebuilt_messages handler is correct.

2. test_prompt_cache_ttl_propagation.py: made the AST guard loop-aware —
   retry-loop sites must break, outer-loop sites must continue (the old
   assertion pinned the bug in (1)). Mutation-checked both directions.

3. test_failover_identity.py: added `model` to the SimpleNamespace agent
   fixture — _redecorate_prompt_cache_for_provider now reads agent.model
   for the per-destination TTL clamp (2 CI failures).

4. prompt_caching.py / agent_runtime_helpers.py: single source of truth
   for the alibaba-family provider set — ALIBABA_FAMILY_PROVIDERS lives
   in prompt_caching and anthropic_prompt_cache_policy imports it, so the
   cache-policy opt-in and the TTL clamp can never desync.

5. auxiliary_client.py: threaded the configured tier into
   _replan_synchronous_cache_sections via new configured_cache_ttl()
   (no live agent on that path) — the aux half of #84733's report also
   stopped regressing 1h to 5m. Guarded by
   TestAuxFallbackReplanThreadsTtl (mutation-checked).

6. Dropped the redundant `or "5m"` at the two threaded call sites —
   effective_cache_ttl already resolves None to "5m", and the `or`
   masked the cache-disabled (None) semantics.
2026-08-13 15:24:47 +05:30
webtecnica 9a5cf83541 fix(agent): propagate prompt-cache TTL to MoA/aux, clamp Qwen 1h, re-preflight on failover (#84733) 2026-08-13 15:24:47 +05:30
Teknium 4ea2a0e546 Revert "Inspired by Perplexity Computer: Model Council mode for Mixture of Agents"
This reverts commit 8d9e18d40b.
2026-08-12 21:50:35 -07:00
Hermes Agent 8d9e18d40b Inspired by Perplexity Computer: Model Council mode for Mixture of Agents
Adds a 'council' synthesis style to MoA (per preset via synthesis_style,
one-shot via the new /council command on CLI + gateway). Reference models
answer independently; the aggregator chairs the deliberation and produces
a user-facing report of consensus, per-model disagreements (with the
differing assumptions behind them), unique contributions, and a
recommendation with an explicit confidence level.

Inspired by Perplexity's Model Council rollout to Perplexity Computer
(changelog 08/04/26): pick a board of 2-8 models, run them independently,
synthesize where they agree/disagree and what each uniquely surfaces.
2026-08-12 19:44:09 -07:00
SeoYeonKim 9acf0db889 feat(plugins): add cache-safe system prompt sections
Salvage the plugin-owned static prompt idea from PR #51589 into the constrained #64167 contract: stable IDs, deterministic placement, bounded fail-open rendering, and full-prompt resume recovery without new session columns.

Co-authored-by: Topher Ross <biz@topherross.com>
2026-08-12 16:34:58 -07:00
kshitij 2afa4be932 fix: trim comments and fix sibling pop site in summary path
Trim verbose comments in conversation_loop.py and run_agent.py to 2 lines
each. Fix the same bug class in the compression summary path at
chat_completion_helpers.py: remove _thinking_prefill from the explicit
pop tuple and move the generic underscore-key sweep to after
_drop_thinking_only_and_merge_users, so the drop pass can recognize
prefill stubs there too.
2026-08-10 10:01:44 +05:30
Jefftree 97ced4bce2 fix(agent): keep the thinking-prefill marker so the drop pass can strip trailing stubs 2026-08-10 10:01:44 +05:30
kshitij e4b2a90dad refactor: follow-up for salvaged PR #81692
- warn (not debug) on final text-turn flush failure: a failure here
  reopens the exact #81641 data-loss window with _persist_session as
  the only remaining retry, unlike the verify siblings which retry
  in-loop; include session id for triage
- trim the flush-site comment to sibling proportion, pointing to the
  test module for the full incident narrative
- test: assert _persist_session presence before indexing, so a wiring
  change fails with a clean assertion instead of ValueError from max()
2026-08-10 09:38:28 +05:30
joaomarcos 6c2c77efba fix(agent): persist completed text turns before the loop exits (#81641)
A pure-text assistant turn (finish_reason=stop) had no durable write of
its own. Its answer reached the user through the streaming / interim
display path, which is display-only and never touches state.db, and the
first durable write was finalize_turn's _persist_session — after the
loop exits and behind post-turn work that can include micro-compaction's
aux-LLM call.

Anything that ended the process or tore the session down inside that
window lost a reply the user had already been shown. On a remote
(non-loopback) backend the window is easy to hit: WS 1006 closures drive
ws_orphan_reap teardown, and affected sessions ended up with user rows
and zero assistant rows in state.db.

The neighbouring exits of the same loop already close this gap:

  * the tool-call exit flushes the assistant(tool_calls) block before
    handing control to _execute_tool_calls (#49045)
  * the verify-on-stop and pre_verify exits flush final_msg before
    appending their nudge (#65919 §7)

Apply that same idiom to the ordinary text exit rather than adding a new
persistence mechanism. The intrinsic _DB_PERSISTED_MARKER dedup makes the
later _persist_session a no-op for this row, so no duplicate rows and no
extra write — the same write, just earlier.

Unlike the tool-call exit, a failed flush must not abort the turn: no
side effect runs after this point and the answer is already produced, so
the failure is logged and _persist_session remains the retry.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 09:38:28 +05:30
kshitij 6bdeb2df24 fix: move ghost filter before alternation repair + promote scaffold constant
Move the legacy ghost-row filter from inside the api_messages loop to
BEFORE repair_message_sequence_with_cursor. Dropping a ghost assistant
row between two user messages creates user→user which the repair can
now fix (previously the repair ran first and missed it).

Promote '[This response was interrupted by a user correction.]' to
module-level _INTERRUPT_SCAFFOLD_MARKER constant — used in both
_apply_active_turn_redirect (checkpoint_parts) and the ghost filter,
so they can never drift.

Update ghost-row test: the two consecutive user messages are now
merged by repair, so check for content as substring.
2026-08-09 21:30:54 +05:30
HexLab98 0072969c18 fix(agent): drop legacy interrupt-scaffold ghost rows from API replay
Sessions already poisoned by the incomplete #73146 else branch still replay
hidden assistant rows whose content is the raw interrupt scaffold. Skip those
rows when building provider messages so old state.db history cannot keep
seeding the echo loop.
2026-08-09 21:30:54 +05:30
HexLab98 68d1aea1ae fix(agent): keep interrupt scaffold off the tool-tail redirect placeholder
The incomplete #73146 else branch still wrote the interrupt checkpoint into
the placeholder assistant row. Mid-tool steers then replayed that scaffold as
the model's own prior reply, which it echoed into a self-replicating ghost
loop. Carry the scaffold only on the user correction's api_content, matching
the assistant-tail branch.
2026-08-09 21:30:54 +05:30