Commit Graph

3725 Commits

Author SHA1 Message Date
Aldo 6ff9426d22 fix(anthropic): track the current fast-mode model matrix (Opus 4.8 / Opus 5)
The speed=fast allowlist still gates on Opus 4.6, but the fast-mode
matrix has changed twice since it was written (verified against the
live docs, platform.claude.com/docs/en/build-with-claude/fast-mode):

- Opus 4.8 and Opus 5 SUPPORT fast mode (research preview, Claude API
  only — not Bedrock/Vertex/Foundry).
- Opus 4.6 LOST fast mode on 2026-06-29. The parameter does not error:
  requests silently run at standard speed and bill standard rates
  (usage.speed: 'standard'). Today's allowlist therefore shows 4.6
  users a fast toggle that does nothing, while denying it to the two
  models that actually support it.
- Opus 4.7 never had it and hard-400s (unchanged).
- Dedicated '…-fast' model ids (OpenRouter's claude-opus-4.8-fast)
  select fast inference via the model field and are explicitly
  excluded from the param gate.

Both gates move in lock-step as before: the adapter param gate
(agent.anthropic_adapter._supports_fast_mode) and the CLI toggle gate
(hermes_cli.models._is_anthropic_fast_model). Docstrings now record
the history in both directions so the next matrix change has context.

## How to test

scripts/run_tests.sh tests/agent/test_anthropic_adapter.py tests/cli/test_fast_command.py -- -q

113 tests pass. The updated predicate/matrix tests fail against the
previous allowlist (verified by stashing the source changes). Tested
on Linux (aarch64).
2026-09-02 05:33:13 -07:00
Teknium 238b6c1ab9 fix(compression): persist the anti-thrash recovery deadline so gateway agent rebuilds cannot block a session forever
The #14694 recovery clock (`_anti_thrash_recovery_deadline`) was a
process-local `time.monotonic()` value zeroed in `bind_session_state()`.
The gateway rebuilds the AIAgent (and its ContextCompressor) on every
cache eviction, so each fresh compressor bound to a durably tripped
session row (#69872) re-armed a full 300s window and the half-open probe
never fired — a long messaging conversation above the threshold stayed
blocked permanently.

Persist the deadline as a wall-clock epoch in a new
`sessions.compression_recovery_deadline REAL` column (declarative column
reconciliation; SCHEMA_VERSION 26 -> 27) with
`SessionDB.get/set_compression_recovery_deadline`. The compressor loads it
in `bind_session_state()` and writes it on change only via
`_set_anti_thrash_recovery_deadline()`. A fresh compressor with no stored
deadline still starts a full window blocked (#54923 restart contract); one
that loads an armed deadline resumes that window. Backward clock jumps are
bounded to one window. The 300s window is unchanged.

Minimal salvage of #100185 (the probe-lease/fencing state machine and
model_config-blob storage were not carried).

Refs #100185
Co-authored-by: Komzpa <me@komzpa.net>
2026-09-02 04:14:10 -07:00
Teknium bc71b8bc95 fix(compression): anchor on the LAST intent row — newer user turn outranks older steer (#100053 follow-up)
Follow-up to the salvaged #100114 commit. Its two-pass anchor selection
scanned steers first and real user rows second, so a transcript shaped
[user A, tool(steer B), ..., user C] anchored the already-consumed steer B
over the newer real request C — the same replay class the PR set out to
fix. Replace it with one reversed positional scan that picks whichever
intent-bearing row is last (real role=user or steer-bearing role=tool),
and make the compressed-transcript steer check count only role=tool rows
(the only place the runtime delivers a steer), so a summary quoting the
marker cannot masquerade as live intent.

Adds S1/S2/S3 regression tests (steer dropped by compaction, steer
surviving in tail, newer user turn after steer) plus alternation and
use-exactly-once assertions.
2026-09-02 04:12:12 -07:00
finn763 f40333a80e fix(agent): preserve busy steer during compression and avoid replaying historical user request
Compression with display.busy_input_mode: steer embeds the follow-up
as an out-of-band marker inside the latest role=tool result. The
post-compression user-turn preservation path only classified
non-scaffolding role=user rows as real intent, so a compressed
transcript that contained no role=user row would discard the steer
and clone an older historical role=user message as the new active
turn, re-activating a previously consumed request.

Fix _ensure_compressed_has_user_turn to (1) treat a compressed
transcript that already carries a steer marker as having user intent,
and (2) prioritize the latest steer payload from the original
transcript over historical user cloning, inserting it as a proper
role=user turn via _insert_real_user_anchor. This preserves the
actual current intent exactly once and never turns history into new
input.

Closes #100053
2026-09-02 04:12:12 -07:00
Teknium 32fe129324 perf(bot-mode): cold DM hops skip the live /models probe; relay replies land within 250ms
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.

- model_metadata: memoize successful remote /models probes on disk
  (cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
  in-memory cache, so authority semantics are unchanged (reconciliation
  still lands within 5 minutes) but the answer is shared across
  processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
  250ms instead of every 2s — up to 2s of dead air on every relayed reply.

Nothing here changes turn ordering: DMs and group rounds stay serial.

Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
2026-09-02 03:42:01 -07:00
kshitijk4poor 2adb1a4ea6 refactor(agent): clone the review snapshot once at the spawn chokepoint
Move the structural clone from the four call sites (auto review, codex
runtime, CLI /refine, gateway /refine) into AIAgent._spawn_background_review,
which every review path — immediate, idle-queue deferred, requeued — passes
through. Callers can no longer forget it, and the private helper is no longer
imported across hermes_cli/ and gateway/ package boundaries.

Tests now bind the real chokepoint (capturing at _spawn_background_review_now)
so they still fail if the clone is removed.
2026-09-02 13:28:41 +05:30
notkisk 619ca3011e fix(agent): isolate background review snapshots 2026-09-02 13:28:41 +05:30
Teknium 37f3ba110a fix(auth): auto-heal single-use OAuth grants already forked across profiles (#100339)
The clone-strip and root-write-through in the previous commit stop NEW forks
but leave installs that forked before upgrading in the broken state: each
profile keeps its own copy of the root grant, whichever profile rotated last
holds the only live refresh token, and root plus every sibling still hit
invalid_grant on their next refresh. The PR body asked those users to
re-auth at root and hand-edit profiles/*/auth.json; this makes it automatic.

`heal_forked_single_use_oauth_grants(provider)` (hermes_cli/auth.py) runs at
the top of a profile's `load_pool()` for SINGLE_USE_REFRESH_POOL_PROVIDERS.
Under the profile lock then the root lock it matches each profile OAuth row
to its root counterpart by lineage — same pool id (preserved by both fork
paths), same JWT account identity, same token material, else same provider +
same client (Anthropic pkce grants carry no claims) — keeps the copy with the
freshest rotation (`expires_at_ms` / `last_refresh` / JWT exp), writes it into
ROOT when root's is older, and strips the profile copy (pool rows, the
`providers.<id>` device-code block for Codex/xAI, and a profile-local
`.anthropic_oauth.json`) so the profile borrows root from then on. Root's
singleton and its hermes_pkce row are kept in step so root's own re-seed
cannot resurrect the spent pair.

Guarantees: idempotent (mtime-keyed clean mark skips the locked scan on the
per-call hot path); one INFO line per healed profile; API-key rows untouched;
a row with no root counterpart (root lost its grant, or an independent
account whose claims differ) is never deleted; only the two auth.json files
the root fallback already reads are touched — no environ/secret-scope reads.
`hermes auth list` / `hermes auth status <provider>` print the heal note.

Live repro (real imports, temp root + forge/atlas each holding a pre-fix
verbatim copy, forge already rotated RT0->RT1 into its own file, fake
single-use token endpoint): before — atlas None, forge AT2 (only in forge),
root None; server log 4x REUSE of spent RT0. After — forge's load heals to
root and rotates there, atlas and root select AT2, profiles/*/auth.json hold
no anthropic rows, server log exactly one ROTATE and zero REUSE.
2026-09-02 00:58:29 -07:00
Teknium 3038493ee6 fix(auth): never fork single-use OAuth grants across profiles (#100339)
Anthropic / Codex / xAI OAuth refresh tokens are single-use: a grant copied
into a second auth.json is one credential with two owners, and the first
profile to refresh it revokes the pair for every sibling (invalid_grant /
refresh_token_reused). Two code paths forked grants that way:

1. `hermes profile create --clone-all` and the dashboard/TUI
   `mirror_credentials` flow copied auth.json (+ .anthropic_oauth.json)
   verbatim. Both now run `strip_cloned_single_use_oauth_grants()`, which
   drops OAuth rows for SINGLE_USE_REFRESH_POOL_PROVIDERS, the matching
   `providers.<id>` device-code blocks, and the PKCE singleton file; API
   keys are still copied. The clone reads the root grant through the
   existing credential-pool root fallback.

2. A named profile with no local rows BORROWS the root grant via
   `read_credential_pool()`'s fallback, but every persist
   (`CredentialPool._persist`, `load_pool` reseed, `remove_index`) wrote the
   rows into the profile's own auth.json — materializing a fork on the first
   rotation. `persist_pool_entries()` now routes borrowed single-use rows
   back to the root store (update-only, under the root lock; never falls
   back to a local copy). A borrowed `hermes_pkce` rotation commits its
   singleton to the root `.anthropic_oauth.json`, the borrower never prunes
   root-seeded rows it cannot see the backing file for, and
   `hermes -p <profile> auth add` persists only the profile's own rows.

Live repro (real imports, temp root + profiles, fake single-use token
endpoint): before — first profile rotation RT0->RT1 in profile only; root
and sibling then hit `invalid_grant`, `resolve_anthropic_token()` -> None.
After — rotation lands in root; root and both siblings select AT1, no reuse.

Direction per Teknium: stop cloning OAuth into profiles (ONE grant at root,
children inherit via context) rather than making clones survive. Supersedes
the clone-strip/root-write-through half of #100389 and the init-refresh idea
in #100703 (an expired-but-refreshable row already refreshes on select()).

Closes #100339
Co-authored-by: HexLab98 <liruixinch@outlook.com>
2026-09-02 00:58:29 -07:00
teknium1 c83ea9bed7 test(agent): pin the reasoning-off continuation to exactly one request; document its prompt-cache cost
The one-shot reasoning-off retry changes a request parameter that is part
of the provider cache key on config-sensitive providers (Anthropic renders
thinking/effort into the prompt; OpenAI lists reasoning.effort as
prefix-affecting), so that request is a deliberate single cache miss.
Pin the bound: the request AFTER it must carry the configured reasoning
again and the system prompt must be byte-identical across the whole retry
sequence. Sabotage-verified (sticky flag -> test fails on request 3).
Docstring on _consume_ephemeral_reasoning_off states the cost honestly.
2026-09-02 00:55:42 -07:00
Teknium 2a0605a807 fix(agent): reasoning-off continuation reaches the wire on the legacy chat path; reset one-shot flag per turn
Follow-up to the #99622 salvage:
- agent/transports/chat_completions.py: the legacy (no provider profile)
  chat_completions path always re-emitted extra_body.reasoning with
  enabled=True, so both reasoning_effort: none and the one-shot
  length-continuation override went out as {enabled: true, effort: none}.
  Honor enabled=False / effort=none the way the profile path does.
- agent/conversation_loop.py: reset agent._ephemeral_reasoning_off at
  turn start so a flag armed by an interrupted/errored turn can never
  strip thinking from the next turn's first request.
- User-facing hints now name the real slash command (/reasoning); the
  /thinkon//thinkoff commands do not exist.
- tests: wire-level regression (continuation request carries
  reasoning.enabled=false) and a stale-flag turn-scope test.
2026-09-02 00:55:42 -07:00
AlexGabbia fb76fb0526 fix(agent): thinking-only length truncations no longer wedge continuations
GLM-5.3-flash on ollama-cloud with reasoning_effort=high can spend the ENTIRE
output cap on reasoning delivered in a separate field and return
finish_reason=length with no visible content (verified live: max_tokens=4096,
completion_tokens=4096, content empty).

The length-continuation path handled that shape badly:
  1. the empty response was appended as an interim assistant fragment,
     poisoning the transcript until the pre-call sanitizer healed it
     (observed 3+ healings per turn on the reporting user's session);
  2. every continuation re-ran with thinking ON, re-deriving the whole
     thinking budget against a growing context, so 4 attempts still produced
     nothing and the turn died with 'Response remains truncated after 4
     continuation attempts'.

Now:
  - interim assistant fragments with no visible content are never appended
    (whichever way they got empty);
  - a thinking-only truncation sets a one-shot reasoning-off override that
    build_api_kwargs consumes for the next request, so the continuation
    writes the answer instead of re-thinking it;
  - the ceiling exit clears a pending override and, when every fragment was
    empty, returns an actionable final_response instead of an invisible None.
2026-09-02 00:55:42 -07:00
Teknium 25d954c2cf fix(guardrails): hard stops catch replays, never legitimate iteration
Before turning hard stops on for unattended platforms, make sure they cannot
cut off normal work:

- Edit -> re-run is progress. A successful mutating call (write_file/patch,
  a green terminal/execute_code, browser actions, job/message/cron/memory/
  skill mutations) marks progress for every failing signature still being
  counted this turn; the next identical retry restarts its streak instead
  of accumulating toward exact_failure_block_after. A pure replay never
  mutates anything between attempts, so it is still blocked at 5.
- Distinct red commands are diagnosis. For FAILURE_TOLERANT_TOOL_NAMES
  (terminal, execute_code, process pollers, browser_navigate, web_extract)
  same_tool_failure_halt_after warns but never halts.
- subagent and api_server keep the warn-only default: both are supervised
  task loops with a live parent/client and do real edit -> re-run work.

Live A/B (real AIAgent platform=telegram, real patch+terminal, 8 rounds of
patch -> red check -> patch ...):
  unmitigated branch: HALTED at round 6 (repeated_exact_failure_block)
  this commit:        COMPLETED all 8 rounds, final answer delivered
Loop shapes still stopped: identical failing read_file 8 calls,
identical successful terminal 5 calls (vs 602 on main).
Six new tests pin these flows; all fail on the unmitigated version.
2026-09-02 00:26:57 -07:00
Teknium 76648a7faf fix(guardrails): identical-call streaks hard-stop any tool on unattended platforms
Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.

- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
  streak raises a halt (identical_call_streak_halt) at
  hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
  stay exempt; a changed result resets the streak; warning-only sessions are
  unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
  every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
  for pollers, or when results change.

Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
  identical failing read_file   main: 602 API calls, budget exhausted
                                branch: 8 calls, repeated_exact_failure_block
  identical successful terminal main: 602 API calls, budget exhausted
                                branch: 5 calls, identical_call_streak_halt
2026-09-02 00:26:57 -07:00
benbenwyb cd2d3089fb fix(agent): guard repeated skill reads
Treat skill_view and skills_list as idempotent read-only tools so the existing no-progress guardrail can warn or block repeated identical skill loads. This prevents large skill outputs from being re-added to the context in tool loops.

Add regression coverage for repeated skill_view results under hard-stop guardrails.
2026-09-02 00:26:57 -07:00
João Vitor Cunha 384fc4bf83 fix(guardrails): preserve interactive platform defaults 2026-09-02 00:26:57 -07:00
João Vitor Cunha ee2147f9e6 fix: hard stop tool loops on non-interactive platforms 2026-09-02 00:26:57 -07:00
Teknium 9de9d7613c fix(compression): keep hygiene turn-hold worker's commit admission so thinking-model summaries are adopted, not burned
The 10s hygiene_max_turn_hold_seconds budget (#92318) releases the arriving
user turn while the summary model is still streaming. For thinking summary
models (DeepSeek-V4-Flash etc.) whose reasoning prefix alone exceeds 10s,
the abandonment path ALWAYS cancelled the commit fence — 100% of the summary
attempt (including the full thinking prefix) was discarded on every turn,
permanently disabling auto-compression while paying the summary model 10s
of thinking per turn, and the flat 60s retry-after then blocked the
agent-side preflight from a fresh chance.

Structural fix (maintainer-chosen direction in #97963): decouple the turn
from the compression instead of holding the turn longer or making the hold
progress-aware (which would reintroduce the #90845 frozen-turn bug):

- CompressionCommitFence gains mark_commit_watermark_fenced() /
  commit_watermark_fenced; compress_context marks the fence right after
  capturing get_active_message_watermark() under the durable compression
  lock (#75316/#87484) — the property that makes a LATE commit safe: rows
  appended after compression start survive both commit paths verbatim as
  cloned concurrent tail (archive_and_compact watermark= and
  publish_compression_child watermark/watermark_ceiling).
- gateway hygiene turn-hold handler: when the fence is watermark-fenced,
  the detached worker (already kept alive via
  _defer_agent_cleanup_until_future_done) KEEPS its commit admission; the
  user's turn proceeds on the uncompressed transcript at the same 10s
  budget, and the summary is adopted at the worker's own watermark-fenced
  commit boundary. Unfenced workers are cancelled exactly as before —
  never worse than the status quo.
- No retry-after is armed while the kept-admission attempt runs (it would
  block preflight adoption via the same-session cooldown); re-attempt
  spacing is covered by the durable compression lock
  (_session_has_compression_in_flight). If the worker ends WITHOUT
  committing, a done-callback restores the flat non-escalating 60s
  retry-after; a successful adoption resets the hygiene failure streak.
  The streak never advances for a deferral either way.
- Docs: configuration.md hygiene_max_turn_hold_seconds one-liner updated
  to describe deferred adoption and the thinking-model case;
  config_defaults.py comment updated. Knob stays config.yaml-only.

Invariants preserved:
- 10s user-latency cap stays hard (#90845/#92318):
  test_session_hygiene_turn_hold_budget_abandons_streaming_wait passes
  UNMODIFIED (its worker is not watermark-fenced, so it pins the cancel
  path through the public surface).
- Stale-clobber impossible: adoption only rides commits bounded by the
  start watermark; the fence still gates admission and unfenced/late
  results are discarded.

New regression tests (tests/gateway/test_session_hygiene_turnhold_adoption.py):
- watermark-fenced worker keeps admission, late summary is committed,
  turn still released at the budget, no cooldown while running,
  streak reset on adoption;
- kept-admission worker that ends without committing restores the flat
  turn-hold retry-after (<=120s, names turn-hold, streak untouched);
- unfenced worker still cancelled and discarded (status quo).
Sabotage-verified: disabling the keep-admission branch fails the two new
adoption tests and leaves the unfenced-cancel test green.

Fixes #97963
2026-09-01 23:56:23 -07:00
Teknium 30c9d40974 fix(compression): stop the Codex and Anthropic aux summary streams at the host deadline too (#99692)
PR #99779 gave the streamed chat.completions consumer the host's absolute
compression deadline. The two wires that consume their streams internally
still ran on their own, always-larger budgets after the host gave up:

- Codex Responses: clamp the re-armable watchdog's hard ceiling to the
  published host deadline, so a live (re-arming) stream is severed the
  instant the host stops waiting instead of at max(600s, 4x timeout).
- Anthropic Messages: the per-event hook now raises at the host deadline
  and on an explicit hard cancel; create_anthropic_message lets that
  TimeoutError abandon the stream (the with-block closes it) instead of
  swallowing it as a callback failure.

Sabotage-verified: without the Codex clamp the new deadline test hangs past
its 25s harness cutoff; without the Anthropic hook the three Anthropic tests
fail.
2026-09-01 23:56:06 -07:00
joaomarcos 904e5bb572 fix(compression): stop the summary stream at the host's own deadline
CompressionCommitFence.set_total_ceiling_seconds documents its deadline as
"shared by the host and worker", but only the host ever read it. The worker's
streamed summary bounds itself with _aux_stream_total_ceiling() instead —
max(600, 4 * aux_timeout) — which is >= the host's total ceiling for every
configured timeout AND starts counting later (after pool admission,
_serialize_for_summary, prompt build and TTFT). A stream that outlives its
abandoned host is therefore not an edge case; it is the guaranteed outcome of
every total-ceiling timeout.

8207862212 closed the first half: a cancelled fence now releases the
compression owner, freeing its pool slot and session lease. Its own comment
leaves the second half open — the isolated provider daemon that holds the
socket keeps streaming "until the auxiliary stream's longer absolute ceiling
expires". With the #99692 reporter's auxiliary.compression.timeout: 600 that
is 2400s of an orphaned ~500K-token summary the fence is already guaranteed to
refuse, and because the session never shrank, every following turn stacks a
fresh orphan on top of the last.

Publish the fence's deadline as an absolute monotonic instant
(CompressionCommitFence.deadline_monotonic) and give the auxiliary layer the
return leg it was missing: aux_stream_deadline() installs it thread-locally,
_ChatStreamAccumulator.feed() stops the stream once it passes, and
_run_protected_sync_provider_call propagates it onto the provider daemon
(thread-locals do not cross that boundary, so an owner-thread-only install
would be inert on exactly the path large-session compression takes).

Absolute, not relative: the deadline is unaffected by however long dispatch and
TTFT took before the accumulator was constructed. Checked as well as — not
instead of — the existing ceiling, so every caller without a host deadline is
byte-for-byte unchanged, and the "timed out" phrasing keeps _is_timeout_error
classification identical to a request timeout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EKrRS7LVgyHf2WQkEahSwu
2026-09-01 23:56:06 -07:00
DragonnZhang 9514d354ca fix(tool-search): validate deferred tool_call arguments against the concrete schema before dispatch
The generic tool_call(name, arguments: object) bridge hides a deferred tool's
real parameter schema from provider-native validation. Before this, only
top-level required-key absence was checked, so invalid enums, wrong types,
nested required fields and forbidden extra properties reached the handler
or MCP server. Now the call is coerced (same coerce_tool_args path normal
dispatch uses) and validated with the schema's declared JSON Schema draft;
failures return the path, constraint and parameters schema so the model
repairs the call in one round-trip. Fails open on missing/malformed schemas,
external $ref, or missing jsonschema.

Fixes #73175

Salvaged from #73179 onto current main (post core-tool deferral #97979).
Co-authored-by: teknium1 <teknium@nousresearch.com>
2026-09-01 23:30:33 -07:00
Teknium c5c9aa8d44 fix(gateway): hygiene compaction keeps the profile secret scope under multiplexing
Session-hygiene compaction ran _compress_context on a bare
loop.run_in_executor(None, ...) worker. Under gateway.multiplex_profiles the
profile secret scope and HERMES_HOME override are ContextVars installed by
the per-turn _profile_runtime_scope, and a bare worker starts with an empty
Context — so the summary model's get_secret(<PROVIDER>_API_KEY) failed
closed with UnscopedSecretError on EVERY hygiene pass and compaction
silently degraded to a lossy truncation (#100849 debug bundle:
'Failed to generate context summary: get_secret(SURPLUS_API_KEY) called
with no profile secret scope active').

- gateway/run.py: run both hygiene executor hops (detached-agent path and
  codex app-server path) inside copy_context().run, keeping the default
  executor so a fence-cancelled hung summary never occupies a gateway
  agent-work slot.
- agent/context_compressor.py: UnscopedSecretError is a missing-credential
  class failure — abort and preserve the session instead of dropping the
  middle window for a placeholder summary (same carve-out as 401/402/403).
- tools/daemon_pool.py: correct the salvaged docstrings — stdlib
  ThreadPoolExecutor only propagates contextvars from 3.14; nothing is
  stripped from the bundled runtime.
- tests: hygiene worker inherits caller ContextVars (fails on bare
  run_in_executor); UnscopedSecretError classified as access failure.

Live A/B (real get_secret in a run_in_executor worker, multiplex on, profile
.env scope installed): main -> UnscopedSecretError; fixed -> scoped value.
2026-09-01 22:28:52 -07:00
AgentLinker cac9db7caf fix(codex): retired stream requests must not synthesize a completed response
When a watchdog (TTFB / stream-idle / stale-call) force-closes a Codex
Responses request, the worker thread can still be draining SSE frames.
`_consume_codex_event_stream` returns `status=terminal_status`, which defaults
to `"completed"`, and its only truncation guard is
`if not saw_terminal and not output`. A mid-stream kill leaves
`saw_terminal=False` but `output`/text non-empty, so the partial text came back
as a `finish_reason=stop` response and got persisted as a finished assistant
turn — a long reply just stops mid-sentence with no error surfaced.

Observed as a long generation dying at `1. Create (6/6)` and never emitting its
end marker, with the truncated text already stored in state.db.

Fix: publish a per-request retirement token so the worker can tell it has been
retired.

- `agent/chat_completion_helpers.py`: `interruptible_api_call` installs
  `agent._active_codex_stream_request_token` before handing off to the worker
  (codex_responses only) and clears it at all four kill sites plus the worker's
  own `finally`. Retirement is cleared BEFORE `_close_request_client_once`,
  which can raise — every other call site wraps it in try/except, and a leaked
  token would let a later worker mistake itself for the owning attempt. The
  request-local `_codex_request_retired` mirror also swallows the transport
  error our own force-close causes, so the worker's local error cannot replace
  the watchdog's retryable TimeoutError (same split as `_request_cancelled`).
- `agent/codex_runtime.py`: `run_codex_stream` captures the token and raises
  `TimeoutError` from `interrupt_check` when it no longer owns the request —
  raising rather than breaking, because a break returns the partial `final`.
  The four stream callbacks also drop post-retirement frames so an abandoned
  attempt cannot stream tokens into the live turn's bubble (the gateway caches
  AIAgent instances per session).

`TimeoutError` is not an httpx / ConnectionError / RuntimeError subclass, so it
passes through the four `except` clauses around the consume call untouched.
No token installed (auxiliary callers such as `handle_max_iterations` drive
`_run_codex_stream` directly) means every check passes — behavior unchanged.

Tests: 5 new cases. Retirement raises instead of returning partial output;
post-retirement deltas stop reaching callbacks; the no-token path keeps its
existing terminal-frame tolerance; the watchdog installs and clears the token;
non-codex api_modes install nothing. A `_LazyCreateStream` helper is needed
because `_FakeCreateStream` materializes events in __init__, which would run
the retirement side effect before consumption starts.
2026-09-01 22:14:06 -07:00
David Metcalfe c905c2b4b5 fix(agent): honor model.streaming: false as a non-streaming escape hatch (#72901)
The conversation loop has forced stream=True for every turn — subagents
included — since #3120 (always-prefer-streaming for liveness health
checking). Self-hosted OpenAI-compatible backends with broken streaming
tool-call paths (e.g. vLLM --tool-call-parser qwen3_xml + reasoning
parser + MTP) can leak tool-call markup into plain text and return zero
tool_calls, so delegated tasks silently no-op instead of executing.

model.streaming was never a real config key, so users could not opt out.
Seed agent._disable_streaming from model.streaming: false at init; the
loop already routes that flag to the non-streaming path (the same path
used when a provider rejects streaming at runtime). Default stays
streaming-on, preserving #3120's behavior for everyone else. Orthogonal
to display.streaming (token rendering).

Tests: config->flag seeding (patched loader + real config.yaml E2E),
legacy string model section, multi-agent config propagation.
2026-09-01 22:14:06 -07:00
Teknium bd7cdd7c53 Merge origin/main into core-tool-deferral (resolve show_tip test seam onto the check_tips_enabled gate) 2026-09-01 21:49:14 -07:00
Teknium c5b99a3ee5 fix(agent): delegated children and cron turns stream again — inline, no worker
should_use_direct_api_call() contexts (gateway cron turns #62151, delegate_task
children #60203) were short-circuited onto the NON-streaming wire because the
interrupt worker wedges inside their nested thread pools. That dropped every
liveness property streaming provides: edge proxies kill the silent POST
(z.ai HTTP 524 — three retries later the child dies as "max_iterations"), and
the non-stream stale watchdog cannot tell a reasoning model's thinking phase
from a hung provider, so children die at exactly stale_timeout (#100260).

Keep those contexts on interruptible_streaming_api_call. The request now runs
INLINE on the conversation thread (no worker → the deadlock class stays
closed) while the existing poll loop — 30s heartbeat, stale-stream detector,
cross-thread interrupt abort — moves onto a monitor thread that only ever
aborts sockets, never dispatches (same shape as direct_api_call's watchdog
timer). Interactive sessions are unchanged: worker + poll loop as before.

should_use_direct_api_call() itself is untouched; only what it routes to.

Live A/B (real SSE server, real AIAgent.run_conversation):
  before: subagent/cron wire stream=None, request on conversation thread
  after:  subagent/cron wire stream=True, request on conversation thread
          cli unchanged (stream=True, spawned worker)
  inline stale detector kills a one-chunk-then-silence stream at budget;
  AIAgent.interrupt() from another thread unwinds the inline stream in 0.6s.

Co-authored-by: Expri-commits <184641533+Expri-commits@users.noreply.github.com>
2026-09-01 21:42:19 -07:00
Teknium 0ebe70d574 fix(agent): long-context tier recovery also rechecks the rebuilt request
The Anthropic long-context 429 handler restarts on row count alone,
the same shape #100614 fixed in the generic overflow handler. Arm the
same provider-overflow recovery flag there so the rebuilt request is
measured against the reduced window before the provider is retried.

The 413 (byte-scored) and output-cap (max_tokens) handlers are a
different yardstick and are left as-is.
2026-09-01 21:36:09 -07:00
Gille bdc46f5c09 fix(agent): recheck compressed requests after overflow 2026-09-01 21:36:09 -07:00
Teknium c0495c6bce fix(cli): context meter no longer sawtooths on reasoning models — show durable transcript, not last-request replay
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.

- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
  provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
  delta estimate (stale reasoning excluded on all but the newest assistant
  message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
  figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.

Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
2026-09-01 15:34:03 -07:00
Jeffrey Quesnelle c56f8cdd48 Merge pull request #100667 from NousResearch/feat/local-models-squash
feat: local models — managed llama.cpp runtime with one-click desktop  setup
2026-09-01 17:53:28 -04:00
Pedro Fontana b3576a29c3 Merge pull request #97354 from NousResearch/fix/nous-org-model-policy
fix(nous): honour the org model policy in the model pickers
2026-09-01 18:20:18 -03:00
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
Mariano Nicolini d7520b2822 fix(aux): seed the shared Nous catalog entry with the pickers' arguments 2026-09-01 16:34:48 -03:00
Teknium 419232d49b fix(codex): extend Happy-Eyeballs racing to Codex OAuth/auth clients; pin async native racing
#94388 (salvage of #70007) added RFC 8305 IPv6/IPv4 connection racing for
the direct synchronous chatgpt.com/backend-api/codex chat transport only.
Per the #13834 residual list, the auxiliary Codex paths were still serial:

- hermes_cli/auth.py Codex OAuth clients (token refresh at
  auth.openai.com/oauth/token, device-code login, token exchange, usage
  probe) each built plain httpx.Client()s — on broken-but-advertised IPv6
  every connect eats the full timeout per AAAA before IPv4 is tried, so
  auth fails where the official Codex CLI (which races) works.
- The async transport (async_mode=True in build_keepalive_http_client)
  had no explicit racing wired.

Changes:
- agent/process_bootstrap.py: add enable_happy_eyeballs_on_client() —
  installs the existing _HappyEyeballsSyncBackend on a ready-built sync
  httpx.Client's direct transports (default transport + mounts), skipping
  proxy-backed pools (HTTPProxy/SOCKSProxy: TCP connect goes to the proxy
  host, out of scope). Export it.
- hermes_cli/auth.py: add _codex_http_client() wrapper and use it for the
  five Codex OAuth/probe endpoints. Best-effort: falls back to default
  serial behavior if the backend can't be installed.
- Async transport: verified httpcore's AnyIOBackend already implements
  RFC 8305 natively via anyio.connect_tcp(happy_eyeballs_delay=0.25) —
  no custom backend needed. Documented in build_keepalive_http_client and
  pinned by tests (contract test on the anyio signature + a live
  regression test where a blackholed 100::1 IPv6 addr hangs and local
  IPv4 wins in ~250ms instead of the serial connect timeout).

network.force_ipv4 is unaffected: it patches socket.getaddrinfo below
all these layers and keeps working as the interim workaround.

Refs #13834; follows #94388 (9cce8725).
2026-09-01 12:08:11 -07:00
fangliquanflq a8ddb231aa fix(redaction): gate ambiguous assignment values 2026-09-01 12:07:16 -07:00
Lakshya Agarwal 428e084dcd feat(web): add Tavily web search and extract provider
This commit re-introduces the Tavily provider, which supports both search and content extraction capabilities, which was removed in #99199.
2026-09-01 10:56:49 -07:00
rainbowgits 622883bad7 fix(agent): accept marker-only finish_reason after stream supersession
A superseded writer was fencing the payload-empty terminal chunk, so
completed streams were mislabeled as mid-stream drops.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:27:06 -07:00
kshitijk4poor b81383ec21 fix(auth): compare pool-identity callers against all candidate keys
Two callers of get_custom_provider_pool_key compared against its single
preferred key and broke when the pool held the other identity:

- _prune_replaced_custom_model_config_credentials skipped only the
  preferred key, so a keyed provider's own legacy-named pool
  (custom:b.ai) was false-pruned of its current model_config credential
  when the preferred key resolved to the bare slug (b-ai).
- _seed_custom_pool seeded only when the pool key equaled the preferred
  key, so a legacy-named pool stopped being seeded from model.api_key.

Both now compare against the full custom_provider_pool_key_candidates
set. Also drops a redundant get_custom_provider_pool_key call from
_try_resolve_from_custom_pool (it returned candidates[0], doubling the
config traversal) and updates the two test files that monkeypatched the
removed module attribute.

Follow-up to #100413.
2026-09-01 22:42:27 +05:30
xxxigm 0bee5ff408 fix(auth): look up keyed custom providers by durable pool slug
hermes auth add stores providers.<key> credentials under the config
slug, but runtime only tried custom:<display-name> and then sent the
no-key-required placeholder. Try the slug first, keep the legacy
namespace as fallback, and thread provider_key/key_env through named
custom resolution.
2026-09-01 22:42:27 +05:30
Teknium 93591eccb5 fix(proxy): harden SSE DONE tracker — spec multi-line data joins, truthy lastOne, EOF-write guard
Follow-ups on the salvaged cluster:
- sse_done.py: dispatch SSE events at blank-line boundaries and join
  consecutive data: lines per the SSE spec (a split JSON event no longer
  reads as two malformed fragments that disable synthesis)
- accept integer/string-truthy lastOne sentinels (1 / "true") in both the
  proxy tracker and the agent stream reader
- server.py: guard the [DONE] append against client hangup at EOF and
  widen the interrupt tuple with OSError
- contributor email mappings for loulanyue and jon-nielsen
2026-09-01 10:12:21 -07:00
Jon Nielsen d304422b3d fix(streaming): extract finish_reason/usage before content-shape continues
vLLM >= 0.1.dev20051 merges finish_reason into the final content chunk.
When the SSE-echo guard is engaged at that moment (GLM-family tokenizers
emit standalone ':' / ' id' tokens mid-prose), the guard's content-shape
continue paths swallow the terminal chunk and finish_reason is never
captured, so a complete stream is misclassified as a mid-stream drop and
retried.

Extract finish_reason/usage at the top of the chunk loop body, before any
content-shape continue; the late tail-side extraction becomes redundant.

Addresses a second root cause of #94614 (the consume-gate fence is
covered by #94625; usage-side classification by #91376).
2026-09-01 10:12:21 -07:00
loulanyue 66d42e0dba fix(stream): do not misclassify stream with final usage chunk as mid-stream drop (#91373)
When stream_options={'include_usage': True} is requested, OpenAI-compliant
providers (e.g. vLLM, OpenAI, DeepSeek) emit a final usage-only chunk with
empty choices (choices=[]) and no finish_reason.

If the preceding text chunks did not explicitly set finish_reason, the
check in _call_chat_completions evaluated _text_only_dropped_no_finish to
True and returned a partial-stream stub with finish_reason='length'. The
conversation loop then assumed the connection was cut off and injected a
spurious continuation nudge, causing the model to rewrite the full answer.

Require usage_obj is None in _text_only_dropped_no_finish so streams that
delivered valid usage metadata complete cleanly with finish_reason='stop'.
2026-09-01 10:12:21 -07:00
rainbowgits ce7f805869 fix(proxy): append SSE [DONE] when Nous streams omit the sentinel
Complete Portal streams can finish with finish_reason/lastOne and clean
EOF without data: [DONE], which strict OpenAI clients treat as truncation.
Normalize at the hermes proxy boundary after clean EOF only.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:12:21 -07:00
joaomarcos ecdbcef7af fix(compression): roll the live transcript back when an in-place compaction commit fails (#99477)
`archive_and_compact()` is atomic: when it raises, every pre-compaction row is
still `active = 1` and the compacted set was never inserted. The rotation branch
already rolled the live transcript back to `messages_before_compression` in that
case, but the in-place branch — the default (`compression_in_place` defaults to
True) — did not, so `compress_context()` handed the caller the uncommitted
compacted list.

That list is marker-swept by `_strip_persistence_markers` (#57491) and the
post-commit `stamp_db_persisted_markers` (#98450) never ran, so the next
append-only flush treated the whole compacted transcript as new and INSERTed it
on top of the rows it was supposed to replace. The active set then held the
summary AND the turns it summarized: the next resume reloaded both, the token
count went up, preflight fired again, and every failed attempt appended another
copy of the protected head plus tail.

The in-place rollback mirrors the rotation branch and is gated on
`split_status != "in_place_committed"`, which is assigned on the statement
immediately after the atomic commit returns, so a committed compaction can never
be rolled back into a mismatch of the opposite sign.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EKrRS7LVgyHf2WQkEahSwu
2026-09-01 09:28:31 -07:00
overtoneblue e17276c7b4 agent: forward first stream chunk timestamp to post_api_request hook
interruptible_streaming_api_call already records first_chunk_at in its
per-attempt stream diagnostics (agent.stream_diag) for failure telemetry,
but the value was dropped on the success path. Stash it on the agent at
stream completion and forward it as first_chunk_at in the existing
post_api_request plugin-hook payload, alongside started_at/ended_at.

Consumers (observability plugins, shell hooks) can now derive TTFB
(first_chunk_at - started_at) and true generation throughput
(output_tokens / (ended_at - first_chunk_at)) without any new
instrumentation in the hot path.

Backward compatible: existing hook subscribers ignore unknown kwargs.
2026-09-01 08:30:45 -07:00
CoolStar aaca343110 fix(agent): preserve Bedrock redacted reasoning replay 2026-09-01 08:30:26 -07:00
Teknium f709bd88b6 feat(skills): render the configured create dir in every instruction that names the path
The skill_manage tool schema description, prompt-builder docs, and the
skills docs page now derive the creation path from skills.create_dir
(display_skill_create_dir()) instead of hardcoding ~/.hermes/skills/ —
so pointing the config at e.g. /opt/brain/skills changes what the agent
is told everywhere, with no SOUL.md fights or read-only chmod tricks.
Adds config default + docs section + 16 tests (incl. a read-only
profile-skills-dir scenario).
2026-09-01 07:32:45 -07:00
giwaov 42c2838674 feat(skills): skills.create_dir routes agent-created skills to a configured directory
Salvaged from PR #13996 (@giwaov, issue #13963), modernized onto current
main: config key renamed to skills.create_dir (per PR #81002's naming),
resolution centralized in agent/skill_utils.get_skill_create_dir() with
~/${VAR} expansion and HERMES_HOME-relative paths, and the directory is
folded into get_all_skills_dirs() so created skills are discovered,
trusted, findable, and patchable like local ones. Out-of-root creations
report their absolute path instead of crashing relative_to().
2026-09-01 07:32:45 -07:00
Teknium 51609a35f6 fix(auth): purge silent OpenRouter paid-default adoption (#81952 class fix)
Three kills at the shared chokepoints:

1. resolve_provider() now REFUSES env-key/pool auto-adoption of openrouter
   while the active config.yaml is corrupt (AuthError code=corrupt_config).
   A broken config falls back to DEFAULT_CONFIG, so tier-2 found no
   model.provider and tier-3/4 silently adopted the PAID openrouter provider
   against the user's real (unparseable) intent. New probe:
   hermes_cli.config.get_active_config_parse_failure(), recorded in the
   existing _warn_config_parse_failure() funnel keyed by (mtime_ns, size) —
   a fixed file clears the block immediately. Explicit provider requests
   are untouched.

2. auxiliary lane built-in OpenRouter fallback model is now a :free SKU
   (nvidia/nemotron-3-ultra-550b-a55b:free) instead of the paid
   google/gemini-3.6-flash. User-configured auxiliary.openrouter_model is
   honored untouched (paid-lane warning retained).

3. env->pool ingestion of OPENROUTER_API_KEY now logs a WARNING (once per
   process per provider) when a credential is newly ingested — ingestion
   itself stays allowed.

Fixes #81952 (silent-paid-default half; sibling PR covers the
non-interactive fail-closed guard).
2026-09-01 07:00:38 -07:00
kshitijk4poor 18a76be124 refactor(anthropic): drop dead allow_alias param, dedupe mcp wire normalization
Simplify-pass follow-ups on the salvaged alias logic:
- _to_oauth_wire_name's allow_alias kwarg was never passed False by any
  call site — removed.
- The _claimed_wire_names set comprehension re-implemented the same
  mcp_/mcp__/bare normalization ladder as _to_oauth_wire_name; both now
  share _normalize_to_mcp_wire. Behavior identical: aliased names are
  bare, so the collision probe's _MCP_TOOL_PREFIX + aliased equals
  _normalize_to_mcp_wire(aliased).
2026-09-01 16:27:20 +05:30