The speed=fast allowlist still gates on Opus 4.6, but the fast-mode
matrix has changed twice since it was written (verified against the
live docs, platform.claude.com/docs/en/build-with-claude/fast-mode):
- Opus 4.8 and Opus 5 SUPPORT fast mode (research preview, Claude API
only — not Bedrock/Vertex/Foundry).
- Opus 4.6 LOST fast mode on 2026-06-29. The parameter does not error:
requests silently run at standard speed and bill standard rates
(usage.speed: 'standard'). Today's allowlist therefore shows 4.6
users a fast toggle that does nothing, while denying it to the two
models that actually support it.
- Opus 4.7 never had it and hard-400s (unchanged).
- Dedicated '…-fast' model ids (OpenRouter's claude-opus-4.8-fast)
select fast inference via the model field and are explicitly
excluded from the param gate.
Both gates move in lock-step as before: the adapter param gate
(agent.anthropic_adapter._supports_fast_mode) and the CLI toggle gate
(hermes_cli.models._is_anthropic_fast_model). Docstrings now record
the history in both directions so the next matrix change has context.
## How to test
scripts/run_tests.sh tests/agent/test_anthropic_adapter.py tests/cli/test_fast_command.py -- -q
113 tests pass. The updated predicate/matrix tests fail against the
previous allowlist (verified by stashing the source changes). Tested
on Linux (aarch64).
The #14694 recovery clock (`_anti_thrash_recovery_deadline`) was a
process-local `time.monotonic()` value zeroed in `bind_session_state()`.
The gateway rebuilds the AIAgent (and its ContextCompressor) on every
cache eviction, so each fresh compressor bound to a durably tripped
session row (#69872) re-armed a full 300s window and the half-open probe
never fired — a long messaging conversation above the threshold stayed
blocked permanently.
Persist the deadline as a wall-clock epoch in a new
`sessions.compression_recovery_deadline REAL` column (declarative column
reconciliation; SCHEMA_VERSION 26 -> 27) with
`SessionDB.get/set_compression_recovery_deadline`. The compressor loads it
in `bind_session_state()` and writes it on change only via
`_set_anti_thrash_recovery_deadline()`. A fresh compressor with no stored
deadline still starts a full window blocked (#54923 restart contract); one
that loads an armed deadline resumes that window. Backward clock jumps are
bounded to one window. The 300s window is unchanged.
Minimal salvage of #100185 (the probe-lease/fencing state machine and
model_config-blob storage were not carried).
Refs #100185
Co-authored-by: Komzpa <me@komzpa.net>
Follow-up to the salvaged #100114 commit. Its two-pass anchor selection
scanned steers first and real user rows second, so a transcript shaped
[user A, tool(steer B), ..., user C] anchored the already-consumed steer B
over the newer real request C — the same replay class the PR set out to
fix. Replace it with one reversed positional scan that picks whichever
intent-bearing row is last (real role=user or steer-bearing role=tool),
and make the compressed-transcript steer check count only role=tool rows
(the only place the runtime delivers a steer), so a summary quoting the
marker cannot masquerade as live intent.
Adds S1/S2/S3 regression tests (steer dropped by compaction, steer
surviving in tail, newer user turn after steer) plus alternation and
use-exactly-once assertions.
Compression with display.busy_input_mode: steer embeds the follow-up
as an out-of-band marker inside the latest role=tool result. The
post-compression user-turn preservation path only classified
non-scaffolding role=user rows as real intent, so a compressed
transcript that contained no role=user row would discard the steer
and clone an older historical role=user message as the new active
turn, re-activating a previously consumed request.
Fix _ensure_compressed_has_user_turn to (1) treat a compressed
transcript that already carries a steer marker as having user intent,
and (2) prioritize the latest steer payload from the original
transcript over historical user cloning, inserting it as a proper
role=user turn via _insert_real_user_anchor. This preserves the
actual current intent exactly once and never turns history into new
input.
Closes#100053
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.
- model_metadata: memoize successful remote /models probes on disk
(cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
in-memory cache, so authority semantics are unchanged (reconciliation
still lands within 5 minutes) but the answer is shared across
processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
250ms instead of every 2s — up to 2s of dead air on every relayed reply.
Nothing here changes turn ordering: DMs and group rounds stay serial.
Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
Move the structural clone from the four call sites (auto review, codex
runtime, CLI /refine, gateway /refine) into AIAgent._spawn_background_review,
which every review path — immediate, idle-queue deferred, requeued — passes
through. Callers can no longer forget it, and the private helper is no longer
imported across hermes_cli/ and gateway/ package boundaries.
Tests now bind the real chokepoint (capturing at _spawn_background_review_now)
so they still fail if the clone is removed.
The clone-strip and root-write-through in the previous commit stop NEW forks
but leave installs that forked before upgrading in the broken state: each
profile keeps its own copy of the root grant, whichever profile rotated last
holds the only live refresh token, and root plus every sibling still hit
invalid_grant on their next refresh. The PR body asked those users to
re-auth at root and hand-edit profiles/*/auth.json; this makes it automatic.
`heal_forked_single_use_oauth_grants(provider)` (hermes_cli/auth.py) runs at
the top of a profile's `load_pool()` for SINGLE_USE_REFRESH_POOL_PROVIDERS.
Under the profile lock then the root lock it matches each profile OAuth row
to its root counterpart by lineage — same pool id (preserved by both fork
paths), same JWT account identity, same token material, else same provider +
same client (Anthropic pkce grants carry no claims) — keeps the copy with the
freshest rotation (`expires_at_ms` / `last_refresh` / JWT exp), writes it into
ROOT when root's is older, and strips the profile copy (pool rows, the
`providers.<id>` device-code block for Codex/xAI, and a profile-local
`.anthropic_oauth.json`) so the profile borrows root from then on. Root's
singleton and its hermes_pkce row are kept in step so root's own re-seed
cannot resurrect the spent pair.
Guarantees: idempotent (mtime-keyed clean mark skips the locked scan on the
per-call hot path); one INFO line per healed profile; API-key rows untouched;
a row with no root counterpart (root lost its grant, or an independent
account whose claims differ) is never deleted; only the two auth.json files
the root fallback already reads are touched — no environ/secret-scope reads.
`hermes auth list` / `hermes auth status <provider>` print the heal note.
Live repro (real imports, temp root + forge/atlas each holding a pre-fix
verbatim copy, forge already rotated RT0->RT1 into its own file, fake
single-use token endpoint): before — atlas None, forge AT2 (only in forge),
root None; server log 4x REUSE of spent RT0. After — forge's load heals to
root and rotates there, atlas and root select AT2, profiles/*/auth.json hold
no anthropic rows, server log exactly one ROTATE and zero REUSE.
Anthropic / Codex / xAI OAuth refresh tokens are single-use: a grant copied
into a second auth.json is one credential with two owners, and the first
profile to refresh it revokes the pair for every sibling (invalid_grant /
refresh_token_reused). Two code paths forked grants that way:
1. `hermes profile create --clone-all` and the dashboard/TUI
`mirror_credentials` flow copied auth.json (+ .anthropic_oauth.json)
verbatim. Both now run `strip_cloned_single_use_oauth_grants()`, which
drops OAuth rows for SINGLE_USE_REFRESH_POOL_PROVIDERS, the matching
`providers.<id>` device-code blocks, and the PKCE singleton file; API
keys are still copied. The clone reads the root grant through the
existing credential-pool root fallback.
2. A named profile with no local rows BORROWS the root grant via
`read_credential_pool()`'s fallback, but every persist
(`CredentialPool._persist`, `load_pool` reseed, `remove_index`) wrote the
rows into the profile's own auth.json — materializing a fork on the first
rotation. `persist_pool_entries()` now routes borrowed single-use rows
back to the root store (update-only, under the root lock; never falls
back to a local copy). A borrowed `hermes_pkce` rotation commits its
singleton to the root `.anthropic_oauth.json`, the borrower never prunes
root-seeded rows it cannot see the backing file for, and
`hermes -p <profile> auth add` persists only the profile's own rows.
Live repro (real imports, temp root + profiles, fake single-use token
endpoint): before — first profile rotation RT0->RT1 in profile only; root
and sibling then hit `invalid_grant`, `resolve_anthropic_token()` -> None.
After — rotation lands in root; root and both siblings select AT1, no reuse.
Direction per Teknium: stop cloning OAuth into profiles (ONE grant at root,
children inherit via context) rather than making clones survive. Supersedes
the clone-strip/root-write-through half of #100389 and the init-refresh idea
in #100703 (an expired-but-refreshable row already refreshes on select()).
Closes#100339
Co-authored-by: HexLab98 <liruixinch@outlook.com>
The one-shot reasoning-off retry changes a request parameter that is part
of the provider cache key on config-sensitive providers (Anthropic renders
thinking/effort into the prompt; OpenAI lists reasoning.effort as
prefix-affecting), so that request is a deliberate single cache miss.
Pin the bound: the request AFTER it must carry the configured reasoning
again and the system prompt must be byte-identical across the whole retry
sequence. Sabotage-verified (sticky flag -> test fails on request 3).
Docstring on _consume_ephemeral_reasoning_off states the cost honestly.
Follow-up to the #99622 salvage:
- agent/transports/chat_completions.py: the legacy (no provider profile)
chat_completions path always re-emitted extra_body.reasoning with
enabled=True, so both reasoning_effort: none and the one-shot
length-continuation override went out as {enabled: true, effort: none}.
Honor enabled=False / effort=none the way the profile path does.
- agent/conversation_loop.py: reset agent._ephemeral_reasoning_off at
turn start so a flag armed by an interrupted/errored turn can never
strip thinking from the next turn's first request.
- User-facing hints now name the real slash command (/reasoning); the
/thinkon//thinkoff commands do not exist.
- tests: wire-level regression (continuation request carries
reasoning.enabled=false) and a stale-flag turn-scope test.
GLM-5.3-flash on ollama-cloud with reasoning_effort=high can spend the ENTIRE
output cap on reasoning delivered in a separate field and return
finish_reason=length with no visible content (verified live: max_tokens=4096,
completion_tokens=4096, content empty).
The length-continuation path handled that shape badly:
1. the empty response was appended as an interim assistant fragment,
poisoning the transcript until the pre-call sanitizer healed it
(observed 3+ healings per turn on the reporting user's session);
2. every continuation re-ran with thinking ON, re-deriving the whole
thinking budget against a growing context, so 4 attempts still produced
nothing and the turn died with 'Response remains truncated after 4
continuation attempts'.
Now:
- interim assistant fragments with no visible content are never appended
(whichever way they got empty);
- a thinking-only truncation sets a one-shot reasoning-off override that
build_api_kwargs consumes for the next request, so the continuation
writes the answer instead of re-thinking it;
- the ceiling exit clears a pending override and, when every fragment was
empty, returns an actionable final_response instead of an invisible None.
Before turning hard stops on for unattended platforms, make sure they cannot
cut off normal work:
- Edit -> re-run is progress. A successful mutating call (write_file/patch,
a green terminal/execute_code, browser actions, job/message/cron/memory/
skill mutations) marks progress for every failing signature still being
counted this turn; the next identical retry restarts its streak instead
of accumulating toward exact_failure_block_after. A pure replay never
mutates anything between attempts, so it is still blocked at 5.
- Distinct red commands are diagnosis. For FAILURE_TOLERANT_TOOL_NAMES
(terminal, execute_code, process pollers, browser_navigate, web_extract)
same_tool_failure_halt_after warns but never halts.
- subagent and api_server keep the warn-only default: both are supervised
task loops with a live parent/client and do real edit -> re-run work.
Live A/B (real AIAgent platform=telegram, real patch+terminal, 8 rounds of
patch -> red check -> patch ...):
unmitigated branch: HALTED at round 6 (repeated_exact_failure_block)
this commit: COMPLETED all 8 rounds, final answer delivered
Loop shapes still stopped: identical failing read_file 8 calls,
identical successful terminal 5 calls (vs 602 on main).
Six new tests pin these flows; all fail on the unmitigated version.
Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.
- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
streak raises a halt (identical_call_streak_halt) at
hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
stay exempt; a changed result resets the streak; warning-only sessions are
unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
for pollers, or when results change.
Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
identical failing read_file main: 602 API calls, budget exhausted
branch: 8 calls, repeated_exact_failure_block
identical successful terminal main: 602 API calls, budget exhausted
branch: 5 calls, identical_call_streak_halt
Treat skill_view and skills_list as idempotent read-only tools so the existing no-progress guardrail can warn or block repeated identical skill loads. This prevents large skill outputs from being re-added to the context in tool loops.
Add regression coverage for repeated skill_view results under hard-stop guardrails.
The 10s hygiene_max_turn_hold_seconds budget (#92318) releases the arriving
user turn while the summary model is still streaming. For thinking summary
models (DeepSeek-V4-Flash etc.) whose reasoning prefix alone exceeds 10s,
the abandonment path ALWAYS cancelled the commit fence — 100% of the summary
attempt (including the full thinking prefix) was discarded on every turn,
permanently disabling auto-compression while paying the summary model 10s
of thinking per turn, and the flat 60s retry-after then blocked the
agent-side preflight from a fresh chance.
Structural fix (maintainer-chosen direction in #97963): decouple the turn
from the compression instead of holding the turn longer or making the hold
progress-aware (which would reintroduce the #90845 frozen-turn bug):
- CompressionCommitFence gains mark_commit_watermark_fenced() /
commit_watermark_fenced; compress_context marks the fence right after
capturing get_active_message_watermark() under the durable compression
lock (#75316/#87484) — the property that makes a LATE commit safe: rows
appended after compression start survive both commit paths verbatim as
cloned concurrent tail (archive_and_compact watermark= and
publish_compression_child watermark/watermark_ceiling).
- gateway hygiene turn-hold handler: when the fence is watermark-fenced,
the detached worker (already kept alive via
_defer_agent_cleanup_until_future_done) KEEPS its commit admission; the
user's turn proceeds on the uncompressed transcript at the same 10s
budget, and the summary is adopted at the worker's own watermark-fenced
commit boundary. Unfenced workers are cancelled exactly as before —
never worse than the status quo.
- No retry-after is armed while the kept-admission attempt runs (it would
block preflight adoption via the same-session cooldown); re-attempt
spacing is covered by the durable compression lock
(_session_has_compression_in_flight). If the worker ends WITHOUT
committing, a done-callback restores the flat non-escalating 60s
retry-after; a successful adoption resets the hygiene failure streak.
The streak never advances for a deferral either way.
- Docs: configuration.md hygiene_max_turn_hold_seconds one-liner updated
to describe deferred adoption and the thinking-model case;
config_defaults.py comment updated. Knob stays config.yaml-only.
Invariants preserved:
- 10s user-latency cap stays hard (#90845/#92318):
test_session_hygiene_turn_hold_budget_abandons_streaming_wait passes
UNMODIFIED (its worker is not watermark-fenced, so it pins the cancel
path through the public surface).
- Stale-clobber impossible: adoption only rides commits bounded by the
start watermark; the fence still gates admission and unfenced/late
results are discarded.
New regression tests (tests/gateway/test_session_hygiene_turnhold_adoption.py):
- watermark-fenced worker keeps admission, late summary is committed,
turn still released at the budget, no cooldown while running,
streak reset on adoption;
- kept-admission worker that ends without committing restores the flat
turn-hold retry-after (<=120s, names turn-hold, streak untouched);
- unfenced worker still cancelled and discarded (status quo).
Sabotage-verified: disabling the keep-admission branch fails the two new
adoption tests and leaves the unfenced-cancel test green.
Fixes#97963
PR #99779 gave the streamed chat.completions consumer the host's absolute
compression deadline. The two wires that consume their streams internally
still ran on their own, always-larger budgets after the host gave up:
- Codex Responses: clamp the re-armable watchdog's hard ceiling to the
published host deadline, so a live (re-arming) stream is severed the
instant the host stops waiting instead of at max(600s, 4x timeout).
- Anthropic Messages: the per-event hook now raises at the host deadline
and on an explicit hard cancel; create_anthropic_message lets that
TimeoutError abandon the stream (the with-block closes it) instead of
swallowing it as a callback failure.
Sabotage-verified: without the Codex clamp the new deadline test hangs past
its 25s harness cutoff; without the Anthropic hook the three Anthropic tests
fail.
CompressionCommitFence.set_total_ceiling_seconds documents its deadline as
"shared by the host and worker", but only the host ever read it. The worker's
streamed summary bounds itself with _aux_stream_total_ceiling() instead —
max(600, 4 * aux_timeout) — which is >= the host's total ceiling for every
configured timeout AND starts counting later (after pool admission,
_serialize_for_summary, prompt build and TTFT). A stream that outlives its
abandoned host is therefore not an edge case; it is the guaranteed outcome of
every total-ceiling timeout.
8207862212 closed the first half: a cancelled fence now releases the
compression owner, freeing its pool slot and session lease. Its own comment
leaves the second half open — the isolated provider daemon that holds the
socket keeps streaming "until the auxiliary stream's longer absolute ceiling
expires". With the #99692 reporter's auxiliary.compression.timeout: 600 that
is 2400s of an orphaned ~500K-token summary the fence is already guaranteed to
refuse, and because the session never shrank, every following turn stacks a
fresh orphan on top of the last.
Publish the fence's deadline as an absolute monotonic instant
(CompressionCommitFence.deadline_monotonic) and give the auxiliary layer the
return leg it was missing: aux_stream_deadline() installs it thread-locally,
_ChatStreamAccumulator.feed() stops the stream once it passes, and
_run_protected_sync_provider_call propagates it onto the provider daemon
(thread-locals do not cross that boundary, so an owner-thread-only install
would be inert on exactly the path large-session compression takes).
Absolute, not relative: the deadline is unaffected by however long dispatch and
TTFT took before the accumulator was constructed. Checked as well as — not
instead of — the existing ceiling, so every caller without a host deadline is
byte-for-byte unchanged, and the "timed out" phrasing keeps _is_timeout_error
classification identical to a request timeout.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EKrRS7LVgyHf2WQkEahSwu
The generic tool_call(name, arguments: object) bridge hides a deferred tool's
real parameter schema from provider-native validation. Before this, only
top-level required-key absence was checked, so invalid enums, wrong types,
nested required fields and forbidden extra properties reached the handler
or MCP server. Now the call is coerced (same coerce_tool_args path normal
dispatch uses) and validated with the schema's declared JSON Schema draft;
failures return the path, constraint and parameters schema so the model
repairs the call in one round-trip. Fails open on missing/malformed schemas,
external $ref, or missing jsonschema.
Fixes#73175
Salvaged from #73179 onto current main (post core-tool deferral #97979).
Co-authored-by: teknium1 <teknium@nousresearch.com>
Session-hygiene compaction ran _compress_context on a bare
loop.run_in_executor(None, ...) worker. Under gateway.multiplex_profiles the
profile secret scope and HERMES_HOME override are ContextVars installed by
the per-turn _profile_runtime_scope, and a bare worker starts with an empty
Context — so the summary model's get_secret(<PROVIDER>_API_KEY) failed
closed with UnscopedSecretError on EVERY hygiene pass and compaction
silently degraded to a lossy truncation (#100849 debug bundle:
'Failed to generate context summary: get_secret(SURPLUS_API_KEY) called
with no profile secret scope active').
- gateway/run.py: run both hygiene executor hops (detached-agent path and
codex app-server path) inside copy_context().run, keeping the default
executor so a fence-cancelled hung summary never occupies a gateway
agent-work slot.
- agent/context_compressor.py: UnscopedSecretError is a missing-credential
class failure — abort and preserve the session instead of dropping the
middle window for a placeholder summary (same carve-out as 401/402/403).
- tools/daemon_pool.py: correct the salvaged docstrings — stdlib
ThreadPoolExecutor only propagates contextvars from 3.14; nothing is
stripped from the bundled runtime.
- tests: hygiene worker inherits caller ContextVars (fails on bare
run_in_executor); UnscopedSecretError classified as access failure.
Live A/B (real get_secret in a run_in_executor worker, multiplex on, profile
.env scope installed): main -> UnscopedSecretError; fixed -> scoped value.
When a watchdog (TTFB / stream-idle / stale-call) force-closes a Codex
Responses request, the worker thread can still be draining SSE frames.
`_consume_codex_event_stream` returns `status=terminal_status`, which defaults
to `"completed"`, and its only truncation guard is
`if not saw_terminal and not output`. A mid-stream kill leaves
`saw_terminal=False` but `output`/text non-empty, so the partial text came back
as a `finish_reason=stop` response and got persisted as a finished assistant
turn — a long reply just stops mid-sentence with no error surfaced.
Observed as a long generation dying at `1. Create (6/6)` and never emitting its
end marker, with the truncated text already stored in state.db.
Fix: publish a per-request retirement token so the worker can tell it has been
retired.
- `agent/chat_completion_helpers.py`: `interruptible_api_call` installs
`agent._active_codex_stream_request_token` before handing off to the worker
(codex_responses only) and clears it at all four kill sites plus the worker's
own `finally`. Retirement is cleared BEFORE `_close_request_client_once`,
which can raise — every other call site wraps it in try/except, and a leaked
token would let a later worker mistake itself for the owning attempt. The
request-local `_codex_request_retired` mirror also swallows the transport
error our own force-close causes, so the worker's local error cannot replace
the watchdog's retryable TimeoutError (same split as `_request_cancelled`).
- `agent/codex_runtime.py`: `run_codex_stream` captures the token and raises
`TimeoutError` from `interrupt_check` when it no longer owns the request —
raising rather than breaking, because a break returns the partial `final`.
The four stream callbacks also drop post-retirement frames so an abandoned
attempt cannot stream tokens into the live turn's bubble (the gateway caches
AIAgent instances per session).
`TimeoutError` is not an httpx / ConnectionError / RuntimeError subclass, so it
passes through the four `except` clauses around the consume call untouched.
No token installed (auxiliary callers such as `handle_max_iterations` drive
`_run_codex_stream` directly) means every check passes — behavior unchanged.
Tests: 5 new cases. Retirement raises instead of returning partial output;
post-retirement deltas stop reaching callbacks; the no-token path keeps its
existing terminal-frame tolerance; the watchdog installs and clears the token;
non-codex api_modes install nothing. A `_LazyCreateStream` helper is needed
because `_FakeCreateStream` materializes events in __init__, which would run
the retirement side effect before consumption starts.
The conversation loop has forced stream=True for every turn — subagents
included — since #3120 (always-prefer-streaming for liveness health
checking). Self-hosted OpenAI-compatible backends with broken streaming
tool-call paths (e.g. vLLM --tool-call-parser qwen3_xml + reasoning
parser + MTP) can leak tool-call markup into plain text and return zero
tool_calls, so delegated tasks silently no-op instead of executing.
model.streaming was never a real config key, so users could not opt out.
Seed agent._disable_streaming from model.streaming: false at init; the
loop already routes that flag to the non-streaming path (the same path
used when a provider rejects streaming at runtime). Default stays
streaming-on, preserving #3120's behavior for everyone else. Orthogonal
to display.streaming (token rendering).
Tests: config->flag seeding (patched loader + real config.yaml E2E),
legacy string model section, multi-agent config propagation.
should_use_direct_api_call() contexts (gateway cron turns #62151, delegate_task
children #60203) were short-circuited onto the NON-streaming wire because the
interrupt worker wedges inside their nested thread pools. That dropped every
liveness property streaming provides: edge proxies kill the silent POST
(z.ai HTTP 524 — three retries later the child dies as "max_iterations"), and
the non-stream stale watchdog cannot tell a reasoning model's thinking phase
from a hung provider, so children die at exactly stale_timeout (#100260).
Keep those contexts on interruptible_streaming_api_call. The request now runs
INLINE on the conversation thread (no worker → the deadlock class stays
closed) while the existing poll loop — 30s heartbeat, stale-stream detector,
cross-thread interrupt abort — moves onto a monitor thread that only ever
aborts sockets, never dispatches (same shape as direct_api_call's watchdog
timer). Interactive sessions are unchanged: worker + poll loop as before.
should_use_direct_api_call() itself is untouched; only what it routes to.
Live A/B (real SSE server, real AIAgent.run_conversation):
before: subagent/cron wire stream=None, request on conversation thread
after: subagent/cron wire stream=True, request on conversation thread
cli unchanged (stream=True, spawned worker)
inline stale detector kills a one-chunk-then-silence stream at budget;
AIAgent.interrupt() from another thread unwinds the inline stream in 0.6s.
Co-authored-by: Expri-commits <184641533+Expri-commits@users.noreply.github.com>
The Anthropic long-context 429 handler restarts on row count alone,
the same shape #100614 fixed in the generic overflow handler. Arm the
same provider-overflow recovery flag there so the rebuilt request is
measured against the reduced window before the provider is retried.
The 413 (byte-scored) and output-cap (max_tokens) handlers are a
different yardstick and are left as-is.
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.
- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
delta estimate (stale reasoning excluded on all but the newest assistant
message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.
Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
#94388 (salvage of #70007) added RFC 8305 IPv6/IPv4 connection racing for
the direct synchronous chatgpt.com/backend-api/codex chat transport only.
Per the #13834 residual list, the auxiliary Codex paths were still serial:
- hermes_cli/auth.py Codex OAuth clients (token refresh at
auth.openai.com/oauth/token, device-code login, token exchange, usage
probe) each built plain httpx.Client()s — on broken-but-advertised IPv6
every connect eats the full timeout per AAAA before IPv4 is tried, so
auth fails where the official Codex CLI (which races) works.
- The async transport (async_mode=True in build_keepalive_http_client)
had no explicit racing wired.
Changes:
- agent/process_bootstrap.py: add enable_happy_eyeballs_on_client() —
installs the existing _HappyEyeballsSyncBackend on a ready-built sync
httpx.Client's direct transports (default transport + mounts), skipping
proxy-backed pools (HTTPProxy/SOCKSProxy: TCP connect goes to the proxy
host, out of scope). Export it.
- hermes_cli/auth.py: add _codex_http_client() wrapper and use it for the
five Codex OAuth/probe endpoints. Best-effort: falls back to default
serial behavior if the backend can't be installed.
- Async transport: verified httpcore's AnyIOBackend already implements
RFC 8305 natively via anyio.connect_tcp(happy_eyeballs_delay=0.25) —
no custom backend needed. Documented in build_keepalive_http_client and
pinned by tests (contract test on the anyio signature + a live
regression test where a blackholed 100::1 IPv6 addr hangs and local
IPv4 wins in ~250ms instead of the serial connect timeout).
network.force_ipv4 is unaffected: it patches socket.getaddrinfo below
all these layers and keeps working as the interim workaround.
Refs #13834; follows #94388 (9cce8725).
A superseded writer was fencing the payload-empty terminal chunk, so
completed streams were mislabeled as mid-stream drops.
Co-authored-by: Cursor <cursoragent@cursor.com>
Two callers of get_custom_provider_pool_key compared against its single
preferred key and broke when the pool held the other identity:
- _prune_replaced_custom_model_config_credentials skipped only the
preferred key, so a keyed provider's own legacy-named pool
(custom:b.ai) was false-pruned of its current model_config credential
when the preferred key resolved to the bare slug (b-ai).
- _seed_custom_pool seeded only when the pool key equaled the preferred
key, so a legacy-named pool stopped being seeded from model.api_key.
Both now compare against the full custom_provider_pool_key_candidates
set. Also drops a redundant get_custom_provider_pool_key call from
_try_resolve_from_custom_pool (it returned candidates[0], doubling the
config traversal) and updates the two test files that monkeypatched the
removed module attribute.
Follow-up to #100413.
hermes auth add stores providers.<key> credentials under the config
slug, but runtime only tried custom:<display-name> and then sent the
no-key-required placeholder. Try the slug first, keep the legacy
namespace as fallback, and thread provider_key/key_env through named
custom resolution.
Follow-ups on the salvaged cluster:
- sse_done.py: dispatch SSE events at blank-line boundaries and join
consecutive data: lines per the SSE spec (a split JSON event no longer
reads as two malformed fragments that disable synthesis)
- accept integer/string-truthy lastOne sentinels (1 / "true") in both the
proxy tracker and the agent stream reader
- server.py: guard the [DONE] append against client hangup at EOF and
widen the interrupt tuple with OSError
- contributor email mappings for loulanyue and jon-nielsen
vLLM >= 0.1.dev20051 merges finish_reason into the final content chunk.
When the SSE-echo guard is engaged at that moment (GLM-family tokenizers
emit standalone ':' / ' id' tokens mid-prose), the guard's content-shape
continue paths swallow the terminal chunk and finish_reason is never
captured, so a complete stream is misclassified as a mid-stream drop and
retried.
Extract finish_reason/usage at the top of the chunk loop body, before any
content-shape continue; the late tail-side extraction becomes redundant.
Addresses a second root cause of #94614 (the consume-gate fence is
covered by #94625; usage-side classification by #91376).
When stream_options={'include_usage': True} is requested, OpenAI-compliant
providers (e.g. vLLM, OpenAI, DeepSeek) emit a final usage-only chunk with
empty choices (choices=[]) and no finish_reason.
If the preceding text chunks did not explicitly set finish_reason, the
check in _call_chat_completions evaluated _text_only_dropped_no_finish to
True and returned a partial-stream stub with finish_reason='length'. The
conversation loop then assumed the connection was cut off and injected a
spurious continuation nudge, causing the model to rewrite the full answer.
Require usage_obj is None in _text_only_dropped_no_finish so streams that
delivered valid usage metadata complete cleanly with finish_reason='stop'.
Complete Portal streams can finish with finish_reason/lastOne and clean
EOF without data: [DONE], which strict OpenAI clients treat as truncation.
Normalize at the hermes proxy boundary after clean EOF only.
Co-authored-by: Cursor <cursoragent@cursor.com>
`archive_and_compact()` is atomic: when it raises, every pre-compaction row is
still `active = 1` and the compacted set was never inserted. The rotation branch
already rolled the live transcript back to `messages_before_compression` in that
case, but the in-place branch — the default (`compression_in_place` defaults to
True) — did not, so `compress_context()` handed the caller the uncommitted
compacted list.
That list is marker-swept by `_strip_persistence_markers` (#57491) and the
post-commit `stamp_db_persisted_markers` (#98450) never ran, so the next
append-only flush treated the whole compacted transcript as new and INSERTed it
on top of the rows it was supposed to replace. The active set then held the
summary AND the turns it summarized: the next resume reloaded both, the token
count went up, preflight fired again, and every failed attempt appended another
copy of the protected head plus tail.
The in-place rollback mirrors the rotation branch and is gated on
`split_status != "in_place_committed"`, which is assigned on the statement
immediately after the atomic commit returns, so a committed compaction can never
be rolled back into a mismatch of the opposite sign.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EKrRS7LVgyHf2WQkEahSwu
interruptible_streaming_api_call already records first_chunk_at in its
per-attempt stream diagnostics (agent.stream_diag) for failure telemetry,
but the value was dropped on the success path. Stash it on the agent at
stream completion and forward it as first_chunk_at in the existing
post_api_request plugin-hook payload, alongside started_at/ended_at.
Consumers (observability plugins, shell hooks) can now derive TTFB
(first_chunk_at - started_at) and true generation throughput
(output_tokens / (ended_at - first_chunk_at)) without any new
instrumentation in the hot path.
Backward compatible: existing hook subscribers ignore unknown kwargs.
The skill_manage tool schema description, prompt-builder docs, and the
skills docs page now derive the creation path from skills.create_dir
(display_skill_create_dir()) instead of hardcoding ~/.hermes/skills/ —
so pointing the config at e.g. /opt/brain/skills changes what the agent
is told everywhere, with no SOUL.md fights or read-only chmod tricks.
Adds config default + docs section + 16 tests (incl. a read-only
profile-skills-dir scenario).
Salvaged from PR #13996 (@giwaov, issue #13963), modernized onto current
main: config key renamed to skills.create_dir (per PR #81002's naming),
resolution centralized in agent/skill_utils.get_skill_create_dir() with
~/${VAR} expansion and HERMES_HOME-relative paths, and the directory is
folded into get_all_skills_dirs() so created skills are discovered,
trusted, findable, and patchable like local ones. Out-of-root creations
report their absolute path instead of crashing relative_to().
Three kills at the shared chokepoints:
1. resolve_provider() now REFUSES env-key/pool auto-adoption of openrouter
while the active config.yaml is corrupt (AuthError code=corrupt_config).
A broken config falls back to DEFAULT_CONFIG, so tier-2 found no
model.provider and tier-3/4 silently adopted the PAID openrouter provider
against the user's real (unparseable) intent. New probe:
hermes_cli.config.get_active_config_parse_failure(), recorded in the
existing _warn_config_parse_failure() funnel keyed by (mtime_ns, size) —
a fixed file clears the block immediately. Explicit provider requests
are untouched.
2. auxiliary lane built-in OpenRouter fallback model is now a :free SKU
(nvidia/nemotron-3-ultra-550b-a55b:free) instead of the paid
google/gemini-3.6-flash. User-configured auxiliary.openrouter_model is
honored untouched (paid-lane warning retained).
3. env->pool ingestion of OPENROUTER_API_KEY now logs a WARNING (once per
process per provider) when a credential is newly ingested — ingestion
itself stays allowed.
Fixes#81952 (silent-paid-default half; sibling PR covers the
non-interactive fail-closed guard).
Simplify-pass follow-ups on the salvaged alias logic:
- _to_oauth_wire_name's allow_alias kwarg was never passed False by any
call site — removed.
- The _claimed_wire_names set comprehension re-implemented the same
mcp_/mcp__/bare normalization ladder as _to_oauth_wire_name; both now
share _normalize_to_mcp_wire. Behavior identical: aliased names are
bare, so the collision probe's _MCP_TOOL_PREFIX + aliased equals
_normalize_to_mcp_wire(aliased).