init_agent (2711 LOC) becomes a ~280-line ordered orchestrator over
_resolve_api_mode / _finalize_routing / _init_* / _build_client /
_load_tools / _parse_compression_config -> CompressionSettings /
_resolve_context_length / _build_context_engine / ... phase helpers.
_build_client is further split per wire mode (_init_anthropic_client,
_init_moa_client, _init_bedrock_client, _init_openai_client with
_explicit_client_kwargs / _routed_client_kwargs). Statement order and
every side effect on the agent are preserved (AST body-parity checked
against origin/main).
Dedupe/dead code: drop _relay_moa_reference_event/_moa_reference_output_allowed
(zero callers; only their own test) and their test file; alias
_normalize_route_base_url; _parse_config_int replaces three copies of the
strict int parser; _cfg_flag replaces four inline truthy-set checks;
_client_kwargs_from_routed + _fallback_entries replace duplicated
routed-client/fallback-entry blocks; _warn_invalid_config_int unifies the
three log+stderr invalid-int warnings (byte-identical text);
_bedrock_region_from_url; _memory_provider_init_kwargs; the
host->default_headers if/elif chain becomes the _HOST_DEFAULT_HEADERS
dispatch table; callback params assigned from _CALLBACK_PARAMS.
Comments/docstrings hand-compacted to their rationale (invariants,
ordering, failure modes kept; issue numbers and narrative dropped).
test_pre_compress_checkpoint_contract source-check repointed at the
CompressionSettings field names.
Verified: tests/run_agent (2066 passed) + all agent_init-referencing tests
(1134 passed), get_tool_definitions() byte-identical vs origin/main, import
smokes for cli/run_agent/gateway.run/hermes_cli.main/agent.conversation_loop/
tui_gateway.server.
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.
Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:
- tools/terminal_scope.py: ContextVar holding the routed profile's
COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
resolves ONLY from it - an omitted key yields the defined default,
never os.environ. Unreadable/malformed policy installs a refusal
scope; terminal_tool / execute_code refuse instead of running under
ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
_profile_runtime_scope, tui_gateway session/build/turn scopes, cron
per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
(_get_env_config, _resolve_container_task_id shared key, orphan
reaper lifetime, degraded mode), gateway/platforms/base.py docker
media translation (volumes, shared key, persistence), runtime_cwd /
agent_init / skill_utils / code_execution_tool / file_tools cwd
anchors, prompt_builder / browser_tool / env_probe backend checks,
gateway footer, @-refs and slash-command cwd. env_probe resolves the
backend in the caller's context, since the probe worker thread does
not inherit the ContextVar.
Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.
Fixes#68559Fixes#94200Fixes#101132Fixes#95470
Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
Adds two bounded fast modes on top of the static /fast toggle, default OFF:
- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
window; requests inside it carry the provider fast param, later tool-loop
requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
user/assistant/tool history).
agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.
resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.
Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.
Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes#64785, #74730.
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
The conversation loop has forced stream=True for every turn — subagents
included — since #3120 (always-prefer-streaming for liveness health
checking). Self-hosted OpenAI-compatible backends with broken streaming
tool-call paths (e.g. vLLM --tool-call-parser qwen3_xml + reasoning
parser + MTP) can leak tool-call markup into plain text and return zero
tool_calls, so delegated tasks silently no-op instead of executing.
model.streaming was never a real config key, so users could not opt out.
Seed agent._disable_streaming from model.streaming: false at init; the
loop already routes that flag to the non-streaming path (the same path
used when a provider rejects streaming at runtime). Default stays
streaming-on, preserving #3120's behavior for everyone else. Orthogonal
to display.streaming (token rendering).
Tests: config->flag seeding (patched loader + real config.yaml E2E),
legacy string model section, multi-agent config propagation.
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.
- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
delta estimate (stale reasoning excluded on all but the newest assistant
message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.
Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
The #17929 init-time fallback block was nested inside the
'_explicit not in {auto, openrouter, custom}' guard, so an exhausted
openrouter credential pool skipped fallback_providers entirely and
AIAgent.__init__ raised 'No LLM provider configured' — surfacing on
Telegram only as the generic 'unexpected error' message.
2026-08-23 outage (~22:09-23:50, 60 occurrences incl. a cron job):
single-entry openrouter pool hit daily free-tier quota; chain had a
healthy local Ollama entry that never got tried because provider was
the default 'openrouter'.
Un-nest the block so any primary without usable credentials walks the
chain before failing; providers explicitly chosen by name keep their
dedicated missing-key diagnostic when both primary AND chain fail.
Regression tests cover the openrouter-exhausted path both with and
without a usable chain entry.
The native generateContent adapter never runs uncapped: when
model.max_tokens is unset it sends maxOutputTokens=65,535
(GEMINI_DEFAULT_MAX_OUTPUT_TOKENS) because Gemini treats an omitted cap
as a low internal default. The context compressor's trigger is
pct×(window − max_tokens), and constructing it with max_tokens=None
reserved 0 — so on a 128K Gemma window the trigger landed at 98,304
while the real safe input budget was 65,537, and the provider 400'd
before compaction fired.
Live repro (real imports, temp HERMES_HOME, native Gemini base_url,
window=131072, max_tokens unset):
before: compressor.max_tokens=None, threshold_tokens=98304,
wire maxOutputTokens=65535 → trigger ABOVE the safe budget
after: compressor.max_tokens=65535, threshold_tokens=64000 → below it
Scoped to the native Gemini wiring (provider names + native base_url via
is_native_gemini_base_url; the /openai compat endpoint is excluded). The
generic provider-default reservation gap remains tracked in #63839.
Reported by @Artemonim in #57275 (residual claim 4).
model.ollama_num_ctx is resolved AFTER the context compressor is
constructed, so a config that sets only ollama_num_ctx (without
model.context_length) ran every request at the smaller served num_ctx
while the compressor still targeted the probed GGUF window (e.g. 256K
Gemma metadata). The compaction trigger then sat several times above the
window the server actually serves and never fired — reproducing the
original #57275 'blows past the limit' symptom on current main.
Live repro (real imports, temp HERMES_HOME, config = {model:
{ollama_num_ctx: 65536}}, probed window 262144):
before: _ollama_num_ctx=65536, compressor.context_length=262144,
threshold_tokens=196608 (300% of the served window)
after: compressor.context_length=65536, threshold below the window
The clamp is one-directional (a num_ctx larger than the resolved window
never inflates the compressor) and reuses update_model() so every
threshold-derived budget recalibrates. Overlaps #60103 (silent-clamp
dead zone) — this is the init-order half.
Reported by @Artemonim in #57275 (residual claim 3).
Use a distinct runtime_capabilities field on agents, preserve compatibility with earlier snapshots, and resolve the canonical direct OpenAI endpoint when a cross-provider switch omits base_url. Keep ambiguous proxy routes fail-closed.
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
- Add rolling status bar metrics:
- cache hit ratio (◈) delta since model/compression reset
(hit = cache_read / prompt, verified against live logs)
- avg latency (◷) and throughput (↑ t/s) over last 10 API calls
(deque in agent, displayed in wide bar only)
- Add display.tui_statusbar_fields list to filter segments:
model, ctx, ctx_bar, cache_hit, latency, tps, compressions,
bg_tasks, bg_processes, bg_subagents, goal, duration, prompt,
idle, focus, yolo, stash, battery, title
Missing/null -> all enabled (backward compat). Unknown keys ignored.
Title gated via right-align; stash/battery also gated.
- Wide bar (≥76 cols) respects fields, narrow/medium filtered,
overflow trim preserved. Battery also respects display.battery.
No private data; mock data in tests.
Test: pytest tests/cli/test_cli_status_bar.py etc. 68 passed,
check-windows-footguns clean.
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.
- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
with a structural base-message identity check that fails closed on any
transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
compaction (codex_runtime), session reset/switch (run_agent), plus the
fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
estimation fallback (first request of a session).
An ACP client talks to a CLI over subprocess stdio: it returns a plain
completion object rather than an iterable stream, and it does not implement
the Responses API surface. Both exclusions spelled out `acp://copilot`, so the
next ACP client silently inherited the wrong defaults — a Responses upgrade
its shim cannot serve, and a streaming call that tries to iterate a
`SimpleNamespace`.
Match on the `acp://` scheme instead. `acp+tcp://` was already handled this
way; copilot-acp's behaviour is unchanged, and the new tests pin that a
non-ACP URL still upgrades, so this is not a blanket opt-out.
The legacy tail budget scales as threshold×target_ratio, which was designed
around 128K windows at a 50% trigger (~13K tail). On modern big-window
models with raised thresholds it silently hoards: a 1M-window session at
threshold 0.85 keeps a 170K-token verbatim tail (255K soft ceiling) out of
EVERY compaction, so a 540K manual /compress lands at ~290K and every
subsequent turn re-ships the hoard. Nobody chooses this; it is an artifact
of the formula outside its design envelope.
Lean mode (#87326, compaction-v2) was built for exactly this and its recall
was validated in the before/after eval (evals/compaction/results/): clamped
2.5%-of-window tail (10K floor / 25K cap), continuity carried by the
upgraded summary (digests, anchor index, verbatim user messages,
session_search recovery pointers). This flips the DEFAULT to lean; explicit
'tail_mode: legacy' in config keeps the old behavior exactly.
Also fixes a latent bug the flip exposed: update_model() re-assigned the
LEGACY formula directly when recomputing budgets, silently reverting a lean
compressor to the hoard on every mid-session model switch. The recompute
now routes through the mode-aware tail_token_budget property (regression
test included).
Surfaces: context_compressor.py defaults + getattr fallbacks, agent_init
parse default, DEFAULT_CONFIG, gateway _CACHE_BUSTING_CONFIG_KEYS gains
compression.tail_mode (mode changes now evict cached gateway agents like
target_ratio changes do), user + developer docs. Tests: 3 new default
contracts, legacy tests pinned explicitly, feasibility-skip scenario pinned
to legacy (under lean its payloads correctly become compressible).
E2E counterfactual (real imports, 1M window @ 0.85):
main default: legacy, tail 170,000 (ceiling 255,000)
head default: lean, tail 25,000 (ceiling 37,500)
head legacy: 170,000 (opt-out intact)
update_model to 400K: 10,000 (lean preserved across switch)
Independent review caught a compaction authority the gate missed:
post-turn micro-compaction (turn_finalizer -> _micro_compact) absorbs the
oldest exchanges into a rolling summary with no pre-compress checkpoint
hook in its path, and both compression.checkpoint_required and
compression.micro_compact could be enabled together — assistant evidence
could vanish into a summary the checkpoint filter later excludes, without
ever reaching the durable provider.
- agent_init: checkpoint_required forces micro-compaction off (warned),
mirroring the native-compaction suppression
- turn_finalizer: defense-in-depth guard at the call site (attribute is
plain mutable state a future path could flip on a live agent)
- behavioral regression test with a sabotage control (gate off proves the
harness reaches the call site; gate armed proves zero calls)
- docs + config example mention the suppression; stale v1 test header fixed
The fail-closed gate lived only in compress_context(), but two native
lossy owners compact without ever crossing it (review on #93996):
- codex app-server: in "native"/"off" auto-compaction mode (native is the
default) Hermes preflight is skipped and the codex agent compacts its
own thread inside run_turn() — the compress_context() rejection was
unreachable. init_agent now refuses checkpoint_required together with
api_mode=codex_app_server (BLOCKED_MISSING_PREREQUISITE, extracted as a
testable guard), and run_codex_app_server_turn() fails closed as
defense in depth before a turn can reach the codex-owned boundary.
- Responses server-side native compaction:
native_compaction_context_management() now returns None while the gate
is armed, so context_management never goes on the wire and the
checkpoint-aware Hermes compressor stays authoritative. The suppression
is logged once per process, not silently applied.
Regressions: checkpoint_required + app-server raises before run_turn()
(the session is never created); checkpoint_required keeps
context_management off the wire while the plain configuration still
produces it; the init guard refuses exactly the incompatible pair. Docs
and cli-config.yaml.example describe both bindings.
Refs #93986
Context compression is intentionally lossy. Deployments that archive
transcript evidence to an external durable store before compaction had no
way to guarantee the archive actually happened: MemoryManager.on_pre_compress
swallows provider failures by design, so a failed archive silently degraded
into data loss.
This adds an opt-in, provider-agnostic checkpoint contract:
- memory_provider: PRE_COMPRESS_CHECKPOINT_API_VERSION = 1; providers opt in
by advertising pre_compress_checkpoint_api_version. Version 0 keeps the
historical best-effort hook semantics.
- memory_manager: supports_pre_compress_checkpoint() capability probe;
on_pre_compress(require_checkpoint=True) propagates checkpoint-provider
failures and raises when no capable provider completed the checkpoint.
- conversation_compression: new compression.checkpoint_required config key
(default false, documented in cli-config.yaml.example). When enabled,
compaction fails closed with BLOCKED_MISSING_PREREQUISITE (the
uncompressed transcript is preserved) unless a checkpoint-capable provider
confirms the durable checkpoint. Providers receive normalized direct
user/assistant evidence: tool rows, system messages, tool-call wrappers,
and prior compaction summaries are filtered host-side into one stable
contract. codex_app_server compaction is rejected under the gate because
it exposes no truthful pre-compaction transcript boundary.
- hermes_state: persistent _compressed_summary column (declarative schema
migration via _reconcile_columns) so summary provenance survives process
restarts; only the resume model history carries the marker, keeping
get_messages_as_conversation on its existing contract.
- gateway: the lossy hygiene/auto-compact paths load the memory provider
(skip_memory=False) so a required checkpoint also guards those rewrites.
The gate arms only on an explicit boolean True (bare-MagicMock agents in
existing tests have truthy auto-attributes). Default behavior is unchanged:
checkpoint_required=false preserves best-effort semantics for all existing
providers. Contract tests, including a restart round-trip of the summary
marker, in tests/agent/test_pre_compress_checkpoint_contract.py.
Refs #93986
Follow-up structural pass on the review fix:
- Runtime provider, auxiliary resolution, model validation
(hermes_cli/models.py), live discovery (bedrock_model_ids_or_none),
and the Mantle URL/SigV4 fallbacks all resolve their region through
resolve_bedrock_runtime_region() — one canonical implementation of the
config-first priority instead of three hand-rolled copies.
- agent_init: drop the 'if "client_kwargs" in locals()' guard by
initializing client_kwargs unconditionally at the top of the else
branch; the Mantle kwargs hook is a documented no-op for non-Mantle
base URLs.
Route Bedrock-hosted OpenAI GPT-5.5 through the Bedrock Mantle OpenAI Responses endpoint with SigV4 request signing. Keep native Bedrock Converse and Claude Bedrock routing unchanged, and add picker/runtime regression coverage.
Cron jobs were constructed with skip_memory=True and a hard 'memory'
toolset denial, so MEMORY.md/USER.md never loaded and the memory tool was
stripped even from per-job enabled_toolsets. That was inconsistent with
kanban/delegate/gateway agents (which all get memory) and forced users
into hacky bypasses.
- cron/scheduler.py: skip_memory=False on the cron AIAgent; drop 'memory'
from _resolve_cron_disabled_toolsets; remove _strip_cron_memory_toolset
and its call sites
- agent/agent_init.py: update stale comment referencing the cron denylist
- tests: flip pinning tests to the new contract (memory enabled, per-job
memory toolset kept, user-level denylist still wins)
- docs: cron-internals + automate-with-cron no longer claim cron has no
persistent memory
Cron already sets skip_memory=True and denylists the memory toolset.
The default cron toolset still names memory, so init treated that as a
request and built MemoryStore. MEMORY.md then landed in the job prompt.
Treat a denylisted toolset as not requested, and strip memory from the
cron enabled list. Flush agents that actually want the memory tool are
unchanged (#65429).
The Zen relay serves *-free models (x-preview-f-free / Ox Alpha) ONLY
anonymously: any Authorization bearer it doesn't recognize is a 401
'Invalid API key' — including our no-key-required placeholder and valid
OpenCode GO subscription keys. The Go relay doesn't serve the free tier
at all ('Model x is not supported'). So the free model failed for every
Hermes user: keyless setups got the placeholder bearer, and OpenCode
subscribers sent a Go key to a relay that rejects it.
Fix (class-wide for all 8 current *-free Zen slugs, not just Ox Alpha):
- hermes_cli/models.py: is_opencode_zen_free_model / opencode_zen_free_runtime
/ opencode_zen_free_headers — one shared policy: free slugs pin to the
Zen relay with a keyless placeholder and an empty Authorization header
that overrides the OpenAI SDK's 'Bearer <key>'.
- runtime_provider.py: free slugs route through the keyless runtime before
the credential-pool/explicit/api_key paths (no key required; Go
selections heal to Zen). Paid models still fail closed without a key.
- agent_init.py + auxiliary_client.py: the placeholder key swaps in the
empty-Authorization headers at both client-build chokepoints.
Verified live (2026-08-21): anonymous chat/completions 200 incl. tools,
streaming, parallel; bad bearer 401; full E2E AIAgent turn with a real
terminal tool round-trip completes keyless under both opencode-zen and
opencode-go providers. Sabotage run: routing tests fail without the fix.
Normalize malformed memory config during initialization and bind per-target write permissions to the session MemoryStore so direct and staged writes cannot update a disabled built-in store.
Reuse the built-in store predicate during agent initialization and evaluate the config-backed memory tool check immediately after edits instead of applying the generic external-probe TTL.
Builds on @fattchris resolve_turn_limit salvage (#67696): flips the default
from a numeric cap to unlimited across all construction paths (CLI, agent_init,
run_agent subagents), adds inf/infinity/null to the unlimited spellings, and
sets DEFAULT_CONFIG agent.max_turns to null. The turn cap caused more problems
than it solved (silent mid-task truncation).
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.
Four pieces, and they only make sense together:
· an in-page engine that inventories what is interactable and performs the
verb, injected as source because it has to run inside the guest page;
· the preview.act.request bridge from the gateway into the pane;
· drive_preview, for acting: elements, click, type, scroll, press, and the
pane's own back/forward/reload;
· annotate_preview, for marking without acting.
Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.
Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.
Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
Composio eval traces showed Hermes wasting turns re-issuing identical tool
calls (same tool, same args, same result — 3x/4x in one run) and ending
turns by announcing an action it never took. Two conservative, config-gated
guards (agent.stall_guards, default true):
- Identical-call loop breaker: ToolCallGuardrailController.observe_identical_call
tracks the consecutive streak of (tool, canonical args, result-hash); on
the 3rd identical call a compact one-line notice is appended to that tool
RESULT at construction time (cache-safe — tool results are append-only).
Never blocks the call. Pollers (process, *_get_result, *_poll) are exempt
via STALL_GUARD_REPEATABLE_TOOLS. Streak resets on any different call,
changed result, or new turn. Observed on the raw result before the
tool-loop warning suffix so its changing count can't defeat matching.
- Said-continue-but-stopped recovery: trailing_continue_intent() detects a
short reply ENDING on an announced next action ('Let me now…', 'I will
now…', 'Next, I…'); the conversation loop feeds it into the EXISTING
intent-ack continuation path (same interim-assistant + user-nudge
mechanism, same codex_ack_continuations cap of 2), preserving message
alternation — no parallel recovery machinery.
Config: agent.stall_guards in DEFAULT_CONFIG; docs in configuration.md;
unit tests for streak/allowlist/reset/gate and detector pos/neg cases.
Un-fences OPENAI_MODEL_EXECUTION_GUIDANCE from the gpt/codex/grok substring
check and gives it its own injection gate, independent of
tool_use_enforcement, controlled by config.yaml `agent.execution_guidance`
(auto/true/false/list — same semantics as tool_use_enforcement). The "auto"
list (EXECUTION_GUIDANCE_MODELS) now also covers deepseek, kimi, qwen, glm,
minimax, mimo, and mistral.
Composio agentic-eval traces showed Hermes+DeepSeek/Kimi failing where
competitors passed: financial math done in prose, no read-back after
external writes, malformed identifiers "repaired", completeness claimed
despite count mismatches. The discipline block existed but those models
never received it.
The block is extended with compact clauses distilled from that analysis:
- external-write read-back (tool-call success is not task success; internal
file edits already confirmed by the tool are not re-verified)
- count reconciliation (declared totals/has_more are hard assertions)
- literal preservation (never normalize identifiers that fail a stated
format; lookup success does not validate a malformed token)
- retry-differently (empty/partial/suspiciously narrow results get a
broader retry before concluding)
- completion gated on verification (done = every named acceptance
criterion verified, never a plausible subset)
The todo tool description now encourages enumeration-as-checklist for
"all N items" tasks and gates completed status on verified work, never
intent.
Guidance is chosen once at session start keyed on model name, so the
system prompt stays byte-stable for the life of a conversation.
Supersedes/absorbs prior contributor proposals: #20588, #35087, #41874
(MiMo), #53847 (GLM tool-calls-as-text stall).
Co-authored-by: Mat-London <56627804+Mat-London@users.noreply.github.com>
Co-authored-by: intelac <8803887+intelac@users.noreply.github.com>
Co-authored-by: 6ylqq <51219463+6ylqq@users.noreply.github.com>
Co-authored-by: tauros1983 <267660491+tauros1983@users.noreply.github.com>
The identical 6-line try/except block for reading model.reasoning_echo
from config appeared in both agent_init.py (init) and
agent_runtime_helpers.py (switch_model). Extracted into
AIAgent._read_reasoning_echo_from_config() static method — net -1 LOC.
Address review feedback on PR #76503:
1. Init-time primary snapshot (agent_init.py:2756) was missing
reasoning_echo_flag — after fallback recovery the flag was
restored as False even when model.reasoning_echo: true was set.
2. Switch transaction snapshot (agent_runtime_helpers.py:2284) was
missing _reasoning_echo_flag — a failed client rebuild during
switch_model would leave the old provider with the new provider
echo policy.
Both omissions now fixed. No test regressions (56 passed).
Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
Add model.reasoning_echo (default false) and per-fallback-entry
reasoning_echo to preserve assistant reasoning_content when
replaying history to custom providers and OpenAI-compatible gateways
that proxy thinking-mode models (Kimi K3, GLM-5.2, DeepSeek, etc.)
but are not matched by the built-in host-based _REASONING_ECHO_RULES.
The flag is per-active-provider, not a global toggle:
- Primary: read from model.reasoning_echo at init and switch_model
- Fallback: set by try_activate_fallback from the fallback entry
- Restore: restore_primary_runtime copies the switch_model snapshot
Unlike PR #76019 global agent.reasoning_echo toggle, the
per-provider flag travels with the active provider — falling back to
a strict provider (Mistral, Groq, Cerebras) correctly strips
reasoning_content even when the primary had the flag enabled,
because the flag is False for the strict fallback.
Complements PR #27361 (dynamic detection) which fires after the first
API response; this PR covers turn-1 and history-replay-on-fresh-session
where dynamic detection has not fired yet.
Closes#76018
Refs: #27297, #27361, #76019
Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
One generic tool in the desktop_ui toolset: discover what is on screen,
highlight an element with narration, or hand the user a paged tour. No tour
content lives in the code — the agent authors each one live, which is what
makes 'how does this work?' answerable as a walkthrough instead of a wall
of text.
Rides the existing blocking-prompt bridge (tour.request/.respond) like
read_preview, so it works on every connection topology.
Implement Claude Opus review findings for Meta API support:
- Document in agent/agent_init.py that provider="meta" without an api.meta.ai URL falls through to chat_completions by design (URL-driven wire selection).
- Comment on suppression guard in hermes_cli/runtime_provider.py noting api.meta.ai is handled by _detect_api_mode_for_url.
- Replace inline __import__ with top-of-module import in tests/hermes_cli/test_model_switch_openai_api_mode.py.
- Rename test_meta_retention_not_sent_when_overridden -> test_meta_retention_override_wins in tests/agent/transports/test_meta_codex_cache.py.
- Add test in tests/agent/test_meta_agent_init.py for provider="meta" fallback without api.meta.ai URL.
- Add test in tests/agent/test_auxiliary_client.py for prompt_cache_retention: "24h" under _CodexCompletionsAdapter.
Source: Claude Opus review findings for feat/meta-api-support.
Relocate host_mandated_api_mode check from top of api_mode cascade to
fallback else branch so URL-based provider-slug rewrites (e.g.
api.anthropic.com -> provider='anthropic') always run first. Previously
the mandate branch set api_mode for api.anthropic.com without rewriting
provider, leaving provider='' and causing credential_pool_matches_provider
to fail closed and discard anthropic-scoped pools (#63425 regression
introduced in 8f60e8263).
The mandate is now a true fallback for hosts without an elif branch
(api.meta.ai -> codex_responses for 93-99% prompt-cache hits vs 0% on
chat, plus future mandates) with lazy import + try/except preserved.
Add regression tests: provider=None + api.anthropic.com URL implies
provider='anthropic'/api_mode='anthropic_messages' and preserves an
anthropic credential pool; provider=None + api.meta.ai URL implies
codex_responses.
- hermes_cli/providers.host_mandated_api_mode: add exact-hostname clause for
api.meta.ai → codex_responses (measured 0% cache on /chat/completions vs
93-99% on /responses with retention); update docstring.
- hermes_cli/runtime_provider._detect_api_mode_for_url: mirror clause for
api.meta.ai (exact hostname, #32243) to keep runtime resolver in lockstep.
- agent/agent_init: call host_mandated_api_mode early in api_mode cascade
(after explicit api_mode wins, before provider-name specials) via lazy
import; single source of truth, preserves user override.
- agent/transports/codex._default_prompt_cache_retention_for_request: return
24h for api.meta.ai unconditionally; build_kwargs setdefault preserves
override; Bedrock branch untouched.
- cli-config.yaml.example: add commented providers.meta example (api_mode
auto-detected).
- website/docs/developer-guide/adding-providers.md: list Meta alongside
Codex/xAI as codex_responses native provider with retention note.
- tests: add hermetic behavior-contract suites for mandate, retention,
content-addressed prompt_cache_key, reasoning passthrough, AIAgent init,
usage cache reporting, model-switch override, and config roundtrip; extend
test_model_switch_openai_api_mode with meta cases.
Per project policy, .env / HERMES_* env vars are reserved for
credentials; behavioural settings belong in config.yaml. Replaces
HERMES_DETERMINISTIC_EMPTY_GUARD and
HERMES_EMPTY_RETRY_COST_THRESHOLD_USD with an additive
agent.empty_response_guard section:
agent:
empty_response_guard:
enabled: true # false = legacy fixed 3-retry behaviour
cost_threshold_usd: 0.25 # per-attempt cost that halves the budget
- hermes_cli/config_defaults.py: new documented subsection under agent
(additive key, no config-version bump needed).
- agent/empty_response_guard.py: resolve_guard_settings() maps the
section to (enabled, threshold) with fail-open tolerance for
malformed values; guard_enabled()/_cost_threshold_usd() now read the
init-resolved agent attributes instead of os.environ.
- agent/agent_init.py: resolves the section once at init into
agent._empty_guard_enabled / agent._empty_guard_cost_threshold_usd,
following the existing tool_use_enforcement extraction pattern.
- Tests updated to config-attr injection; new TestResolveGuardSettings
covering malformed sections, YAML string booleans, bad thresholds,
and a DEFAULT_CONFIG sync check; new integration test proving
enabled:false restores the legacy 1+3-call behaviour.
Requested by isak-ialogics on PR #75115.
Live desktop E2E caught a write-ordering bug the automated E2E missed:
tui_gateway applies pending_title to state.db AFTER the first turn, but
the system prompt builds at turn START — the DB-title gate saw nothing
and the Bot Chat was cached protocol-less forever. The gateway now
hands the agent its intended title at construction and the gate checks
the hint first, DB second (CLI/messaging-gateway paths unchanged).
Live-verified on the running desktop: fresh bot's Bot Chat persisted
with the protocol section, handle, and roster in its system prompt;
regular sessions and SOUL.md untouched.
Replaces the plugin-side SOUL.md protocol append: on Bot-Mode-managed
installs (any profile carrying ui_meta['hermes-bots']) the prompt builder
injects the "Messaging other agents" section into every session of every
profile — including headless `hermes -p <bot> chat` sessions a teammate
starts — so bot handoffs work without mutating user-authored SOUL files.
- tools/bot_mode_probe.py: silent-when-unmanaged probe, cached per
(process, home), keyed off the agent's OWN home (not ambient
HERMES_HOME); silent when SOUL.md already carries the legacy section
- agent/system_prompt.py + agent_init.py + config_defaults.py: wired as
agent.bot_mode_protocol (default True), stable tier, byte-stable
across rebuilds (E2E-verified against the real build_system_prompt)
- tui_gateway profiles.list gains bot_mode_protocol capability flag;
the bundled plugin gates ALL SOUL protocol writes on it (backfill,
composeSoul, Edit save) — older gateways keep the SOUL-append path
- overhead: ~916 bytes, only on Bot-Mode installs; zero elsewhere
Supersedes the SOUL backfill half of Hermes-Bot-Mode#99 (credit
@kaduxo — the handle fix, `hermes profile list` correction, and
idempotent-append guards from that PR ship in the bundled plugin).
#87326 shipped the lean-compaction capability on the compressor; this adds
the config.yaml surface (compression.tail_mode: legacy|lean, default
legacy), DEFAULT_CONFIG entry, and docs on both the dev-guide compression
page and the user-guide configuration page, with zh-Hans parity.
Promote _split_model_config_default to hermes_cli/config.py as the single
shared helper for flattening dict-valued model.default/model.model config.
All 8 defense-in-depth sites now route through it instead of inlining
their own isinstance checks with inconsistent key orders.
Changes:
- Add split_model_config_default() to hermes_cli/config.py (public)
- cli.py: _split_model_config_default delegates to shared helper
- Fix key extraction order: agent_runtime_helpers.py was reversed
(default->model); now consistent (model->default) across all sites
- Remove provider-as-model-name fallback from main.py, oneshot.py,
model_tools.py, cli.py — provider is a routing key, not a model ID
- Add 'name' to _normalize_root_model_keys flattening loop and
_has_nested_default detection to cover the deprecated model.name alias
Tests: 118 passed + 1 skipped (cli_init, managed_scope, config).
E2E: 31/31 passed (config chokepoint, managed scope, crash site,
negative cases, edge cases).
A dict-valued model.default (e.g. {provider:..., model:...}) in config.yaml
was leaking into agent.model and crashing the agent at init:
AttributeError: 'dict' object has no attribute 'lower'
agent/agent_runtime_helpers.py: anthropic_prompt_cache_policy
This manifested on the Telegram gateway as an infinite reset loop: every
turn built an agent with model=dict, crashed during init, the gateway
treated the failed turn as a session needing reset, and /reset rebuilt the
agent and crashed again.
Coerce dict -> string at every model-resolution entry point so the value
is normalized once and never reaches a .lower() call as a dict:
- agent/agent_runtime_helpers.py: anthropic_prompt_cache_policy (the crash site)
- agent/agent_init.py: configured default model resolution
- cli.py: CLI config model + _normalize_model_for_provider
- hermes_cli/main.py: _has_any_provider_configured
- hermes_cli/oneshot.py: _run_agent model resolution
- hermes_cli/runtime_provider.py: _get_model_config default handling
- model_tools.py: _resolve_active_context_length
1. Use open_credentialed_url() instead of bare urlopen() in
templates.py apply_template() and probe_existing_customization().
Both send Authorization: Bearer headers; bare urlopen forwards
credentials on cross-origin redirects. The codebase has
open_credentialed_url() in hermes_cli/urllib_security.py that
strips credentials on cross-origin redirects — used by 4 other
modules.
2. Guard unavailable_reason() with the dedup set check before
calling it. The gateway builds a fresh AIAgent per message, so
without this guard unavailable_reason() (which calls _load_config()
→ stat + file read + JSON parse, and _check_local_runtime() →
importlib probes) runs on every gateway turn for an unavailable
provider, even though the warning is deduped after the first.
3. Move INDICATOR_GLYPH from Hindsight's eye emoji to a generic
brain (🧠) in core (agent/memory_provider.py). Hindsight overrides
with its own _HINDSIGHT_GLYPH (👁️) in recall_status() and
_emit_saving_indicator(). Other memory providers no longer inherit
Hindsight's brand mark as the default glyph.