Unpinned zero-credential installs now pick Exa or Parallel by the
parity of the per-process random session id (stable within a process,
even split fleet-wide) instead of always favoring Parallel. An explicit
hermes tools selection (web.backend / per-capability keys) bypasses the
split entirely; the runner-up vendor stays in the walk as fallback.
Live E2E: 6 fresh processes split 3/3 between vendors, each performed
a real keyless search via its picked endpoint; explicit pin verified.
Un-fences OPENAI_MODEL_EXECUTION_GUIDANCE from the gpt/codex/grok substring
check and gives it its own injection gate, independent of
tool_use_enforcement, controlled by config.yaml `agent.execution_guidance`
(auto/true/false/list — same semantics as tool_use_enforcement). The "auto"
list (EXECUTION_GUIDANCE_MODELS) now also covers deepseek, kimi, qwen, glm,
minimax, mimo, and mistral.
Composio agentic-eval traces showed Hermes+DeepSeek/Kimi failing where
competitors passed: financial math done in prose, no read-back after
external writes, malformed identifiers "repaired", completeness claimed
despite count mismatches. The discipline block existed but those models
never received it.
The block is extended with compact clauses distilled from that analysis:
- external-write read-back (tool-call success is not task success; internal
file edits already confirmed by the tool are not re-verified)
- count reconciliation (declared totals/has_more are hard assertions)
- literal preservation (never normalize identifiers that fail a stated
format; lookup success does not validate a malformed token)
- retry-differently (empty/partial/suspiciously narrow results get a
broader retry before concluding)
- completion gated on verification (done = every named acceptance
criterion verified, never a plausible subset)
The todo tool description now encourages enumeration-as-checklist for
"all N items" tasks and gates completed status on verified work, never
intent.
Guidance is chosen once at session start keyed on model name, so the
system prompt stays byte-stable for the life of a conversation.
Supersedes/absorbs prior contributor proposals: #20588, #35087, #41874
(MiMo), #53847 (GLM tool-calls-as-text stall).
Co-authored-by: Mat-London <56627804+Mat-London@users.noreply.github.com>
Co-authored-by: intelac <8803887+intelac@users.noreply.github.com>
Co-authored-by: 6ylqq <51219463+6ylqq@users.noreply.github.com>
Co-authored-by: tauros1983 <267660491+tauros1983@users.noreply.github.com>
Update the sibling tests that pinned the old use_gateway-writing
contract: image/video selector and reconfigure rows now assert the
single provider string ('nous' managed / 'fal' BYOK) plus legacy-key
popping, the stt/video picker writes drop the use_gateway expectation,
the web_server managed-browser select asserts the persisted 'nous'
cloud_provider, and explicit-local STT pins no-cloud-fallback against
a stored raw-config selection.
Two real gaps the CI-red sibling tests exposed:
- read_selection() treated EVERY raw stt.provider: local as the legacy
DEFAULT_CONFIG seed and reported no-selection — but the seed never
reached config.yaml (save_config strips schema defaults), so a
picker- or hand-written local pick was silently discarded and the
autodetect ladder could route an explicit local user to cloud STT.
A raw 'local' is now a genuine selection; the merged-view ambiguity
note replaces the over-broad shim (mirror comment updated in
nous_subscription._selected_provider and _get_provider).
- _reconfigure_provider was half-migrated: the tts/stt/browser/web
branches and the managed-category fallthrough still wrote
use_gateway flags and vendor names for managed rows. They now write
the single provider string ('nous' for managed rows) and pop the
legacy key, matching _write_provider_config.
New tests/tools/test_strict_provider_selection.py covers read_selection
semantics (legacy use_gateway interpretation, seeded stt local, empty
strings, browser.backend vs cloud_provider) and the three strict
behaviors per category: managed 'nous' selection wins over present
direct keys, a vendor selection with missing credentials raises the
selection-naming error with NO managed call, and never-configured
installs keep today's autodetect. Updated the tests that pinned the old
credential-first precedence (TTS resolver gateway override, STT silent
managed fallback, web invalid-backend reroute, video_gen picker writes).
Sabotage-verified: reverting the image FAL strict switch makes the new
managed-selection tests fail.
The chat_completions chokepoint fix (ultra->max for every model,
cherry-picked from #89509) has siblings with the same bug shape:
- codex.py: ultra->max was gated on gpt-5.6 only; now baseline for all
Responses-API models (backend-specific branches still override).
- Kimi top-level reasoning_effort: K3 accepts low/high/max only —
'medium' and upper-ladder levels were dropped to the medium default
(400s on K3, ladder inversion on K2). Full ladder mapped per family,
mirroring the kimi-coding plugin's K3 map.
- TokenHub: 'minimal' fell through to the 'high' default (asked least,
got most); full ladder now mapped onto low/medium/high.
- auxiliary_client Responses path: ultra->max alongside the existing
minimal->low clamp.
- custom provider plugin: ultra capped at max instead of forwarded
verbatim to GLM/vLLM/SGLang backends that reject it.
- copilot plugin: ad-hoc downgrade rules replaced with the shared
clamp_reasoning_effort_to_supported ladder walk so ultra/max resolve
to the strongest supported level instead of medium (#74295).
Sabotage-verified: new sibling-site tests fail 6/10 without the fixes.
Hermes' internal effort vocabulary extends the wire set with ultra
(documented by /reasoning as none..xhigh|max|ultra). OpenAI-compatible
wires — OpenRouter chief among them — accept exactly
max|xhigh|high|medium|low|minimal|none and reject the extension with
HTTP 400, so an ultra configured while the default model was Anthropic
worked (the Anthropic adapter maps its own levels) but leaked
untranslated the moment a per-job override pinned a non-Anthropic
model, failing every call for that job.
The wire-compat chokepoint for this transport previously mapped
ultra to max only for gpt-5.6; generalize the cap to every model.
test_no_config_no_credentials_returns_none pinned 'resolved provider
must be is_available()' — stale now that the keyless tier resolves
Parallel/Exa with is_available()=False + is_keyless_available()=True.
Accept keyed OR keyless-capable results (env-leak detection intact).
Exa and Parallel now each render as two picker rows in hermes tools —
'Free (keyless)' and 'Paid (API key)'. Selection persists to
web.provider_tier.<name>:
- free: always the anonymous public endpoint, even with a key set
- paid: always the keyed SDK path; missing key errors instead of
silently downgrading to the free tier (is_keyless_available also
returns False so the auto-fallback walk can't route there)
- unset: auto (key present -> paid, else keyless)
Mechanism: get_setup_schema() gains a 'variants' list the picker
flattens into sibling rows sharing one web_backend; selection writes
the tier via both _write_provider_config sites; active-row detection
matches the tier (auto mirrors use_keyless). Routing goes through a
single use_keyless() chokepoint shared by search+extract in both
providers.
Live E2E: tier=free with a fake key present searched keyless OK (a
keyed call would have 401'd); tier=paid without key errored naming
PARALLEL_API_KEY; picker rows verified for both vendors x both tiers.
With zero web credentials configured, web_search/web_extract previously
resolved to the nonfunctional firecrawl sentinel and errored. Now the
backend resolution walks a strictly-last keyless tier: Parallel's and
Exa's public anonymous MCP endpoints (the same free tiers opencode ships
as its default search path).
- plugins/web/keyless_mcp.py: minimal JSON-RPC tools/call client for
mcp.exa.ai + search.parallel.ai (SSE + plain JSON parsing, typed
errors, per-process random session id, no user identifiers)
- WebSearchProvider.is_keyless_available(): separate weaker tier that
never leaks into is_available(), so keyed setups are never pre-empted
- Exa/Parallel providers: route to keyless endpoints when their key is
absent; keyed SDK path unchanged
- registry + _get_backend(): keyless walk (parallel -> exa) strictly
after every keyed/importable candidate; check_web_api_key() lights
the tools up on zero-credential installs
- web.keyless_fallback config key (default true) to disable the tier
- docs: web-search.md + configuration.md
E2E-verified against both live endpoints from an isolated HERMES_HOME
(search + extract via the real dispatchers, disable-flag negative path).
Review polish from the 3-angle pass on the final stack:
- The warning now says WHICH shape leaked (top-level, extra_body, or both).
Relay injects top-level while request_overrides typically inject via
extra_body, so the shape identifies the offending middleware when
debugging.
- Fold the 'always returns a fresh mapping' assertion into the parametrized
real-endpoint test (the caller mutates the result with stream=True, so the
copy contract is load-bearing on no-drop paths too) and drop the
SimpleNamespace stub test it strictly subsumes. The nested-preserve stub
stays: the parametrized test only exercises top-level retention.
The wire guard only removed the top-level prompt_cache_retention kwarg, but
the OpenAI SDK merges extra_body into the outgoing JSON body, so a nested
extra_body.prompt_cache_retention reaches chatgpt.com/backend-api/codex just
the same and still triggers the non-retryable HTTP 400. Both injection
vectors are real and probe-verified: the Relay overlay's 'key not in
baseline' arm admits an interceptor-added extra_body, and
request_overrides={'extra_body': {...}} lands verbatim in build_kwargs
output.
Close the gap in the same helper: strip the nested field too (copy-on-write,
never mutating the caller's mapping), drop extra_body entirely when it
empties, and log the same warning. Compatible endpoints keep nested
retention untouched.
Mutation-verified: removing the extra_body leg fails both new nested tests.
Reported by egilewski's review on #89969.
The salvaged compatibility test stubs `_is_codex_backend=lambda: False` on a
SimpleNamespace, so it proves the helper honors its own boolean but not that
the boolean is right for any real endpoint. A predicate change that widened
the drop onto retention-supporting hosts would keep it green.
Adds a parametrized test that builds a real AIAgent per base URL and asserts
the drop only fires for chatgpt.com/backend-api/codex, while api.meta.ai,
bedrock-mantle.*.api.aws, api.openai.com and a same-host/different-path
backend keep their supported 24h value. Also asserts prompt_cache_key
survives untouched on every endpoint, since retention and cache-key routing
are independent and the guard must not disturb caching.
Verified non-vacuous: relaxing the guard's condition to drop on every
endpoint fails 4 of the 6 cases (Meta, Bedrock, OpenAI, non-codex path).
Drive-by on the guard itself: drop the dead `None` default on the `pop` that
is already gated by an `in` check, and record why the predicate is resolved
via getattr -- run_codex_stream is driven with lightweight stand-in agents
that lack `_is_codex_backend`, so a bare call would raise AttributeError.
Follow-ups on top of the salvaged #82631 surface:
- _select_surface: an unknown model id found in the live /images/models
catalog now ROUTES to the dedicated Image API instead of only logging a
hint — without this, a model picked from the live picker that postdates
the curated snapshot would fall onto chat-completions and fail. Curated
defaults stay pinned to chat (no behaviour change for existing setups);
offline probes still fall back to chat. _HINTED_MODELS removed.
- list_models (OpenRouter): union of the live GET /images/models catalog
(43 models today) and the chat-completions image models, deduped,
defaults first; curated metadata wins for known ids, API names for the
rest. Nous Portal (no /images route) keeps its chat-only catalog.
Offline fallback: static chain + curated Image API snapshot.
- Tests updated/added: unknown-id routing (flipped from the hint-only
pinning test), non-catalog id stays on chat, merged-picker union/dedupe/
order, Nous exclusion.
- Docs: image-generation.md gains the OpenRouter Image API section and an
editing-support row.
Live-verified: picker lists 43 models; generation succeeded through the
dedicated API on google/gemini-3.1-flash-lite-image and on the previously
unreachable black-forest-labs/flux.2-klein-4b (config-selected, no kwarg).
`_spawn_gateway_restart` already reuses an in-flight `hermes gateway
restart` child so a double-clicked button cannot start two racing
restarts. That guard evaporates exactly when it is needed most: the
child exits as soon as it has handed the restart to the supervisor (or
to the running gateway), long before the gateway is actually back, so a
stale cached dashboard frontend re-firing its own restart every few
seconds cleared the guard on every attempt and started a fresh restart
each time.
#89034 measured the result on an s6-supervised container: 77
`gateway-restart started` entries, 17 of them inside one minute. Each
one SIGHUPs a gateway that is still coming up, and killing it
mid-FTS5-write corrupted `state.db` ("database disk image is
malformed", 203x in agent.log) until the operator recreated the file by
hand.
Requests for the same profile within GATEWAY_RESTART_COOLDOWN_SECONDS of
the last spawn are now coalesced onto that spawn and logged, so a storm
produces one restart instead of one per request. The window is fixed
rather than health-gated on purpose: a gateway that never comes back
would leave a health-gated restart action permanently inert, which is a
worse failure than the flood it prevents. The cooldown state is kept
outside `_ACTION_PROCS` because completed action children are reaped out
of that table, and a guard that disappears when the child exits is the
bug being fixed.
Only the *frontend-flood* half of #89034 is addressed here. The s6
`finish` death-cap the report also asks for is a separate change to
`hermes_cli/service_manager.py` with a much larger blast radius, and is
left for a maintainer decision.
open_preview and read_preview could drive the page, but nothing could close
the pane. Same desktop_ui / session-source gate as the rest of the GUI tools.
Follow-up to the salvaged fix. Three parity gaps in the new CLI branch:
- Timeout arm dropped the denial-breaker addendum that the same function's
gateway arm and check_all_command_guards' CLI tail both append, so a
tripped breaker went unreported on a timeout.
- Human deny called _record_denial(), advancing a tally scoped to guardian
LLM DENY verdicts. Neither sibling CLI tail does this, so three
deliberate user denials escalated to breaker hard-stop text.
- The platform-marker half of the leak (HERMES_SESSION_PLATFORM set, no
HERMES_EXEC_ASK) was unpinned; it reaches the same branch.
Adds a platform-marker regression test plus two breaker-parity guards, and
clears the process-global _denial_tally in the shared fixture so a leaked
tally can't bleed the escalated addendum into unrelated assertions.
e37a0321eb fixed _run_approval_gate and check_all_command_guards: when
HERMES_EXEC_ASK (or a session platform marker) leaks into an interactive CLI
process with no gateway notify callback registered, those two functions now
prefer the registered CLI Dangerous Command panel over a silent
pending_approval nobody can see.
check_execute_code_guard — the whole-script gate for execute_code, a
separate function with its own copy of the same notify_cb-less
short-circuit — never got the same treatment. It doesn't even accept an
approval_callback parameter. In the same leaked-ask-mode-into-CLI scenario,
execute_code calls still silently drop into pending_approval with the panel
never shown, even though a CLI callback is registered.
Compute is_cli/approval_callback the same way the two fixed functions do,
and when _should_fall_through_to_cli_approval() says yes, run the same
hook-fire -> prompt_dangerous_approval -> hook-fire -> choice-branch
sequence _run_approval_gate's tail already uses, adapted to this function's
own message/persistence conventions (smart-denied session/permanent
suppression, denial-breaker addendum). Falls back to the existing
pending_approval behavior when no CLI callback is available.
Tests: 4 new cases in tests/tools/test_cli_approval_exec_ask_leak.py
mirroring the existing check_all_command_guards pair (approve/deny/timeout/
session-persistence). Mutation-verified: all 4 fail against the pre-fix code
and pass with it restored.
Neighbor suites: tests/tools/*approval* (291+ tests) and
tests/gateway/{test_approval_prompt_redaction,test_tui_approval_redaction,
test_plaintext_approval_routing,test_discord_exec_approval_content} +
tests/cli/test_cli_approval_ui.py all green. The 7 test_approval_mode_parity
/ test_nonrecursive_verification_artifact_cleanup failures seen in one full
batch run are pre-existing and independent of this change — confirmed by
re-running the identical batch with tools/approval.py stashed back to
pre-fix: the same 7 fail for the same reasons either way.
Note on an adjacent open PR: #65592 also touches check_execute_code_guard,
but an earlier, unrelated region of the function (adding an AST dangerous-
operation scanner to the "not is_gateway and not is_ask" auto-approve
branch). No semantic overlap with this fix's notify_cb-less branch; a small
rebase may be needed depending on merge order.
663fa68cd4 added an unguarded `agent._read_reasoning_echo_from_config()` call
to switch_model's core field swap. The fake agent here carries only the
attributes switch_model touches, so the new call raises AttributeError inside
the rollback-protected block — every field the test asserts on gets restored
to its pre-swap value and all four cases fail on main with
`assert 'opencode-go' == 'moa'`.
Teach the fake the reader, matching the production AIAgent shape.
The re-exec'd child inherits the console, so sys.stdin.isatty() still reported
a terminal and the update asked its local-changes question. By then the parent
shim had exited and the shell had taken the console back, so the prompt could
not be answered and the update sat there forever — worse than the lock it
replaced, because nothing recovers without closing the window.
Spawn the child with stdin closed. It then takes the same path the gateway and
Desktop updates take: honour updates.non_interactive_local_changes, which
stashes by default so nothing is lost, and keep going without asking.
Salvaged from #73811 per the consolidation triage on #76503, adapted to this
PR's per-active-provider model.reasoning_echo design.
Existing reasoning_echo tests hand-set _reasoning_echo_flag; none drives the
real config path init_agent uses:
load_config_readonly().get("model").get("reasoning_echo")
which is wrapped in `except Exception: False`, so a broken read would silently
disable the feature untested. This adds a temp-HERMES_HOME test that resolves a
named custom provider via the real resolve_runtime_provider (asserting
provider == "custom"), materializes the flag from a real config file, and
checks reasoning_content is preserved with the flag on and stripped with it off.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ab7R4NizLNbNyQZShtciKg
Add model.reasoning_echo (default false) and per-fallback-entry
reasoning_echo to preserve assistant reasoning_content when
replaying history to custom providers and OpenAI-compatible gateways
that proxy thinking-mode models (Kimi K3, GLM-5.2, DeepSeek, etc.)
but are not matched by the built-in host-based _REASONING_ECHO_RULES.
The flag is per-active-provider, not a global toggle:
- Primary: read from model.reasoning_echo at init and switch_model
- Fallback: set by try_activate_fallback from the fallback entry
- Restore: restore_primary_runtime copies the switch_model snapshot
Unlike PR #76019 global agent.reasoning_echo toggle, the
per-provider flag travels with the active provider — falling back to
a strict provider (Mistral, Groq, Cerebras) correctly strips
reasoning_content even when the primary had the flag enabled,
because the flag is False for the strict fallback.
Complements PR #27361 (dynamic detection) which fires after the first
API response; this PR covers turn-1 and history-replay-on-fresh-session
where dynamic detection has not fired yet.
Closes#76018
Refs: #27297, #27361, #76019
Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
Detection across every launch variant (argv[0], the zipapp __main__.py, the
main-module spec origin, the ancestor chain) plus the venv scoping that keeps
an unrelated hermes.exe from triggering a hand-off; the re-exec's argv, env
marker, loop guard and both fall-through paths; the pending-rename filter;
and the venv/.venv layout split.
Retires the reboot-deferred quarantine assertion along with the fallback.
Launch-variant cases from #89970 by @Akloenx123, pending-rename cases from
#88121 by @fangliquanflq.
The self-test grew a branch that spawns a Python child so pytest could prove
progress advances during one. It doesn't need to: /progress is answered from
its own runspace, so the existing hold already blocks the main thread, and
the spawn only exercised Invoke-HermesStep, which nothing here changes.
The Windows test now asserts the invariant instead of the self-test's stage
string, and the posix half -- previously untested, and the half that broke --
gets real coverage: serve-ui.py's wire shape, and posix.sh driven end to end
with a stub `hermes` that reports which stage was on screen while it ran.
llms.txt coverage is asserted in Python, but website/ sat on the Python skip
list, so a PR adding a docs page — or regressing the generator — went green
without ever running the test that checks the page is reachable. That is how
the index drifted to 53% coverage unnoticed.
The routing table listed 18 topics and had nothing to say about the rest of
the product, so an agent asked how to get bots to talk to each other answered
that it could not — while user-guide/bot-mode documented four ways to do it.
Point the catch-all at the published index, which is generated from the docs
tree on every build and so cannot fall behind the feature set. website/ is
never packaged, so the URL is the only complete self-knowledge a running
Hermes has; curl covers sessions where the web tools are disabled.
The section list decided membership as well as order, so it drifted as the
docs grew: 109 of 204 pages were absent from the index every LLM reads to
learn what Hermes does — Bot Mode, the desktop app, computer use, web search,
skins, Mixture of Agents, and 22 messaging platforms among them.
Enumerate the docs tree instead. SECTIONS now curates only which pages lead a
section; anything it does not name is absorbed under its path, and a page
matching no section lands in "More" rather than falling out. This also picks
up the three .mdx pages the .md-only glob never saw, points section landing
pages at the directory URL Docusaurus actually serves, and drops a curated row
still aimed at a guide moved to developer-guide/plugins in #59613.
Tests hold both directions against the filesystem rather than the enumerator,
so a page cannot go missing and a link cannot point at a page that moved.
Covers the new --yes/-y flag on , asserting the
parsed value reaches do_uninstall(skip_confirm=True) via the real
main() -> cmd_skills -> skills_command dispatch path. Mirrors the
install-flag test pattern in test_skills_install_flags.py.
The renderer's `tour.request` handler ships in the desktop bundle, but the
tool is offered by the backend, and the two update on different clocks. A
desktop build older than the tour tool receives the event in a renderer with
no branch for it, so `tour.respond` never comes and the agent blocks for the
full 45s deadline — once per action the model tries. A single "give me a
tour" turn (targets, then narrate, then stop) stacked those waits into
minutes of dead air, which is what got reported against #89620.
Hold a session's first action to a deadline a working renderer cannot miss,
and let an unanswered probe mark the bridge unavailable for that session:
later calls return immediately with an error naming the actual fix instead
of stalling again. Once a client has answered, real actions get the full
deadline back, so a preview tour injecting into a live page still works and
one slow action no longer condemns a live client. The verdict lives on the
session record, so it dies with the session and a new one re-probes.
The same five-action sequence goes from ~225s of dead air to a single 10s
probe. Toolset gating is unchanged: removing the tool outright needs a
client capability declared at session.create, which prompt caching means
can only take effect for a new session.