Follow-up to the session-scoping fix: _get_sudo_password_cache_scope()
and _resolve_container_task_id() carried byte-identical copies of the
HERMES_SESSION_KEY lookup (contextvar + os.environ fallback). Collapse
both onto one helper adopting the bare-import convention approval.py
already uses — get_session_env() implements the fallback internally, so
the old try/except could only fire on import failure, where silently
degrading to process-global semantics would reintroduce exactly the
cross-session contamination the fix prevents.
_resolve_container_task_id always returned "default", so _active_environments
shared a single SSHEnvironment across all WebUI sessions. When a user switched
from profile A (ssh_host=10.0.0.1) to profile B (ssh_host=10.0.0.2), the new
session found _active_environments["default"] already set to A's SSHEnvironment
and reused it — silently running every command on the wrong remote host.
Fix: when HERMES_SESSION_KEY is present (set per-session by the WebUI streaming
layer and per-message by the gateway via contextvars), return "session:<key>"
as the cache key instead of "default". Each session now owns its own slot in
_active_environments and always creates an environment from its own profile's
TERMINAL_SSH_HOST / TERMINAL_ENV config.
Behaviour unchanged in CLI mode (no HERMES_SESSION_KEY → still "default").
RL/benchmark task overrides (register_task_env_overrides) are unaffected.
Subagent task_ids inside a WebUI session collapse to "session:<key>" so they
continue to share the parent session's container.
Five new regression tests added to test_shared_container_task_id.py.
Bot Mode agents now DM teammates through a real tool instead of
hand-assembled shell commands. message_agent(target, message) validates
the target against the live roster, applies the sender's attribution
prefix server-side, and delivers over the existing proven transports
(hermes -p ... --query-file for local teammates, hermes peer dm for
peer gateways) as a tracked background process with notify-on-complete
— fire-and-forget, the reply wakes the sender on a later turn.
Containment: the schema is injected per-turn ONLY into a bot's
canonical 'Bot Chat' session on Bot-Mode-managed installs (same gate as
the protocol section); it is never registered in the tool registry or
any toolset, and dispatch re-gates on the session title so a forged
call from any other session refuses. The gate is session-stable, so the
tool list stays byte-identical across turns (prompt-cache safe).
The protocol section is rewritten to teach the tool and now carries the
teammate roster WITH ROLES (Bot Mode title + profile description), so
bots know who does what before picking a recipient. Roles and a
protocol version salt join the capability fingerprint: existing eternal
Bot Chats adopt the v2 protocol + tool with one epoch refresh, and a
rename/description edit refreshes the roster on the next message.
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.
- cron/scheduler.py: token parsing, target resolution (own profile /
named local profile / unknown -> skipped with warning), subprocess
delivery lane with cron.bot_chat_delivery_timeout_seconds (default
600s), preflight exemption, and bot-chat entries in
cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
must exist on this machine (fail at create, not at 3am); deliver schema
documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
token on the profile-scoped create so Desktop-side aliases can never
name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.
Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
The router treated any server-stamped principal as a bound lane, so with the
flag ON every authenticated dashboard/API session lost the legacy browser
backend even when no extension controller ever registered (scope_for_session
returns None -> ControllerUnavailable, no fallback) — while check_fns still
advertised the tools via the legacy OR-gate.
New broker.lane_registered() distinguishes the two cases:
- lane never registered -> generic callers keep the legacy backend
- lane registered (controller offline/ambiguous) -> fail closed, unchanged —
a control-this-tab session never silently jumps to another browser
Also makes the four non-allowlisted wrapped tools (cdp/console/vision/
get_images) behave correctly for never-registered lanes (legacy backend)
while staying fail-closed for registered lanes.
Surfaced during review of PR #85351.
Generic Hermes callers still use the existing browser backend when extension control is disabled or no server-bound controller identity exists.
Once the gateway binds a controller identity, missing scope, disconnect, or capability loss now fail closed instead of silently switching a control-this-tab request to another local or cloud browser. Covers the schema-build to dispatch disconnect race.
Treat unexpected controller transport loss as recoverable until each command's original deadline. Same-identity reconnects refresh transport and capability state, flush deferred cancels before new dispatch, and can complete already-started work.
Keep explicit detach and different controller/browser identity replacement terminal, owner-gate every inbound lifecycle frame, distinguish slow in-flight WebSocket writes from real send failures, and exclude browser-control session identity from shared shell snapshots.
Keep extension control opt-in and preserve existing browser backends unless an exact server-bound controller is available. Centralize protocol and capability admission across API and dashboard transports, make selected-controller results authoritative, bypass stale availability caches only inside bound requests, and serialize structured results for the existing tool contract.
Add a real browser_snapshot route-table/WebSocket E2E, strict admission and ownership regressions, public configuration and protocol documentation, and tests proving feature-off/no-controller compatibility.
Route the model-supplied target through _bound_error_text so a huge
bogus target can't bloat context, and restore the "Use 'memory' or
'user'" hint. Follow-up to HexLab98's review note on the salvage.
Normalize malformed memory config during initialization and bind per-target write permissions to the session MemoryStore so direct and staged writes cannot update a disabled built-in store.
Reuse the built-in store predicate during agent initialization and evaluate the config-backed memory tool check immediately after edits instead of applying the generic external-probe TTL.
The CLI (hermes cron create/edit) routes through cronjob(); removing the
parameter outright broke that lane (CI slices 6/9). The parameter is back
on the function, but CRONJOB_SCHEMA and the registry handler still omit
it — same pattern as the intentional model/provider/base_url omission.
New test proves a hallucinated reasoning_effort arg through the model
dispatch is dropped.
Standing policy: models do not make model-configuration decisions (the
only exception is user-defined profile selection in Bot Mode/kanban).
The per-job reasoning pin stays fully functional via
`hermes cron create/edit --reasoning-effort` and the job store; the
cronjob tool still SURFACES the pin in listings but cannot set it.
A schema-absence test pins the policy.
A cron job can now pin its own reasoning (thinking) effort, independent
of the global agent.reasoning_effort and per-model reasoning_overrides.
Heavy scheduled analyses can run at high while cheap recurring jobs run
at minimal, without touching the fleet-wide default.
- cron/jobs.py: new optional job field, validated at the storage choke
point against the canonical grammar via the shared
hermes_constants.parse_reasoning_effort (spelling-only; capability
clamping stays owned by the provider transports at send time, same as
config-set effort). Empty string clears on update; invalid values
raise ValueError before anything persists. Not a drift-guard axis.
- cron/scheduler.py: _resolve_job_reasoning_config resolves per-job pin
> agent.reasoning_overrides > agent.reasoning_effort at fire time,
after the auth-fallback model swap (the pin is model-independent by
design). A stored value that no longer parses warns and falls back to
config resolution instead of killing the tick.
- tools/cronjob_tools.py: reasoning_effort on BOTH mutation verbs
(create and update), conditional key in _format_job, schema documents
grammar/precedence/transport clamping/clear semantics. Agent-settable,
unlike model/provider pins: it cannot redirect spend to a different
model.
- hermes cron create/edit --reasoning-effort (empty string clears).
- Docs: cron feature page tip + CLI reference rows.
Tests: tests/cron/test_cron_reasoning_effort.py (32) — store contract,
scheduler precedence incl. byte-identical absent-field behavior and
garbage fallback, tool create/update/clear/error paths, schema surface.
Three live findings from rc.4 staging, all on the relay-fronted Slack
path, all with the failure observed in live logs before the fix:
1. Approval-send timeout is AMBIGUOUS, not failed (no re-ask).
send_exec_approval through the connector can time out with the card
already rendered — the connector may ack after the deadline (slow
platform API call, transient backpressure, event-loop stall) — and
the timeout-as-failure path re-sent and produced duplicate cards.
The outcome is now tri-state: sent / failed / ambiguous. Ambiguous =
no re-send, no text fallback; the prompt registration stays armed so
a late tap still resolves. Only a definite send error falls back to
text.
2. pending_approval tool results forbid re-issuing the command.
With one card correctly armed, the agent could still mint a SECOND
card by re-running a rephrased variant of the gated command after
reading the pending_approval tool result (observed live: same
command re-issued in a different form, two cards). The tool message
now instructs: do not re-run/rephrase; wait or report pending.
Applied to both the terminal and execute_code arms.
3. Draft interim AND seal frames carry format_hints.
format_hints are stamped on send, edit, and send_for_platform, but
both draft-frame builders (send_draft interim + _seal_open_draft
seal) shipped bare metadata. A streamed final therefore arrived at
the connector hintless and sealed as a plain code block while
non-streamed sends rendered native markdown blocks (observed live:
language-tagged block on send/edit, downgrade on streamed seal).
Both sites now stamp _with_format_hints_for_chat
(destination-resolved, same pattern as the existing lanes).
Verified live after the fix against the platform's stored message
payload: rich_text_preformatted with language field on a streamed
seal.
Tests: tri-state outcome unit tests (5), draft/seal hint stamping + knobs-
off regression control (2, RED-first), existing format-hints suite intact
(14/14). Mutation-verified: reverting the adapter hunk sends
test_draft_interim_and_seal_frames_carry_hints red; restore -> green.
Boundary sweep (text egress lanes crossing the frame contract): send ✓
(pre-existing) edit ✓ (pre-existing) send_for_platform ✓ (pre-existing)
draft-interim ✓ (this PR) draft-seal ✓ (this PR); task_card lane carries
no text content — exempt.
The one-shot keyless extract rescue (d1eefe6ac) treated ANY whole-batch
failure as a backend outage. A website-policy refusal also arrives as a
failed batch, so blocked URLs were routed through the free-tier ring:
in CI the ring's live fetch attempt returned a result for the wrong URL
or a bare None error, turning test_website_policy reds on main (slices
8/12 and 12/12) — and in production it would fetch content the user
explicitly blocked.
_rescue_extract now partitions policy blocks (blocked_by_policy flag or
policy error text) out of the rescue set: they are preserved verbatim,
only genuine failures ride the ring, and order/merge parity is kept.
Two sabotage-verified regression tests pin the class.
One conflict, gateway/relay/adapter.py send_for_platform: main added the
turn-final draft-seal interception (_sfp_metadata with the _interim_send
marker stripped, seal-or-fall-through); this branch added format-hint
stamping on the same frame. COMPOSED: the plain-send frame now stamps
_with_format_hints_for_platform over _sfp_metadata (the stripped copy),
so both the seal fall-through contract and the cron-lane block hints
hold. Note: the seal frame itself (op:draft final) does not stamp hints
— cron sends are never open drafts, so the flagship path is unaffected;
noted as a connector-PR follow-up for streamed interactive finals.
Every drive_preview action answered with the entire inventory — around 120
elements of ref, role, label, and an up-to-eight-rung `:nth-child` selector
chain. On a real app shell that was ~24.5k characters, re-sent after every
click, so a ten-step task paid for ten copies of a page that had barely moved.
Handles are now durable and legible. An element is named after what it is and
what it says — `btn-sign-in`, `inp-email`, `srch-search-projects` — minted once
per page and never reused, with duplicates disambiguated as `btn-edit`,
`btn-edit-1`. Each one remembers a stable attribute, its role, its accessible
name, and the nearest landmark it sits in, so when a framework destroys the
node and builds a new one the handle moves across and the agent is told
`rebound` rather than being handed a removal it has to react to and an addition
it has to re-read. The re-bind ladder is anchortree's (Apache-2.0), minus its
geometry rung, which can never clear the threshold on its own.
Because the handles hold, the first look at a page returns the inventory and
every look after it returns only what moved. `changed` carries the ref and
whichever of label/value/disabled actually shifted — role and selector are
absent by construction, since a change in either would mean the re-bind ladder
was looking at a different element. A delta gives way to a full re-read when
half the page is new, where there is nothing left to reuse.
The selector column is gone with it. It was 74% of the inventory on an
85-element page, nothing downstream ever read it, and a positional chain is
wrong the moment a sibling appears. An `#id` or `[data-testid]` survives when
the page offers one; everything else is addressed by handle.
Legibility is what makes the delta work rather than a nicety. `+ btn-sign-in`
on turn nine reads on its own, where `+ @e42` sends the model back to an
inventory twenty thousand tokens ago.
Measured on an 85-element app shell: 18,693 -> 4,930 characters for a baseline,
and a steady turn that moved two things costs ~200.
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.
Four pieces, and they only make sense together:
· an in-page engine that inventories what is interactable and performs the
verb, injected as source because it has to run inside the guest page;
· the preview.act.request bridge from the gateway into the pane;
· drive_preview, for acting: elements, click, type, scroll, press, and the
pane's own back/forward/reload;
· annotate_preview, for marking without acting.
Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.
Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.
Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
build_session_key embeds the workspace segment (scope_id) in every Slack
dm/group/thread key, but both cron seed helpers built their SessionSource
without it: the seeded row keyed agent:main:slack:dm:<chat>:<thread> while
a real scoped reply keys agent:main:slack:dm:<team>:<chat>:<thread> — a
row no reply ever resolves to. DMs were rescued only incidentally by the
legacy-key claim-once migration; scoped channels/threads got continuation
amnesia, and identical channel ids in two workspaces could collide.
Capture HERMES_SESSION_SCOPE_ID into the cron origin (_origin_from_env —
the session-context var async_delegation already snapshots), add scope_id
to _seed_cron_thread_session/_seed_cron_channel_session, and pass the
origin's scope at all three seed call sites.
Tests: scoped dm-thread / channel-thread / flat-channel seed-vs-reply key
equality through the real build_session_key, plus a two-workspace
non-collision guard.
When the chosen/keyed backend fails a web_search or web_extract call
(bad key, upstream outage, 5xx, raised exception), that single call
retries on the keyless free-tier ring instead of erroring. The next
call attempts the chosen backend again — no sticky failover, no state.
Resolves the keyed half of #78984/#32159 (keyless half landed in the
ring PR).
- tools/web_tools.py: _rescue_eligible (keyed ring vendors + non-ring
backends eligible; keyless-mode calls excluded — they already walked
the ring), _rescue_search/_rescue_extract (search annotates
rescued_from + backend_error naming the original failure and the
retry-next-call semantics; extract rescues only whole-batch failures,
partial failures pass through untouched; rescue failure preserves the
ORIGINAL backend error with the rescue note appended)
- both dispatchers wrap the provider call: failure-results AND raised
exceptions rescue; ineligible paths re-raise unchanged
- web.keyless_rescue config key (default true; implicitly off when
keyless_fallback is off); docs updated
Live E2E: keyed Tavily with an invalid key 401'd and the call was
served by the real ring with the rescue annotation; a second call
re-attempted Tavily first (statelessness proven); whole-batch extract
rescue returned real page content. 13 new tests; 67 green across the
keyless suites.
Fresh installs with zero web credentials now rotate web_search/
web_extract across FIVE vendors' public free tiers — Exa, Parallel,
Tavily, Firecrawl, Keenable — instead of a 2-vendor 50/50 split, with
next-in-line ring failover on rate limits (multi-hop until a vendor
serves or the ring is exhausted; served_by marks the actual vendor).
- plugins/web/keenable/: new bundled provider (search via /v1/search,
fetch via /v1/fetch; keyed Bearer or keyless with the mandatory
X-Keenable-Title app header). Credit: integration proposed by
Ilya Gusev (Keenable) in #49758; Free/Paid picker rows included.
- keyless_mcp: tavily/firecrawl/keenable keyless search+extract
wrappers, _KEYLESS_RING + per-process round-robin cursor (seeded by
the random session id, advances per unpinned request), pinned-vendor
entry (pin = start there; rotation off), paid-pinned vendors excluded
from the ring entirely.
- Tavily/Firecrawl providers route keyless traffic through the ring;
both are now default-on ring members (no longer selection-gated).
- web_tools/registry: keenable in backend sets, auto-detect, availability
probes; _keyless_preference() delegates to the ring cursor.
- KEENABLE_API_KEY in OPTIONAL_ENV_VARS; docs updated (ring semantics).
Live E2E: all 10 vendorXcapability paths (5 search + 5 extract) served
real results keyless; rotation cycled all five vendors over 5 dispatch
calls; double-throttle failover walked exa->parallel->tavily.
Salvaged from #78434 by @Slobaka (also the issue reporter; earlier than
the competing #78436). hermes doctor no longer paints a green web check
when the explicitly selected provider cannot initialize — web splits
into per-capability rows (web search / web extract) resolved through
the same registry resolvers the dispatchers use, with readiness from a
true availability probe (_provider_is_ready).
Keyless-tier integration on top of the salvage:
- _provider_is_ready counts is_keyless_available() as ready — keyless
mode is a working state, not a misconfiguration (zero-config installs
and selected-keyless Tavily/Firecrawl show ok, not warn)
- Tavily/Firecrawl gain is_keyless_available() (True only when
explicitly selected — they stay out of the zero-config fallback)
- doctor triggers plugin discovery before reading the registry (fresh
doctor processes saw an empty registry and warned on everything)
E2E: searxng-selected-without-URL warns (the #78412 repro);
zero-config, tavily-keyless, firecrawl-keyless all read ok;
parallel pinned paid without a key warns.
With memory.memory_enabled and memory.user_profile_enabled both false,
agent_init never builds a MemoryStore -- but check_memory_requirements()
returned True unconditionally and MEMORY_GUIDANCE was gated only on the
tool being present in valid_tool_names. So the tool shipped in every
request's schema while answering "Memory is not available" on every call,
and the system prompt still told the model to save durable facts there.
Gate both on the config flags, using the store predicate for the tool and
the already-resolved agent state for the guidance (config is not re-read
mid-conversation, so the prompt stays byte-stable). Either flag alone
still backs the tool, so only turning both off removes it.
This lets a user running a third-party provider (Hindsight, Mem0, ...)
turn the built-in files off without paying for the dead surface on every
API call. The provider's own tools are unaffected: hiding the built-in
tool moves the decision onto the toolset gate, and listing memory under
agent.disabled_toolsets remains the only switch that takes those down.
- Updated the Tavily API key description to clarify that it is optional and keyless access is supported.
- Modified the Tavily plugin and provider to handle requests with or without an API key, using Bearer authentication when the key is provided.
- Enhanced documentation to reflect the new keyless functionality and updated environment variable descriptions.
- Added tests to ensure correct behavior for both keyed and keyless requests.
Composio-style MCP servers return un-paginated 22-47K-char payloads that
sail under the generic 100K per-result spillover threshold, bloating
context and ballooning per-turn reasoning time on long conversations.
Competitors cap harder (OpenCode/pi 50KB, Claude Code 30K, Codex ~10K
tokens). Three changes:
- mcp_* tools spill at a tighter 50K default (BudgetConfig.mcp_result_size,
config-overridable via tool_budget.mcp_result_size_chars; pinned and
per-tool overrides still win; capped by the context-scaled default).
- The persisted-output preview now teaches recovery: page the saved file
with read_file or process with execute_code instead of re-requesting the
same data from the remote API.
- Untrusted/MCP string results are scanned (bounded, first 64KB) for
provider-side elision markers ('...N more items', "has_more": true,
'saved to sandbox', data_preview) and get ONE cache-safe incompleteness
notice appended at result-construction time, before untrusted wrapping —
so the model stops treating provider-elided enumerations as complete.
- Hard 2M-char allocation cap in mcp_tool.py (text, error, and
structuredContent paths) so a pathological multi-MB server payload is
bounded before it propagates, while ordinary large results reach
spillover intact. Distilled from #56060/#56072/#56511 (issue #56059);
supersedes their 50K lossy truncation with spillover-friendly semantics.
Docs: configuration.md spillover-budget section + cli-config.yaml.example.
Co-authored-by: Stoltemberg <215755014+Stoltemberg@users.noreply.github.com>
Co-authored-by: AlexFucuson9 <295703459+AlexFucuson9@users.noreply.github.com>
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
Unpinned zero-credential installs now pick Exa or Parallel by the
parity of the per-process random session id (stable within a process,
even split fleet-wide) instead of always favoring Parallel. An explicit
hermes tools selection (web.backend / per-capability keys) bypasses the
split entirely; the runner-up vendor stays in the walk as fallback.
Live E2E: 6 fresh processes split 3/3 between vendors, each performed
a real keyless search via its picked endpoint; explicit pin verified.
Un-fences OPENAI_MODEL_EXECUTION_GUIDANCE from the gpt/codex/grok substring
check and gives it its own injection gate, independent of
tool_use_enforcement, controlled by config.yaml `agent.execution_guidance`
(auto/true/false/list — same semantics as tool_use_enforcement). The "auto"
list (EXECUTION_GUIDANCE_MODELS) now also covers deepseek, kimi, qwen, glm,
minimax, mimo, and mistral.
Composio agentic-eval traces showed Hermes+DeepSeek/Kimi failing where
competitors passed: financial math done in prose, no read-back after
external writes, malformed identifiers "repaired", completeness claimed
despite count mismatches. The discipline block existed but those models
never received it.
The block is extended with compact clauses distilled from that analysis:
- external-write read-back (tool-call success is not task success; internal
file edits already confirmed by the tool are not re-verified)
- count reconciliation (declared totals/has_more are hard assertions)
- literal preservation (never normalize identifiers that fail a stated
format; lookup success does not validate a malformed token)
- retry-differently (empty/partial/suspiciously narrow results get a
broader retry before concluding)
- completion gated on verification (done = every named acceptance
criterion verified, never a plausible subset)
The todo tool description now encourages enumeration-as-checklist for
"all N items" tasks and gates completed status on verified work, never
intent.
Guidance is chosen once at session start keyed on model name, so the
system prompt stays byte-stable for the life of a conversation.
Supersedes/absorbs prior contributor proposals: #20588, #35087, #41874
(MiMo), #53847 (GLM tool-calls-as-text stall).
Co-authored-by: Mat-London <56627804+Mat-London@users.noreply.github.com>
Co-authored-by: intelac <8803887+intelac@users.noreply.github.com>
Co-authored-by: 6ylqq <51219463+6ylqq@users.noreply.github.com>
Co-authored-by: tauros1983 <267660491+tauros1983@users.noreply.github.com>
Two real gaps the CI-red sibling tests exposed:
- read_selection() treated EVERY raw stt.provider: local as the legacy
DEFAULT_CONFIG seed and reported no-selection — but the seed never
reached config.yaml (save_config strips schema defaults), so a
picker- or hand-written local pick was silently discarded and the
autodetect ladder could route an explicit local user to cloud STT.
A raw 'local' is now a genuine selection; the merged-view ambiguity
note replaces the over-broad shim (mirror comment updated in
nous_subscription._selected_provider and _get_provider).
- _reconfigure_provider was half-migrated: the tts/stt/browser/web
branches and the managed-category fallthrough still wrote
use_gateway flags and vendor names for managed rows. They now write
the single provider string ('nous' for managed rows) and pop the
legacy key, matching _write_provider_config.
An explicitly stored browser.cloud_provider that names no registered
plugin now raises the honest selection-naming error instead of warning
and silently auto-detecting; the auto-detect walk (including the managed
gateway entitlement probe) runs only when no cloud_provider key was ever
written. The 'nous' selection routes to the Browser Use provider, whose
config resolver is now a strict switch: 'nous' => managed only, stored
vendor => direct BROWSER_USE_API_KEY only with a selection-naming error
when missing. Camofox is selected via browser.cloud_provider: camofox;
CAMOFOX_URL stays the server ADDRESS only and can no longer override an
explicit different selection (never-configured installs keep the legacy
env-var activation).
Both _resolve_openai_audio_client_config resolvers now switch on the
stored provider string: 'nous' (or legacy use_gateway: true) => managed
openai-audio gateway only, erroring by selection name when unentitled —
the STT twin previously never read the stored gateway intent at all, so
a direct OPENAI_API_KEY silently overrode the Nous Subscription pick;
stored vendor => direct credentials only with a selection-naming error
on missing keys (no silent managed fallback); never-configured keeps the
legacy ladder. DEFAULT_CONFIG stops seeding stt.provider: local, and the
seeded value on existing configs is treated as no-selection so autodetect
keeps working for that installed base.
_get_backend returns the stored web.backend verbatim (mapping the managed
'nous' selection to the firecrawl provider) — unknown names surface the
honest selection-naming error at dispatch instead of silently rerouting
through the credential ladder, which now runs only on never-configured
installs. _get_capability_backend no longer discards an explicit
search/extract backend when its availability probe fails. The firecrawl
client resolves strictly: 'nous' => managed gateway only (unavailable =>
selection-naming error), stored vendor => direct only (no FIRECRAWL key
=> error, never a silent managed fallback billed to Nous).
Add read_selection()/selection_exists()/selection_error() to
tool_backend_helpers: one provider string per category ('nous' = managed
Nous Tool Gateway, vendor name = direct with the user's own credentials,
no key ever written = legacy credential autodetect). Legacy configs are
interpreted at read time only (use_gateway: true => nous); nothing is
migrated on disk, and the DEFAULT_CONFIG-seeded stt.provider: local is
treated as never-configured.
_resolve_managed_fal_gateway / _resolve_managed_fal_video_gateway now
switch on that string: 'nous' routes managed only (unentitled => error
naming the selection), a stored vendor routes direct only (missing
FAL_KEY => error naming FAL_KEY and the selection, no silent managed
reroute), and FAL_KEY presence no longer selects the route. Krea's
model-driven managed interception now requires no stored provider (or
the managed selection) instead of merely provider != krea, and the
image/video registries map the 'nous' selection to the FAL plugin.
With zero web credentials configured, web_search/web_extract previously
resolved to the nonfunctional firecrawl sentinel and errored. Now the
backend resolution walks a strictly-last keyless tier: Parallel's and
Exa's public anonymous MCP endpoints (the same free tiers opencode ships
as its default search path).
- plugins/web/keyless_mcp.py: minimal JSON-RPC tools/call client for
mcp.exa.ai + search.parallel.ai (SSE + plain JSON parsing, typed
errors, per-process random session id, no user identifiers)
- WebSearchProvider.is_keyless_available(): separate weaker tier that
never leaks into is_available(), so keyed setups are never pre-empted
- Exa/Parallel providers: route to keyless endpoints when their key is
absent; keyed SDK path unchanged
- registry + _get_backend(): keyless walk (parallel -> exa) strictly
after every keyed/importable candidate; check_web_api_key() lights
the tools up on zero-credential installs
- web.keyless_fallback config key (default true) to disable the tier
- docs: web-search.md + configuration.md
E2E-verified against both live endpoints from an isolated HERMES_HOME
(search + extract via the real dispatchers, disable-flag negative path).
open_preview and read_preview could drive the page, but nothing could close
the pane. Same desktop_ui / session-source gate as the rest of the GUI tools.
Follow-up to the salvaged fix. Three parity gaps in the new CLI branch:
- Timeout arm dropped the denial-breaker addendum that the same function's
gateway arm and check_all_command_guards' CLI tail both append, so a
tripped breaker went unreported on a timeout.
- Human deny called _record_denial(), advancing a tally scoped to guardian
LLM DENY verdicts. Neither sibling CLI tail does this, so three
deliberate user denials escalated to breaker hard-stop text.
- The platform-marker half of the leak (HERMES_SESSION_PLATFORM set, no
HERMES_EXEC_ASK) was unpinned; it reaches the same branch.
Adds a platform-marker regression test plus two breaker-parity guards, and
clears the process-global _denial_tally in the shared fixture so a leaked
tally can't bleed the escalated addendum into unrelated assertions.
e37a0321eb fixed _run_approval_gate and check_all_command_guards: when
HERMES_EXEC_ASK (or a session platform marker) leaks into an interactive CLI
process with no gateway notify callback registered, those two functions now
prefer the registered CLI Dangerous Command panel over a silent
pending_approval nobody can see.
check_execute_code_guard — the whole-script gate for execute_code, a
separate function with its own copy of the same notify_cb-less
short-circuit — never got the same treatment. It doesn't even accept an
approval_callback parameter. In the same leaked-ask-mode-into-CLI scenario,
execute_code calls still silently drop into pending_approval with the panel
never shown, even though a CLI callback is registered.
Compute is_cli/approval_callback the same way the two fixed functions do,
and when _should_fall_through_to_cli_approval() says yes, run the same
hook-fire -> prompt_dangerous_approval -> hook-fire -> choice-branch
sequence _run_approval_gate's tail already uses, adapted to this function's
own message/persistence conventions (smart-denied session/permanent
suppression, denial-breaker addendum). Falls back to the existing
pending_approval behavior when no CLI callback is available.
Tests: 4 new cases in tests/tools/test_cli_approval_exec_ask_leak.py
mirroring the existing check_all_command_guards pair (approve/deny/timeout/
session-persistence). Mutation-verified: all 4 fail against the pre-fix code
and pass with it restored.
Neighbor suites: tests/tools/*approval* (291+ tests) and
tests/gateway/{test_approval_prompt_redaction,test_tui_approval_redaction,
test_plaintext_approval_routing,test_discord_exec_approval_content} +
tests/cli/test_cli_approval_ui.py all green. The 7 test_approval_mode_parity
/ test_nonrecursive_verification_artifact_cleanup failures seen in one full
batch run are pre-existing and independent of this change — confirmed by
re-running the identical batch with tools/approval.py stashed back to
pre-fix: the same 7 fail for the same reasons either way.
Note on an adjacent open PR: #65592 also touches check_execute_code_guard,
but an earlier, unrelated region of the function (adding an AST dangerous-
operation scanner to the "not is_gateway and not is_ask" auto-approve
branch). No semantic overlap with this fix's notify_cb-less branch; a small
rebase may be needed depending on merge order.
ceabb030f added the Grok Imagine Image 2.0 catalog entry with upscale=True,
violating the Aug 2026 opt-in-only upscaling policy (f06c41522) and breaking
test_upscale_defaults_are_all_off on main, which reddened every PR's slice
12/12.
Tool Search treated desktop_ui and project as plugin catalog, so
read_window_below dropped out of the model-facing array and the HUD
note never fired. Session-gated GUI tools stay off the core list but
remain direct once a session enables them.
Co-authored-by: fangliquan <fangliquan@qq.com>