Key the post-ladder AuthError on the literal `custom` request again, in
addition to the dead runtime shape (provider=custom, empty api_key). The
previous commit widened it to every alias that resolves to custom (ollama,
vllm), which broke `/model <direct-alias>` switching: _creds_for_switched_provider
resolves the alias provider tolerantly and _apply_direct_alias_endpoint
supplies the alias endpoint AFTER that call, so the early raise turned a
working switch into "ollama is not connected"
(tests/hermes_cli/test_models.py::TestLocalOllamaModelDiscovery, red in CI).
The issue's own scope (#111741) is the bare non-routable placeholder; a
credential-less local alias is a legitimate intermediate state there.
Tighten the post-ladder guard from #111744 to the exact failing shape: a
resolved ``custom`` runtime with an EMPTY api_key. Every other custom rung
(named entry, direct alias, local bypass, pool, key_cmd) yields a key, a
callable or the ``no-key-required`` placeholder, so the loopback heuristic and
the has_usable_secret() re-check were dead branches — and keying on the
requested name alone missed aliases that resolve to custom (``ollama``,
``vllm`` with nothing configured), which still returned the credential-less
OpenRouter fallback and died at agent construction as "No LLM provider
configured". The AuthError now names the requested provider, so cron's
runner/preflight, the CLI, the gateway and the TUI gateway (all of which
format AuthError or walk their fallback chain on it) show the culprit.
Tests folded to two invariants: raise-and-name (bare + alias, plus the
OpenRouter-key control from #111750) and the loopback no-auth control.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
hermes_cli/runtime_provider._getenv was a 4-line copy of get_secret(name,
default) or default; it becomes agent.secret_scope.get_secret_str (returns
default only when the secret is genuinely unset, still raises
UnscopedSecretError — a child's unscoped read is a spawn-site bug). The
runtime_provider_backends/_custom siblings call it directly instead of via
the origin module.
tools/tts_tool, tools/transcription_tools and tools/xai_http each carried an
identical get_env_value re-export kept "so tests can patch" it; the seam is
hermes_cli.config.get_env_value, read lazily at call time. Callers
(tts_streaming, tts_tool_providers, transcription_cloud, voice_client_config,
tools_config) go there directly; resolve_provider_secret already defaults to
it so the env_getter kwarg is gone. Tests repointed at the canonical; the two
tests that only proved the shim forwarded are deleted.
Behavior change: none.
Two regressions in the mirror support: (1) the canonical
https://openrouter.ai/api/v1 that `hermes setup` persists under
provider: openrouter was treated as a custom endpoint, dropping the
auth.json credential pool and returning an empty API key; (2) an
unrelated CUSTOM_BASE_URL (which outranks the config mirror) still
received OPENROUTER_API_KEY because key selection tested mirror
eligibility, not the endpoint actually selected. A config URL is a
mirror only when its host is not openrouter.ai, and the mirror key
branch fires only when base_url is the config URL.
Found by independent review before merge.
When config.yaml sets `model.provider: openrouter` together with a
`model.base_url` mirror/proxy, an explicit `--provider openrouter`
request ignored the mirror and sent traffic to the public OpenRouter
endpoint: the config base_url was only trusted for auto/custom, the
credential pool was still consulted (so a pooled key won over the
mirror), and even when the mirror URL was used its host failed the
openrouter.ai match so OPENROUTER_API_KEY was not selected for it.
Trust the config base_url for the explicit openrouter case, treat that
mirror as an OpenRouter context for key selection, and bypass the pool
like the other custom-endpoint cases already do.
Fixes#10622
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
An external-process provider is an agent CLI Hermes drives over stdio rather
than an HTTP endpoint. Three things about it were spelled out for one vendor,
and each was a hard stop for any other:
* ``resolve_provider()`` gates on ``PROVIDER_REGISTRY``. Its auto-extend from
``providers/`` covered api-key providers only, so an external-process profile
never entered it and ``hermes -m <that provider>`` died with "Unknown
provider" before a client was ever built.
* ``resolve_runtime_provider()`` keyed the external-process branch on the
literal ``"copilot-acp"``, so anything else silently fell through to the
OpenRouter default instead of its own runtime.
* ``resolve_external_process_provider_credentials()`` hardcoded the binary
(``copilot``), the argv (``--acp --stdio``), the env var names and the
placeholder api_key — so a third-party provider would have been handed
another vendor's CLI.
Now the profile carries what only the provider knows — ``process_command``,
``process_args``, ``process_command_env_vars``, ``process_args_env_var`` — and
the three core paths key on ``auth_type == "external_process"`` instead of a
name. copilot-acp's values move into its profile verbatim, so
``HERMES_COPILOT_ACP_COMMAND`` / ``COPILOT_CLI_PATH`` /
``HERMES_COPILOT_ACP_ARGS`` and its ``copilot-acp`` api_key placeholder behave
exactly as before; the new tests assert that alongside the out-of-tree case at
every step.
The error for a missing binary now names the provider and its own env override
instead of telling every user to install GitHub Copilot CLI.
Co-Authored-By: Junie <junie@jetbrains.com>
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
Two callers of get_custom_provider_pool_key compared against its single
preferred key and broke when the pool held the other identity:
- _prune_replaced_custom_model_config_credentials skipped only the
preferred key, so a keyed provider's own legacy-named pool
(custom:b.ai) was false-pruned of its current model_config credential
when the preferred key resolved to the bare slug (b-ai).
- _seed_custom_pool seeded only when the pool key equaled the preferred
key, so a legacy-named pool stopped being seeded from model.api_key.
Both now compare against the full custom_provider_pool_key_candidates
set. Also drops a redundant get_custom_provider_pool_key call from
_try_resolve_from_custom_pool (it returned candidates[0], doubling the
config traversal) and updates the two test files that monkeypatched the
removed module attribute.
Follow-up to #100413.
hermes auth add stores providers.<key> credentials under the config
slug, but runtime only tried custom:<display-name> and then sent the
no-key-required placeholder. Try the slug first, keep the legacy
namespace as fallback, and thread provider_key/key_env through named
custom resolution.
Salvaged from PR #84199 by @RickyYii. DirectAlias gains api_key/key_env; the direct-alias override re-resolves credentials against the alias endpoint (host-gated, #28660) and reuses the pre-alias key only on an origin match; oneshot -m <alias> passes the alias key as explicit_api_key; direct-alias branch gains the OLLAMA_API_KEY host gate. Fixes#83612.
Forward normalized custom-provider capabilities on the default gateway path so native compaction does not depend on session rehydration. Document the content trust boundary and cover both lookup and gateway resolution.
Router shipped a minimal /v1/chat/completions compatibility surface
(translated onto Responses) after this PR was written, so the
'does not exist and 404s' wording is stale. Responses remains the
native wire — per-model reasoning-effort validation, reasoning
summaries, and prompt caching live there — so the api.router.com
host mandate is unchanged; only the comments and docs are updated.
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:
- plugins/model-providers/router/: RouterProfile plugin —
api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
runtime_provider._detect_api_mode_for_url: api.router.com ->
codex_responses. The host is Responses-only — POST /v1/chat/completions
does not exist and 404s — so this is a genuine host mandate (exact
hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
hook (tri-state: None=defer, ()=model takes no reasoning params,
tuple=clamp set). Router validates reasoning.effort per model and
returns HTTP 400 invalid-argument on levels outside the model's
published vocabulary, and 400 unsupported_parameter when a
non-reasoning model receives any reasoning field (both verified live).
The profile answers from a cached copy of the catalog's
router.capabilities.reasoning block: cache-only on the hot path,
seeded for free by fetch_models(), disk-mirrored across processes
(/cache/router_catalog.json), background-warmed when cold
— same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
the generic effort-clamp branch (xai/actual/github branches untouched;
profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
rejection, profile registration + auth auto-registry wiring, catalog
parsing, and transport clamp/suppression/fallback paths.
Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.
A session row persists the provider identity a chat actually used. When that
provider is later renamed or removed (e.g. a custom_providers:/providers:
entry deleted, or a provider renamed oldone->newone), Desktop/TUI resume
restores the stale name into agent init and dies with:
agent init failed: Unknown provider '<name>'
while the CLI resumes the same session fine with the configured default.
- runtime_provider: add is_routable_provider() (full resolution chain:
built-in -> providers: -> custom_providers: -> models.dev)
- _stored_session_runtime_overrides: heal a non-routable provider via
canonical_custom_identity (base_url -> model -> configured provider),
drop to the configured default when unrecoverable, and clear the stale
base_url after healing so a dead endpoint cannot override the registry URL
- _start_agent_build: gate deferred-resume overrides on provider routability;
when the stored provider is gone, prefer the model the user picked for THIS
session, else the configured default
- tests: is_routable_provider cases, heal/fallback round-trips, gate checks
Refs #75128
Test-pollution class: runtime_provider is usually imported lazily (inside
switch_model's resolution path), so its first import in a pytest worker can
happen while a test has hermes_cli.config.load_config patched. The
module-level from-import then bound the MagicMock permanently — after the
patch exited, every later caller in the process silently read the dead
test's config. Live victim: MoA aggregator context-length resolution
(resolve_runtime_provider -> AuthError 'Unknown provider custom:example'),
making TestMoAContextLength::test_moa_custom_context_configures_compressor_threshold
fail whenever it shared a process with
TestLocalOllamaModelDiscovery::test_switch_model_on_current_ollama_custom_endpoint_keeps_base_url.
Fix: load_config / get_compatible_custom_providers / normalize_extra_headers
become late-bound delegates resolving hermes_cli.config attributes at call
time. Both patch targets (config.load_config and
runtime_provider.load_config) keep working. Regression tests pin the
late-binding property and fail if the delegates revert to from-imports
(sabotage-verified).
resolve_runtime_provider() documents target_model as the explicit model
override for mid-session switches and auxiliary slots, but the custom
provider path (_resolve_named_custom_runtime) never received it and
silently substituted the provider's configured default_model instead.
This made auxiliary slots such as auxiliary.background_review silently
run the provider's default model rather than the configured one — e.g. an
ocx-proxy slot configured for gemini-flash actually executed
cursor/claude-sonnet-5, hitting upstream rate limits.
Pass target_model through to the custom runtime resolver and prefer it
over the provider's default model in both the pooled and non-pooled
credential paths.
Follow-up structural pass on the review fix:
- Runtime provider, auxiliary resolution, model validation
(hermes_cli/models.py), live discovery (bedrock_model_ids_or_none),
and the Mantle URL/SigV4 fallbacks all resolve their region through
resolve_bedrock_runtime_region() — one canonical implementation of the
config-first priority instead of three hand-rolled copies.
- agent_init: drop the 'if "client_kwargs" in locals()' guard by
initializing client_kwargs unconditionally at the top of the else
branch; the Mantle kwargs hook is a documented no-op for non-Mantle
base URLs.
Address review feedback on #65076:
- Add resolve_bedrock_runtime_region() to agent/bedrock_adapter.py: the
config-first region resolution (bedrock.region in config.yaml, then
AWS_REGION/AWS_DEFAULT_REGION/botocore profile/us-east-1) that the main
runtime resolver uses, exposed as a shared helper.
- Switch auxiliary client resolution (agent/auxiliary_client.py aws_sdk
branch) to the new helper. Previously it derived its region with bare
resolve_bedrock_region() (env-first), so when config.yaml pinned
bedrock.region to a different region than the ambient AWS env, auxiliary
calls (compression, memory, vision) left the primary runtime's region.
Both the AnthropicBedrock/Converse path and the new Mantle OpenAI
Responses path now resolve identically to the main runtime.
- Add regression tests covering the bedrock.region-vs-AWS_REGION mismatch
for both the Claude auxiliary path and the Mantle auxiliary path.
- Update website/docs/guides/aws-bedrock.md: the guide claimed Hermes never
uses the OpenAI-compatible endpoint, which the Mantle route made stale.
Document the triple routing (AnthropicBedrock / Mantle OpenAI Responses /
Converse), the Mantle auth model (bearer token or SigV4), and add the
GPT-5.5/5.6 model IDs to the models table.
Route Bedrock-hosted OpenAI GPT-5.5 through the Bedrock Mantle OpenAI Responses endpoint with SigV4 request signing. Keep native Bedrock Converse and Claude Bedrock routing unchanged, and add picker/runtime regression coverage.