Files
hermes-agent/providers
Teknium 2a95791992 simplify(compat): run_agent/model_tools/toolsets/acp/providers — drop 42 re-exports/aliases, repoint 15 callers + 99 test files
run_agent.py: delete the `# noqa: F401` re-export block (agent.process_bootstrap
OpenAI/_SafeWriter/_get_proxy_*, model_tools get_tool_definitions/
handle_function_call/check_toolset_requirements, FailoverReason,
_qwen_portal_headers/_routermint_headers, session_persistence names,
estimate_request_tokens_rough, ContextCompressor + friends, jittered_backoff,
prompt_builder names, message_sanitization names, tool_dispatch_helpers
names) — 41 names run_agent never used itself — and the `_STREAM_DIAG_HEADERS`
back-compat class alias (no in-tree reader). run_agent now imports only what
it uses (get_toolset_for_tool, is_local_endpoint, coalesce/uniquify tool-call
ids, cleanup_vm/get_active_env from terminal_tool_lifecycle).

agent/*: `_ra().X` late-binds that only reached a re-export now import the
defining module directly (agent_runtime_helpers -> process_bootstrap.OpenAI,
model_tools.handle_function_call, session_persistence._safe_session_filename_component;
agent_init -> model_tools.get_tool_definitions/check_toolset_requirements,
_lazy_headers("agent.client_lifecycle", ...) for qwen/routermint;
system_prompt -> agent.prompt_builder / model_tools directly, dropping its
own _ra() shim and the `_r` parameter threading). `_ra()` stays for
run_agent-resident names (logger, AIAgent, _hermes_home, _set_interrupt, ...).

toolsets.py: remove resolve_multiple_toolsets (shim-only, restored by
34abf954bd); tests/test_toolsets.py pins the same union behavior via
resolve_toolset over each name.

providers/__init__.py: drop the OMIT_TEMPERATURE re-export (no callers via the
package); ProviderProfile stays because __init__ uses it for annotations —
2 tests repointed to providers.base.

agent/iteration_budget.py: drop the "run_agent re-exports the class"
docstring pointer; 4 tests import IterationBudget from its home.

model_tools.py (arg_coercion names), agent/tool_executor.py, and
hermes_cli/cli_session_mixin.py repoints landed via a sibling commit on this
shared worktree.

Callers repointed: gateway/run.py, hermes_cli/cli_chat_turn_mixin.py,
hermes_cli/cli_tui_mixin.py, tui_gateway/session_workdir.py,
agent/transports/codex.py (one-line imports) + comment pointers in
tools/file_state.py, tools/schema_sanitizer.py, scripts/tool_search_livetest.py.
Tests: patch("run_agent.X") / monkeypatch.setattr(run_agent, "X") /
`from run_agent import X` -> defining module across 99 test files.
2026-09-03 13:28:22 -07:00
..

providers/

Registry and ABC for every inference provider Hermes knows about.

Each provider is declared once as a ProviderProfile. Every other layer — auth resolution, transport kwargs, model listing, runtime routing — reads from these profiles instead of maintaining its own parallel data.


Layout

providers/
├── base.py         ProviderProfile dataclass + OMIT_TEMPERATURE sentinel
├── __init__.py     Registry: register_provider(), get_provider_profile(), list_providers()
└── README.md       This file

The profiles themselves live as plugins under plugins/model-providers/<name>/ (bundled in this repo) and $HERMES_HOME/plugins/model-providers/<name>/ (per-user overrides). The registry in providers/__init__.py lazily discovers them the first time any consumer calls get_provider_profile() or list_providers(). See plugins/model-providers/README.md for the plugin contract and examples.


How it wires in

The registry is populated on first access. After that, every downstream layer reads from it:

  • hermes_cli/auth.py extends PROVIDER_REGISTRY with every api-key profile it sees (skipping copilot, kimi-coding, kimi-coding-cn, zai, openrouter, custom — those need bespoke token resolution).
  • hermes_cli/models.py extends CANONICAL_PROVIDERS and calls profile.fetch_models() inside provider_model_ids().
  • hermes_cli/doctor.py adds a /models health check for each auth_type="api_key" profile.
  • hermes_cli/config.py injects every env_var into OPTIONAL_ENV_VARS so the setup wizard knows about it.
  • hermes_cli/runtime_provider.py reads profile.api_mode as a fallback when URL detection finds nothing.
  • agent/model_metadata.py maps hostname → provider via profile.get_hostname().
  • agent/auxiliary_client.py reads profile.default_aux_model first before falling back to the legacy hardcoded dict.
  • agent/transports/chat_completions.py::_build_kwargs_from_profile() invokes profile.prepare_messages(), profile.build_extra_body(), and profile.build_api_kwargs_extras() on every call.
  • run_agent.py passes provider_profile=<ProviderProfile> so the transport takes the profile path instead of the legacy flag path.

Adding a provider

See plugins/model-providers/README.md — drop a new directory there (or under $HERMES_HOME/plugins/model-providers/ for a private plugin).


Hooks you can override on ProviderProfile

Hook Purpose
get_hostname() URL-based detection — default derives from base_url.
prepare_messages(msgs) Provider-specific message preprocessing (Qwen normalises to list-of-parts, injects cache_control).
build_extra_body(**ctx) Provider-specific extra_body (OpenRouter provider prefs, Gemini thinking_config).
build_api_kwargs_extras(**ctx) (extra_body_additions, top_level_kwargs) — Kimi puts reasoning_effort top-level, Qwen splits enable_thinking/thinking_budget.
supported_reasoning_efforts(model) Declared per-model reasoning-effort vocabulary for gateways that 400 on unknown levels (Ramp Router reads its live catalog). None = defer to transport defaults, () = model takes no reasoning params, tuple = clamp target. Must be cache-only — called on the request hot path.
fetch_models(*, api_key) Live catalog fetch — default hits {models_url or base_url}/models with Bearer auth. Override for no-REST providers (Bedrock), OAuth catalogs (Anthropic), or public catalogs (OpenRouter).

Configuration fields

Full reference in providers/base.py dataclass definition.