Router shipped a minimal /v1/chat/completions compatibility surface
(translated onto Responses) after this PR was written, so the
'does not exist and 404s' wording is stale. Responses remains the
native wire — per-model reasoning-effort validation, reasoning
summaries, and prompt caching live there — so the api.router.com
host mandate is unchanged; only the comments and docs are updated.
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:
- plugins/model-providers/router/: RouterProfile plugin —
api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
runtime_provider._detect_api_mode_for_url: api.router.com ->
codex_responses. The host is Responses-only — POST /v1/chat/completions
does not exist and 404s — so this is a genuine host mandate (exact
hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
hook (tri-state: None=defer, ()=model takes no reasoning params,
tuple=clamp set). Router validates reasoning.effort per model and
returns HTTP 400 invalid-argument on levels outside the model's
published vocabulary, and 400 unsupported_parameter when a
non-reasoning model receives any reasoning field (both verified live).
The profile answers from a cached copy of the catalog's
router.capabilities.reasoning block: cache-only on the hot path,
seeded for free by fetch_models(), disk-mirrored across processes
(/cache/router_catalog.json), background-warmed when cold
— same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
the generic effort-clamp branch (xai/actual/github branches untouched;
profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
rejection, profile registration + auth auto-registry wiring, catalog
parsing, and transport clamp/suppression/fallback paths.
Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.