docs: chat-completions is now a compat shim on Router, not a 404
Router shipped a minimal /v1/chat/completions compatibility surface (translated onto Responses) after this PR was written, so the 'does not exist and 404s' wording is stale. Responses remains the native wire — per-model reasoning-effort validation, reasoning summaries, and prompt caching live there — so the api.router.com host mandate is unchanged; only the comments and docs are updated.
This commit is contained in:
@@ -152,10 +152,11 @@ model:
|
|||||||
# router:
|
# router:
|
||||||
# base_url: https://api.router.com/v1
|
# base_url: https://api.router.com/v1
|
||||||
# api_key: ${RAMP_ROUTER_API_KEY}
|
# api_key: ${RAMP_ROUTER_API_KEY}
|
||||||
# # api_mode auto-detected as codex_responses for api.router.com — the host
|
# # api_mode auto-detected as codex_responses for api.router.com — the
|
||||||
# # is Responses-only (/v1/chat/completions 404s). The bundled router
|
# # Responses API is Router's native wire (chat/completions is only a
|
||||||
# # provider covers this; a named custom provider is only needed for a
|
# # compatibility shim). The bundled router provider covers this; a named
|
||||||
# # non-default Router-compatible endpoint.
|
# # custom provider is only needed for a non-default Router-compatible
|
||||||
|
# # endpoint.
|
||||||
|
|
||||||
|
|
||||||
# Command-minted credentials (optional): key_cmd
|
# Command-minted credentials (optional): key_cmd
|
||||||
|
|||||||
@@ -663,8 +663,10 @@ def host_mandated_api_mode(base_url: str = "") -> Optional[str]:
|
|||||||
- api.meta.ai only achieves KV-cache hits on /v1/responses with
|
- api.meta.ai only achieves KV-cache hits on /v1/responses with
|
||||||
prompt_cache_retention; /v1/chat/completions returns 0 cached
|
prompt_cache_retention; /v1/chat/completions returns 0 cached
|
||||||
tokens (measured 0% vs 93-99% on /responses with retention).
|
tokens (measured 0% vs 93-99% on /responses with retention).
|
||||||
- api.router.com (Ramp Router) implements ONLY the Responses API —
|
- api.router.com (Ramp Router) is Responses-native: per-model
|
||||||
POST /v1/chat/completions does not exist on the host and 404s.
|
reasoning-effort validation, reasoning summaries, and prompt
|
||||||
|
caching live on /v1/responses; /v1/chat/completions is only a
|
||||||
|
minimal compatibility shim translated onto it.
|
||||||
- api.anthropic.com / ``…/anthropic`` suffixes speak native Messages.
|
- api.anthropic.com / ``…/anthropic`` suffixes speak native Messages.
|
||||||
- Kimi's ``/coding`` endpoint speaks native Messages.
|
- Kimi's ``/coding`` endpoint speaks native Messages.
|
||||||
- AWS Bedrock runtime hosts speak Converse.
|
- AWS Bedrock runtime hosts speak Converse.
|
||||||
@@ -697,9 +699,11 @@ def host_mandated_api_mode(base_url: str = "") -> Optional[str]:
|
|||||||
# cache-cold (0% vs 93-99% measured). Exact-hostname match per #32243.
|
# cache-cold (0% vs 93-99% measured). Exact-hostname match per #32243.
|
||||||
if hostname == "api.meta.ai":
|
if hostname == "api.meta.ai":
|
||||||
return "codex_responses"
|
return "codex_responses"
|
||||||
# Ramp Router (api.router.com) is Responses-only: the host serves
|
# Ramp Router (api.router.com) is Responses-native: reasoning-effort
|
||||||
# GET /v1/models and POST /v1/responses, and /v1/chat/completions 404s
|
# validation, reasoning summaries, and prompt caching live on
|
||||||
# (docs.router.com/api/endpoint). Exact-hostname match per #32243.
|
# /v1/responses, and /v1/chat/completions is only a minimal
|
||||||
|
# compatibility shim (docs.router.com/api/endpoint). Exact-hostname
|
||||||
|
# match per #32243.
|
||||||
if hostname == "api.router.com":
|
if hostname == "api.router.com":
|
||||||
return "codex_responses"
|
return "codex_responses"
|
||||||
if hostname.startswith("bedrock-runtime.") and base_url_host_matches(base_url, "amazonaws.com"):
|
if hostname.startswith("bedrock-runtime.") and base_url_host_matches(base_url, "amazonaws.com"):
|
||||||
|
|||||||
@@ -163,8 +163,9 @@ def _detect_api_mode_for_url(base_url: str) -> Optional[str]:
|
|||||||
return "codex_responses"
|
return "codex_responses"
|
||||||
if hostname == "api.actual.inc":
|
if hostname == "api.actual.inc":
|
||||||
return "codex_responses"
|
return "codex_responses"
|
||||||
# Ramp Router: Responses-only host — /v1/chat/completions does not
|
# Ramp Router: Responses-native host — /v1/chat/completions is only a
|
||||||
# exist and 404s (docs.router.com/api/endpoint). Mirrors the
|
# minimal compatibility shim, while reasoning and caching support live
|
||||||
|
# on /v1/responses (docs.router.com/api/endpoint). Mirrors the
|
||||||
# host_mandated_api_mode clause in hermes_cli/providers.py so the
|
# host_mandated_api_mode clause in hermes_cli/providers.py so the
|
||||||
# runtime resolver stays in lockstep. Exact hostname per #32243.
|
# runtime resolver stays in lockstep. Exact hostname per #32243.
|
||||||
if hostname == "api.router.com":
|
if hostname == "api.router.com":
|
||||||
|
|||||||
@@ -8,10 +8,14 @@ spend controls server-side.
|
|||||||
|
|
||||||
Wire notes (verified live against api.router.com, Aug 2026):
|
Wire notes (verified live against api.router.com, Aug 2026):
|
||||||
|
|
||||||
* **Responses API only.** Router implements ``GET /v1/models`` and
|
* **Responses API is the native wire.** Router serves ``GET /v1/models``
|
||||||
``POST /v1/responses``; ``POST /v1/chat/completions`` does not exist and
|
and ``POST /v1/responses``; ``POST /v1/chat/completions`` is only a
|
||||||
404s. ``api_mode="codex_responses"`` plus the ``api.router.com`` host
|
minimal compatibility shim (added Aug 2026) that translates onto
|
||||||
mandate in ``hermes_cli/providers.py`` keep every path off the chat wire.
|
Responses. Per-model reasoning-effort validation, reasoning summaries,
|
||||||
|
and prompt caching are Responses-surface features, so
|
||||||
|
``api_mode="codex_responses"`` plus the ``api.router.com`` host mandate
|
||||||
|
in ``hermes_cli/providers.py`` keep every path on the native wire —
|
||||||
|
the same shape as the ``api.openai.com`` mandate.
|
||||||
* **Account-scoped catalog.** Valid model IDs are whatever the key's
|
* **Account-scoped catalog.** Valid model IDs are whatever the key's
|
||||||
``GET /v1/models`` returns (BYOK accounts see extra entries), so this
|
``GET /v1/models`` returns (BYOK accounts see extra entries), so this
|
||||||
profile ships **no** ``fallback_models`` — the picker relies on the live
|
profile ships **no** ``fallback_models`` — the picker relies on the live
|
||||||
|
|||||||
@@ -1,7 +1,8 @@
|
|||||||
"""Behavior contracts for the Ramp Router (api.router.com) provider.
|
"""Behavior contracts for the Ramp Router (api.router.com) provider.
|
||||||
|
|
||||||
Router is Responses-only: the host implements GET /v1/models and
|
Router is Responses-native: the host implements GET /v1/models and
|
||||||
POST /v1/responses, and POST /v1/chat/completions does not exist (404).
|
POST /v1/responses, and /v1/chat/completions is only a minimal
|
||||||
|
compatibility shim translated onto Responses.
|
||||||
These tests pin the host mandate, the runtime URL detection that mirrors
|
These tests pin the host mandate, the runtime URL detection that mirrors
|
||||||
it, and the profile/auth registry wiring — same contract suite shape as
|
it, and the profile/auth registry wiring — same contract suite shape as
|
||||||
tests/hermes_cli/test_meta_prompt_cache.py.
|
tests/hermes_cli/test_meta_prompt_cache.py.
|
||||||
|
|||||||
@@ -35,7 +35,7 @@ The important abstraction is `api_mode`.
|
|||||||
|
|
||||||
- Most providers use `chat_completions`.
|
- Most providers use `chat_completions`.
|
||||||
- Codex and Meta Model API (`api.meta.ai` — Muse Spark) use `codex_responses` (auto-sends `prompt_cache_retention: 24h` for prompt caching; `api.meta.ai` achieves 93–99% cache hits only on `/v1/responses`).
|
- Codex and Meta Model API (`api.meta.ai` — Muse Spark) use `codex_responses` (auto-sends `prompt_cache_retention: 24h` for prompt caching; `api.meta.ai` achieves 93–99% cache hits only on `/v1/responses`).
|
||||||
- Ramp Router (`api.router.com`) also uses `codex_responses` — the host is Responses-only (`/v1/chat/completions` 404s) and validates `reasoning.effort` per model, which the router profile handles by declaring each model's vocabulary from the live catalog (`ProviderProfile.supported_reasoning_efforts`).
|
- Ramp Router (`api.router.com`) also uses `codex_responses` — Responses is Router's native wire (`/v1/chat/completions` is only a minimal compatibility shim), and it validates `reasoning.effort` per model, which the router profile handles by declaring each model's vocabulary from the live catalog (`ProviderProfile.supported_reasoning_efforts`).
|
||||||
- Anthropic uses `anthropic_messages`.
|
- Anthropic uses `anthropic_messages`.
|
||||||
- A new non-OpenAI protocol usually means adding a new adapter and a new `api_mode` branch.
|
- A new non-OpenAI protocol usually means adding a new adapter and a new `api_mode` branch.
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user