cf60ebbdfd264c3fc067b9e3230df8752c50d2ec
3 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
804f8b4732 |
feat(providers): add Ramp Router (router.com) provider plugin
Ramp Router is an OpenAI Responses-compatible LLM gateway at https://api.router.com/v1 that routes each request across upstream providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side fallbacks and spend controls. Nous asked for a PR adding it as a provider, so: - plugins/model-providers/router/: RouterProfile plugin — api_mode=codex_responses, RAMP_ROUTER_API_KEY auth, RAMP_ROUTER_BASE_URL override, live account-scoped catalog via GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and Router's docs mandate runtime catalog reads). - hermes_cli/providers.host_mandated_api_mode + runtime_provider._detect_api_mode_for_url: api.router.com -> codex_responses. The host is Responses-only — POST /v1/chat/completions does not exist and 404s — so this is a genuine host mandate (exact hostname match per #32243, mirroring the api.meta.ai precedent). - providers/base.py: new overrideable supported_reasoning_efforts(model) hook (tri-state: None=defer, ()=model takes no reasoning params, tuple=clamp set). Router validates reasoning.effort per model and returns HTTP 400 invalid-argument on levels outside the model's published vocabulary, and 400 unsupported_parameter when a non-reasoning model receives any reasoning field (both verified live). The profile answers from a cached copy of the catalog's router.capabilities.reasoning block: cache-only on the hot path, seeded for free by fetch_models(), disk-mirrored across processes (/cache/router_catalog.json), background-warmed when cold — same design as the OpenRouter reasoning-caps clamp on the chat path. - agent/transports/codex.py: consult the profile-declared vocabulary in the generic effort-clamp branch (xai/actual/github branches untouched; profiles that do not override the hook see no behavior change). - cli-config.yaml.example + adding-providers.md + providers/README.md: document the provider, the host mandate, and the new hook. - tests: behavior contracts for the host mandate/URL detection/spoof rejection, profile registration + auth auto-registry wiring, catalog parsing, and transport clamp/suppression/fallback paths. Verified live against api.router.com (Aug 2026): one-shot chat, streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning replay on OpenAI-served models, function_call_output follow-up turns on OpenAI- and Fireworks-served models; store:false / prompt_cache_key / include:[reasoning.encrypted_content] / reasoning.summary accepted across backends; effort clamp confirmed to convert a would-be 400 (xhigh on o3) into a successful request via the disk mirror. |
||
|
|
9022804d78 |
feat(providers): make all 33 providers pluggable under plugins/model-providers/
Every provider profile is now a self-contained plugin under plugins/model-providers/<name>/, mirroring the plugins/platforms/ pattern established for IRC and Teams. The ProviderProfile ABC stays in providers/; the per-provider profile data moves out. - plugins/model-providers/<name>/__init__.py calls register_provider() - plugins/model-providers/<name>/plugin.yaml declares kind: model-provider - providers/__init__.py._discover_providers() lazily scans bundled plugins then $HERMES_HOME/plugins/model-providers/<name>/ (user override path) - User plugins with the same name override bundled ones (last-writer-wins in register_provider) - Legacy providers/<name>.py layout still supported for back-compat with out-of-tree editable installs - Hermes PluginManager: new kind=model-provider; skipped like memory plugins (providers/ discovery owns them); standalone plugins with register_provider+ProviderProfile in their __init__.py auto-coerce to this kind (same heuristic as memory providers) - skip_names extended to include 'model-providers' so the general PluginManager doesn't double-scan the category - 4 new tests in tests/providers/test_plugin_discovery.py covering bundled discovery, user override, and general-loader isolation - Docs updated: website/docs/developer-guide/adding-providers.md, provider-runtime.md, providers/README.md, plugins/model-providers/README.md No API break: auth.py / config.py / doctor.py / models.py / runtime_provider.py / model_metadata.py / auxiliary_client.py / chat_completions.py / run_agent.py all still consume providers via get_provider_profile() / list_providers() — they just now see plugin-discovered entries instead of pkgutil-iterated ones. Third parties can now drop a single directory into ~/.hermes/plugins/model-providers/<name>/ to add or override an inference provider without touching the repo. |
||
|
|
20a4f79ed1 |
feat: provider modules — ProviderProfile ABC, 33 providers, fetch_models, transport single-path
Introduces providers/ package — single source of truth for every inference provider. Adding a simple api-key provider now requires one providers/<name>.py file with zero edits anywhere else. What this PR ships: - providers/ package (ProviderProfile ABC + 33 profiles across 4 api_modes) - ProviderProfile declarative fields: name, api_mode, aliases, display_name, env_vars, base_url, models_url, auth_type, fallback_models, hostname, default_headers, fixed_temperature, default_max_tokens, default_aux_model - 4 overridable hooks: prepare_messages, build_extra_body, build_api_kwargs_extras, fetch_models - chat_completions.build_kwargs: profile path via _build_kwargs_from_profile, legacy flag path retained for lmstudio/tencent-tokenhub (which have session-aware reasoning probing that doesn't map cleanly to hooks yet) - run_agent.py: profile path for all registered providers; legacy path variable scoping fixed (all flags defined before branching) - Auto-wires: auth.PROVIDER_REGISTRY, models.CANONICAL_PROVIDERS, doctor health checks, config.OPTIONAL_ENV_VARS, model_metadata._URL_TO_PROVIDER - GeminiProfile: thinking_config translation (native + openai-compat nested) - New tests/providers/ (79 tests covering profile declarations, transport parity, hook overrides, e2e kwargs assembly) Deltas vs original PR (salvaged onto current main): - Added profiles: alibaba-coding-plan, azure-foundry, minimax-oauth (were added to main since original PR) - Skipped profiles: lmstudio, tencent-tokenhub stay on legacy path (their reasoning_effort probing has no clean hook equivalent yet) - Removed lmstudio alias from custom profile (it's a separate provider now) - Skipped openrouter/custom from PROVIDER_REGISTRY auto-extension (resolve_provider special-cases them; adding breaks runtime resolution) - runtime_provider: profile.api_mode only as fallback when URL detection finds nothing (was breaking minimax /v1 override) - Preserved main's legacy-path improvements: deepseek reasoning_content preserve, gemini Gemma skip, OpenRouter response caching, Anthropic 1M beta recovery, etc. - Kept agent/copilot_acp_client.py in place (rejected PR's relocation — main has 7 fixes landed since; relocation would revert them) - _API_KEY_PROVIDER_AUX_MODELS alias kept for backward compat with existing test imports Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com> Closes #14418 |