Commit Graph

187 Commits

Author SHA1 Message Date
Teknium ce54c8e04a refactor(model_switch): collapse duplicate bodies, flatten alias/flag parsing, trim re-exports
- parse_model_flags_detailed: boolean flags via _BOOL_FLAGS table; scope ladder as data
- _load_direct_aliases / _configured_provider_matches / _config_declares_model /
  resolve_alias: single scan or builder per shape instead of parallel branches
- _route_from_model_input split at the MoA/alias seam (-> _route_after_alias)
- re-export block trimmed to the names referenced through hermes_cli.model_switch
  (acp_adapter, cli.py, tui_gateway, inventory, gateway, tests); orphaned trailing
  comment blocks folded into model_switch_providers docstrings
- 2389 -> 2079 lines; largest function _model_sort_key (90 LOC)

Parity: r2m_switch_corpus.py 434 entries byte-identical vs base; 70 model_switch test
files + tests/test_tui_gateway_server.py green.
2026-09-02 16:54:55 -07:00
Teknium 0fd89c0ae5 wip(model_switch): snapshot of in-progress provider-listing extraction + status split (worker died mid-run) 2026-09-02 16:10:32 -07:00
Teknium 3b1ecfc0a1 refactor(cli): decompose list_authenticated_providers and switch_model
model_switch.py (4405 -> 3712):
- list_authenticated_providers (1367 lines) is a thin driver over _PickerBuild
  state and one builder per section (_lap_builtin_rows, _lap_overlay_rows,
  _lap_canonical_rows, _lap_user_provider_rows, _lap_bare_custom_row,
  _lap_custom_provider_rows). Credential checks shared with
  _collect_authed_provider_slugs are single-copy helpers
  (_iter_builtin_candidates, _auth_store_has_provider, _raw_pool_usable,
  _pool_usable, _overlay_has_env_creds, _has_aws_sdk_creds_for_listing);
  live/curated lookup (_live_or_curated_ids, _aws_live_or_curated_ids,
  _nous_picker_model_ids), endpoint discovery (_discover_endpoint_models) and
  config-entry parsing are deduped across sections 1/2/2b/3/3b/4.
- switch_model: _switch_fail, _runtime_creds, _entry_configured_key,
  _ollama_configured_base, _unknown_provider_message, _aggregator_alias_error,
  _aggregator_catalog_match, _config_declares_model,
  _apply_direct_alias_endpoint replace inline blocks; four copies of the
  resolve_runtime_provider unpack collapse to one.
- _model_sort_key: one _flush() for the four repeated version-flush blocks
  (identical on 400k random ids).
Row dicts, ordering, error strings and lazy-import patch points unchanged;
picker rows and the interactive hermes model screen byte-identical vs base.
2026-09-02 13:30:39 -07:00
Solitud1nem 1e8f829f1a fix(model): show copilot-acp in pickers when its executable resolves
The overlay loop in list_authenticated_providers() checks every way a
provider might be authenticated — env keys, the auth store, the
credential pool, even Claude Code's external token files — but never
asks the one question that matters for an external_process provider:
does the executable resolve? copilot-acp has no key or token by design
(the spawned `copilot --acp --stdio` brings its own auth), so has_creds
stayed False and the filter dropped it from every picker. Funny enough,
five lines further down the same loop has a dedicated copilot-acp
branch for fetching its model ids — it just never got a chance to run.

Availability now comes from get_auth_status(), the same source
`hermes model` and the auth status endpoints already use, so the CLI
and GUI agree on what 'configured' means for external-process
providers.

Fixes #63662
2026-09-02 20:51:07 +05:30
Teknium aaa34b0e08 fix(desktop): model picker no longer hardcodes --global; one persist policy for every surface (#90235)
Symptom: picking a model in the Desktop composer for the primary chat
silently rewrote config.yaml (model.default + model.provider) as the
profile default, ignoring model.persist_switch_by_default. A throwaway
pick that resolved to e.g. openai-api (no key) left the profile with an
unusable default on the next launch (#90235).

Root cause: 7d96537bc8 (#86414) made use-model-controls.ts send --global
for every primary-tile pick so a fresh profile would get a persisted
provider instead of falling through to a leftover OPENAI_API_KEY env var.
That put a persistence policy in the client, contradicting the
server-side rule /model uses (resolve_persist_behavior).

Fix:
- resolve_persist_behavior gains one rule, ahead of the --provider
  session-only rule: when neither model.default nor model.provider is
  configured yet, persist. This preserves #86414's first-pick motivation
  for CLI, gateway and Desktop alike. With a default configured, a plain
  pick is session-only unless --global / persist_switch_by_default.
- Desktop primary-tile picks send no scope flag and let the gateway decide.
  Secondary tiles and MoA presets still send --session.
- /model help text in cli.py said "(persists)"; it now matches reality and
  lists --global.
- Docs: desktop.md picker note + slash-commands /model row.

Tests: test_first_pick_persists_then_session_only (fails on main), and the
existing use-model-controls vitest updated to assert the flag-less request.
2026-09-02 05:33:33 -07:00
Teknium 44a57921c5 fix(providers): hide phantom -cn picker rows lit only by shared intl keys; give alibaba-token-plan-cn its own key var
- alibaba-coding-plan-cn / alibaba-token-plan-cn keep the shared intl key vars
  as ordered fallbacks after their dedicated *_CN_API_KEY, so users who set
  ALIBABA_CODING_PLAN_API_KEY / ALIBABA_TOKEN_PLAN_API_KEY for the CN endpoint
  keep working (the PR as filed dropped them).
- list_authenticated_providers hides a '-cn' row whose only lit key vars are
  ones it shares with its non-CN sibling, unless that CN provider is the
  configured model.provider. With only the shared key: one row, not two;
  DASHSCOPE_API_KEY alone: 3 alibaba rows, not 4.
- Docs: environment-variables.md, providers.md.
2026-09-02 05:32:20 -07:00
Pedro Fontana b3576a29c3 Merge pull request #97354 from NousResearch/fix/nous-org-model-policy
fix(nous): honour the org model policy in the model pickers
2026-09-01 18:20:18 -03:00
Teknium 33797073bb fix: harden startup route salvage — aggregator-slug guard, alias credential ownership, oneshot dedup
Follow-ups on top of #87210 (@liuhao1024) and #87246 (@JoaoMarcos44):
- resolve_startup_model_route: aggregator-native slugs stay on the current
  routing aggregator (bare vendor slugs resolve WITHIN the aggregator first);
  URL-bearing aliases resolve via direct_alias_runtime_request so a foreign
  provider label never carries the vendor token to the alias host (#28660);
  route carries the alias's own api_key.
- cli.py: pass current_provider; explicit --api-key wins over alias key.
- Drop #87246's oneshot double-handling (main's oneshot alias+detection path
  already covers it once #87210's detection fix is in) and the PR-body SVG.
- Rewrote/extended startup-route tests for the hardened semantics.
2026-09-01 12:07:35 -07:00
joaomarcos 4c870951e2 fix(cli): resolve startup model routes before provider defaults
Resolve configured aliases and provider/model inputs before HermesCLI attaches the configured default provider. Keep aggregator namespaces intact, cover oneshot startup, and document the supported CLI forms.
2026-09-01 12:07:35 -07:00
liuhao1024 4033f3fc5f fix(models): honor vendor/model prefix and dict model.aliases in provider detection (#87189) 2026-09-01 12:07:35 -07:00
RickyYii 3145986c20 fix(cli): honour model_aliases api_key, stop cross-provider key leak (#83612)
Salvaged from PR #84199 by @RickyYii. DirectAlias gains api_key/key_env; the direct-alias override re-resolves credentials against the alias endpoint (host-gated, #28660) and reuses the pre-alias key only on an origin match; oneshot -m <alias> passes the alias key as explicit_api_key; direct-alias branch gains the OLLAMA_API_KEY host gate. Fixes #83612.
2026-08-31 10:59:45 -07:00
Stephen Chin f245765a6d fix(gateway): preserve capabilities on model switches
Carry normalized provider capabilities through /model results and session overrides so a live model switch does not wait for gateway rehydration.
2026-08-30 05:16:10 -07:00
Stephen Chin 5247a6f07f fix(compaction): clarify runtime capability state
Use a distinct runtime_capabilities field on agents, preserve compatibility with earlier snapshots, and resolve the canonical direct OpenAI endpoint when a cross-provider switch omits base_url. Keep ambiguous proxy routes fail-closed.
2026-08-30 05:16:10 -07:00
Stephen Chin 08c7879ca1 fix(compaction): preserve native capability across runtime switches
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
2026-08-30 05:16:10 -07:00
Jack 5d238be2ca fix(gateway): carry request_overrides through /model session overrides
Follow-up to the previous commit (which fixed the default/fallback
provider path). A mid-session `/model` switch stores a per-session
override bundle in `_session_model_overrides` that omitted
`request_overrides`, and the two consumers
(`_resolve_session_agent_runtime` fast path and
`_apply_session_model_override`) only copied
provider/api_key/base_url/api_mode. So switching *to* a custom provider
via `/model` did not apply its `extra_body`.

- `ModelSwitchResult` gains a `request_overrides` field, derived for the
  switched provider via `_get_named_custom_provider` /
  `_custom_provider_request_overrides` (the same overrides
  `resolve_runtime_provider` surfaces for the default path).
- Both `/model` override-storage sites in slash_commands.py persist it.
- Both consumers apply it; `_apply_session_model_override` also clears a
  stale value when switching to a provider that has none.

Extends tests/gateway/test_turn_request_overrides.py (3 new cases).
2026-08-29 19:13:00 -07:00
CharZhou a9b696c671 fix(model): initialize switch request overrides 2026-08-29 19:13:00 -07:00
CharZhou 863aac9012 fix: preserve named custom provider request_overrides in gateway and /model switches
Carry provider-derived request_overrides through runtime resolution,
fallback projection, session /model state, restart rehydration, and
turn-route merge so named custom providers keep extra_body and related
overrides.
2026-08-29 19:13:00 -07:00
Mariano Nicolini 4d482ed344 refactor(nous): trim comments and drop an unused field 2026-08-28 15:39:23 -03:00
Mariano Nicolini da3c2435e2 fix(nous): only rescue an empty list where emptiness means "filtered out"
The fallback also ran on unavailable_models, which is legitimately empty on a
paid tier, filling the picker with the whole reachable set. Make it opt-in.
2026-08-28 13:34:27 -03:00
Mariano Nicolini 9fc43919cc perf(nous): stop prefetching a catalog nothing reads
The `/model` picker warms `provider_models_cache.json` in parallel before
its serial build loop, and Nous was collected into that prefetch because
the credential scan treats any auth.json providers entry as credentials
regardless of auth type.

Nothing reads the result. The picker's nous branch builds from the curated
list rather than `cached_provider_model_ids`, and Nous cannot reach the
api_key-only unified pathway that would call it. Because the prefetch
forces a refresh it also skips the cache read, so the entry is written and
never read — a live authenticated /v1/models round trip per picker open
for nothing.

Exclude it. Also add the plan this and the preceding commits implement.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:14:38 -03:00
Mariano Nicolini b1ea9196f7 fix(nous): narrow every model list to the org's policy
Four surfaces list Nous models, and none of them was filtered. All four
seed from the docs-hosted curated manifest and union the Portal's
`recommended-models` endpoint; neither source is authenticated, so org
policy had no effect on the model a user picks — which is the model they
then use. The Portal endpoint compounds it, serving one globally
CDN-cached payload for the whole platform, invalidated only by admin
pricing edits and never by a policy change, so it can put a hidden model
straight back into a list.

Narrow all four against the authenticated catalog:

  - `_login_nous`, which chooses the model the session starts on
  - `_model_flow_nous`, the `hermes model` picker
  - `list_authenticated_providers`, the `/model` picker
  - `/api/model/recommended-default`, dashboard onboarding

The list stays curated and curated-ordered — the policy set only ever
subtracts. Replacing a list with the catalog's keys would swap a curated
agentic list for a large alphabetical dump of vendor-prefixed models,
which is the regression the picker's nous branch already exists to avoid.

The `/model` picker's filter sits outside the try that wraps the Portal
union, so a Portal outage still yields a policy-filtered curated list.
`_login_nous` and `_model_flow_nous` also narrow their unavailable lists,
so a policy-hidden model is not offered as a free-tier upsell either.

For an org with no policy — the common case — the filter is a no-op and
every list is what it was.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:13:53 -03:00
Brooklyn Nicholson d83dcb4c3a fix(model): coerce YAML integer provider names before picker/CRUD
PyYAML loads unquoted names like provider: 2070 as int. GET /api/model/options
then called .lower() on the providers dict key and 500ed, and activate/delete
looked up "2070" and missed. Stringify identity fields so Desktop can assign
and remove those endpoints.

Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
2026-08-27 11:31:21 -05:00
Teknium 2a2307e68f feat: keyless providers count as authenticated everywhere — opencode-free appears in /model and desktop pickers with zero setup
A keyless provider has no credential to lack, but every auth-gated
surface treated 'no key' as 'not authenticated', so opencode-free was
invisible in /model, provider:model listing, and the desktop model
pickers unless a user had unrelated OpenCode env vars set.

One policy, three gates, all derived from the HermesOverlay keyless
flag (#91358):
- auth.py get_api_key_provider_status: keyless providers report
  configured/logged_in=True with key_source 'keyless' — flows through
  get_auth_status to every status consumer (hermes status, dashboards,
  list_available_providers).
- model_switch.py list_authenticated_providers: keyless overlay rows
  get has_creds=True before any env/pool/auth-store checks — this is
  the source for /model, the TUI picker, and the desktop
  /api/model/options payload.
- inventory.py explicit-only filter (desktop chat pickers): keyless
  providers are kept — there is nothing to 'explicitly configure', and
  hiding a zero-setup provider defeats its purpose.

E2E (temp HERMES_HOME, all keys stripped): get_auth_status logged_in,
list_available_providers authenticated, picker row with 6 models,
desktop payload default AND explicit_only both include the provider,
and the full switch pipeline (parse free:x-preview-f-free →
switch_model) resolves to the keyless runtime. 4 new tests.
2026-08-21 04:39:31 -07:00
Teknium a83c3915a3 fix(opencode): family-wide provider predicate + reserved tool-name aliases for custom opencode-* providers
Builds on @Lesnak1's #85619 (issue #85589):

- New opencode_provider_family() single-owner predicate in
  hermes_cli/models.py — resolves built-in AND custom family providers
  (opencode-go-bridge, OpenCode-Zen-Custom, ...) case-insensitively.
  Migrated all 8 inlined family checks (models.py x3, runtime_provider.py
  x4 from the salvaged commits) plus 4 sibling sites the PR missed:
  cli.py api_mode sync, agent_runtime_helpers.py double-/v1 guard,
  model_normalize.py flat-namespace strip, model_switch.py base_url
  normalization.
- Responses transport: alias OpenCode-reserved function names
  (web_search, search_files -> hermes_*) on the wire and map them back on
  dispatch — same pattern as the xAI web_search collision fix. Matches
  family providers and any base_url on opencode.ai. Fixes the HTTP 400
  'custom function name X is reserved' half of #85589.
- Tests: custom-provider routing assertions + 5 new transport alias tests.
2026-08-20 20:21:12 -07:00
Teknium 72b7c6c8d1 fix(models): keep discovery sentinels out of the user-facing models mapping
PR #67934 marked auto-discovered catalogs by writing two sentinel keys
INSIDE the user-facing ``models`` mapping of custom provider entries:
``__discovered_model_catalog__`` (written by
_save_discovered_models_to_config) and ``__explicit_model_allowlist__``
(injected by _normalize_custom_provider_entry). Every consumer of that
mapping — pickers, selectors, gateway/agent readers, and the user's own
config.yaml — had to know to filter those keys, and any site that
didn't listed them as phantom model IDs (``__discovered_model_catalog__``
showing up as a selectable "model"). The v11→v12 config migration and
the ACP session-state test caught exactly that leak on main.

Replace the in-mapping sentinels with a single entry-level flag:

- ``models_discovered: true`` now sits next to ``models``/``base_url``
  on the provider entry; the models mapping stays a clean
  ``{model_id: metadata}`` dict with no reserved keys.
- _save_discovered_models_to_config writes the new shape and refreshes
  catalogs it previously discovered (entry-level flag or legacy
  sentinel) instead of treating them as user-curated metadata.
- _normalize_custom_provider_entry no longer injects
  ``__explicit_model_allowlist__``; a dict-shaped models mapping counts
  as an explicit allowlist exactly when the entry is NOT marked
  models_discovered.
- _models_config_is_allowlist takes the discovered flag as a parameter
  (new helper _entry_models_discovered resolves it, including the
  legacy in-mapping sentinel); all call sites updated
  (model_switch.py, model_setup_flows.py, acp_adapter/server.py).
- Backward compat, no config version bump: configs written by a
  pre-fix Hermes (sentinels inside models) still read correctly —
  ``__discovered_model_catalog__: true`` is treated as
  models_discovered, both sentinel keys are stripped from model
  listings, and the next discovery save migrates the entry to the
  clean shape. Covered by a new regression test.

Also restore ``except Exception:`` on the pre-existing guards this PR
had narrowed to specific exception tuples (the resolve_runtime_provider
fallback in switch_model, the picker discovery/cache guards in
list_authenticated_providers, _get_model_config_dict, and
_credential_fingerprint). Those guards were intentionally broad on
main — a failed resolution or probe must degrade to the fallback path,
never crash the model switch. Guards the PR introduced for its own new
probe code keep their authored tuples.

The ACP new_session payload also goes back to
probe_current_custom_provider=False, matching the contract main's
test_new_session_returns_authenticated_cross_provider_model_state pins
(session opens must not block on live-probing the current custom
endpoint).
2026-08-18 14:26:16 -07:00
Vadelma 7fb6b28ec8 fix(models): complete selector parity safeguards
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
Vadelma 81e813507f fix(models): preserve empty catalog and custom model semantics
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
Vadelma 9d2ba6c655 fix(acp): isolate custom catalog identities
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
Vadelma 7e61bd38c6 fix(models): preserve discovery provenance and endpoint identity
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
Vadelma f70e9abce9 fix(models): close final provider discovery edge cases
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
Vadelma fa1bb88e3e feat(models): propagate native discovery across selectors
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
lepetitprince716-prog 7e439dbb1b perf: parallelize provider model-list fetches in model picker
When the 1h provider_models_cache.json TTL lapses, the model picker
serially fetches /v1/models for each authenticated provider. With 10+
providers this stacks to 15-30s of blocking before the picker renders.

Add a parallel prefetch step before the serial picker build loops:
- _collect_authed_provider_slugs(): lightweight credential pre-scan
  that mirrors sections 1/2/2b without fetching model lists
- _prefetch_provider_models_parallel(): ThreadPoolExecutor-based
  concurrent fetch of stale/missing cache entries (max 8 workers)
- update_provider_cache_entry(): thread-safe single-entry cache writer
  with threading.Lock to prevent concurrent write races

Guardrails:
- Skipped when <=3 authed providers (overhead not worth it)
- Skipped when refresh=True (serial path force-refreshes)
- Exception-isolated (falls back to serial path on any failure)
- No behavioral change (same model lists, same picker output)

Closes #80413

(cherry picked from commit 89dddd6cb5d53d73278e0518c375fb5b878e5c6b)
2026-08-15 00:34:29 -07:00
Teknium 4b7b2b0049 fix: widen base-URL hostname identity class to remaining substring sites
Follow-up to #85737, which migrated five provider-identity sites onto
utils.base_url_host_matches()/base_url_hostname(). This completes the class
sweep (never-patch-predicates: one owner, every site) and folds in the two
open contributor PRs attacking individual sites:

- agent/auxiliary_client.py ZAI/Kimi OpenAI-wire rewrite (PR #85715,
  pierrenode): 'bigmodel'/'api.z.ai'/'api.kimi.com' substring checks
  rewrote proxy paths containing those markers.
- hermes_cli/runtime_provider.py Azure endpoint detection (PR #74721,
  RelaxJonh, issue #74312): 'azure.com' substring picked the Azure key
  for non-Azure hosts whose path contained the text.
- run_agent.py: _is_azure_openai_url, _is_copilot_url, Anthropic
  credential-refresh azure guard, _anthropic_preserve_dots host
  allowlist, OpenRouter/mistral reasoning gates.
- agent/chat_completion_helpers.py: nousresearch / nvidia detection.
- agent/conversation_loop.py: GitHub Models 413 hint.
- agent/usage_pricing.py: localhost billing-route detection.
- hermes_cli/model_switch.py: api.openai.com catalog fallback and
  localhost custom-provider detection.
- cli.py: local-model autodetect and Ollama/LM Studio context-length
  hints (port-anchored instead of '11434' in URL).
- tools/mcp_oauth.py: Figma remote-MCP detection.
- tools/skills_hub.py: raw.githubusercontent.com source-URL check.

Regression tests extend tests/hermes_cli/test_base_url_host_identity.py
(azure/copilot/dotted-model/figma proxy-path + lookalike cases) and
tests/agent/test_minimax_auxiliary_url.py (ZAI/Kimi path false positives).

Closes #74312. Salvages #85715 and #74721 with authorship preserved.
2026-08-14 22:04:16 -07:00
kshitij acd8737c10 fix(models): ETag conditional GET, no-network hot-path invariant, mirror URL override for models.dev catalog
Harden the models.dev catalog refresh path (#35838) with three missing
pieces:

1. ETag conditional GET — every network request sends If-None-Match
   with the last-known ETag (persisted alongside the cache file). A 304
   Not Modified re-confirms the existing cache without re-downloading
   the full ~2 MB registry. This makes the 4-hour TTL effectively free
   to maintain.

2. No-network-on-hot-paths invariant — allow_network=False is now the
   default for every query function called on the conversation hot path:
   get_model_capabilities, get_model_info, lookup_models_dev_context,
   _get_provider_models. These are called during vision routing, image
   routing, cost-guard checks, and context-length resolution on every
   turn — they must never block on the network. Interactive flows
   (model picker, model switch) explicitly pass allow_network=True.

3. Mirror URL override — models_dev.url in config.yaml lets deployments
   point at a self-hosted mirror without code changes. Follows the same
   pattern as model_catalog.url.

Additional hardening:
- Cache TTL bumped from 1h to 4h (ETag makes refresh cheap)
- Corrupt/empty disk cache is rejected with a warning instead of being
  served as {} and silently breaking provider/model resolution
- _validate_registry() guards against non-dict and empty-dict payloads

Fixes #35838
2026-08-14 03:31:22 +05:30
Teknium 0569c001d0 fix(model-switch): route switch_model user-provider key reads through the secret scope
Extends the picker fix to the read that actually uses the key: switch_model's
user-provider credential resolution (the ${VAR} api_key expansion and the
key_env fallback at the resolve-credentials step) still read os.environ raw
and passed the result to resolve_runtime_provider as explicit_api_key — so
under multiplex_profiles the actual switch, not just the picker listing,
could adopt another profile's key. Same _scoped_key_env helper, same
fail-closed semantics; identical behavior when multiplexing is off.

Adds end-to-end switch_model tests pinning that an installed scope wins over
the process environment for both read shapes.
2026-08-08 19:17:05 -07:00
Drexuxux 0c97a883af fix(model-switch): read picker key_env through the per-profile secret scope
854007d1c routed the remaining main-agent fallback key reads through
agent.secret_scope so the multiplexed gateway's per-profile scope applies.
list_authenticated_providers - which gateway/slash_commands.py calls
directly for /model - still resolved custom-endpoint and fallback-entry
credentials with raw os.environ.get(key_env), so under multiplex_profiles
one profile's picker reads whatever key the process environment happens to
hold, i.e. another profile's.

  no multiplexing : profileA-key   (unchanged)
  scope installed : profileB-key   (was profileA-key)

Route both reads through a _scoped_key_env() helper over
secret_scope.get_secret(). get_secret is identical to os.getenv when
multiplexing is off, so single-profile deployments are byte-for-byte
unchanged; a fail-closed UnscopedSecretError is treated as "no credential
visible for this profile", which is how the picker already handles a
missing key.

Scope: only the two key_env credential reads. The other environment reads
in that function are provider-presence probes (AWS creds, LM_BASE_URL),
a separate concern.
2026-08-08 19:17:05 -07:00
Teknium b79e83827d fix(model-switch): surface candidates on ambiguous alias instead of guessing
An alias that family-matches multiple catalog models (/model opus) used to
silently pick one via _model_sort_key heuristics. The heuristics have
guessed wrong repeatedly — dated snapshots like claude-opus-4-20250514
parsed as version 20,250,514 and outranked claude-opus-4-8; suffix
tiebreaks landed on the cheapest tier — and every wrong guess silently
switches the user to a model they did not ask for.

resolve_alias now raises AmbiguousAliasError whenever more than one model
matches the alias family; switch_model catches it at all three call sites
(explicit-provider path, current-provider path, authenticated-provider
fallback) and returns a failure result listing the candidates
(best-guess-first ordering, capped at 10) with instructions to pick an
exact name. A single match still resolves automatically, and DIRECT_ALIASES
exact mappings are unaffected.

The date-stamp split from #67571 is kept, demoted from selection logic to
display ordering of the candidate list.

Supersedes the auto-pick approach of #67571; credit to @Sahaun and @GottZ
for the date-stamp parser analysis that this builds on.
2026-08-08 18:35:31 -07:00
Sohom Sahaun 21bc9ba341 fix(model-switch): split YYYYMMDD date stamps from version tuple in _model_sort_key
_model_sort_key treated YYYYMMDD snapshot stamps (e.g.
claude-opus-4-20250514) as version components, so 20250514 > 8
and resolve_alias("opus", "anthropic") returned the wrong model.

Fix: split components ≥ 19_000_101 (smallest plausible date stamp)
out of the version tuple, keeping them as a trailing tiebreaker so
bare IDs sort before their dated snapshots and newer snapshots
before older ones.  Shorter numeric components (mistral-large-2411,
gpt-4-0613) keep their current behavior.  No models.dev dependency
in the sort path.
2026-08-08 18:35:31 -07:00
Austin Pickett e0c3caf3b8 fix(model-picker): serve cached custom-provider catalog on no-probe opens (supersedes #81665, #81556) (#81973)
* fix(model-picker): serve cached custom-provider catalog on no-probe opens

#58183 stopped GUI picker opens from live-probing saved custom
OpenAI-compatible endpoints so a stopped local server could not stall the
picker. It gated the whole discovery block, not just the network call, so
`cached_fetch_api_models()` was skipped too — and with it the catalog an
earlier probe had already written to `provider_models_cache.json`.

A custom endpoint that is not the current provider therefore renders only
the models named in its config entry. A local server with 8 models loaded
shows the 1 model that was saved when the provider was first added, on
every picker open, while an explicit Refresh shows all 8.

Add `cache_only` to `cached_fetch_api_models()`: answer from disk within
the existing stale-serve window, never fetch, never revalidate off-thread,
return None on a miss. Split the three call sites in
`list_authenticated_providers()` into what the user's config permits
(`discover_models`, an explicit `models:` allowlist) and how we may obtain
it, so suppressing the probe now downgrades to a cached read instead of
skipping discovery outright. `discover_models: false` still pins, and a
cache hit no longer writes back to config since the probe that populated
it already did.

The latency win stands: a cold cache is a miss, so picker opens against
offline endpoints still make zero network calls.

* test(model-picker): pin the cached-catalog contract for no-probe opens

Cover both halves of the invariant, since fixing either one alone
reintroduces a bug the other guards against.

`cache_only` on `cached_fetch_api_models()`: a fresh entry and an entry
past its TTL but inside the stale-serve window both serve; an entry beyond
that window, an empty cache, rotated credentials, `force_refresh`, and a
missing base_url are all misses — and none of them fetch or spawn a
background revalidation.

`list_authenticated_providers()` on the GUI path: a non-current endpoint
with a warm cache reports its full catalog across all three provider
shapes (`custom_providers`, `providers:`, bare `provider: custom`) with no
live fetch attempted. A cold cache keeps the configured list and still
makes no network call, which is the #58183 guarantee. `discover_models:
false` keeps pinning, and a cache hit does not write back to config.

* fix: persist discovered custom-provider models in the hermes model flow

The `hermes model` named-custom-provider flow (_model_flow_named_custom)
probes the endpoint and shows the full catalog, but never persists it to the
entry's `models:` list. No-probe surfaces (dashboard, desktop, ACP) call
build_models_payload(..., probe_custom_providers=False) and only render the
configured `models:` list, so a provider added via `hermes model` collapses
to the single `model:` default everywhere except the CLI. OpenAI-compatible
providers added via a probing picker already benefit from
_save_discovered_models_to_config; the CLI flow did not.

Persist the live catalog after a successful probe, mirroring the picker path
in model_switch.py. A failed save is non-fatal.

* fix(model-picker): stop an auto-saved catalog pinning a keyless endpoint

The cached-catalog read added for no-probe picker opens still sat behind
the no-key discovery gate, so it never reached the shape that motivated
it: a keyless local model server.

`bool(api_key) or not has_explicit_models` is a network-cost gate. It
exists so Hermes does not probe an endpoint it cannot authenticate to
when that endpoint already declares its catalog (5f00f36ba, 1039e90b5).
Reading a catalog an earlier probe already paid for costs nothing, so
the gate belongs on the probe, not on discovery as a whole.

Left on the discovery side it re-pins the endpoint it was meant to
spare. A successful probe calls `_save_discovered_models_to_config()`,
which writes a plain list into `models:` — exactly the shape
`_models_config_is_allowlist()` reads back as an explicit user
allowlist. A keyless server therefore froze on the catalog of its first
probe and could never widen again, which is the "lineup changes after
config was written" case. f66319097 already carved the dict shape out of
this trap for the same reason; the list shape is the other door into it.

Move the clause to `_probe_live` at both custom-endpoint sites. Probe
suppression is unchanged — verified byte-identical to main across the
keyed/keyless x declared/undeclared matrix — and `discover_models: false`
remains the documented way to pin a catalog.

* test(model-picker): cover the keyless auto-save pinning trap

Three tests around the gate move, each failing on the code before it:

- a keyless endpoint carrying an auto-saved `models:` list still reads
  its full cached catalog
- the same row, cold cache and probing enabled, still makes zero live
  fetches — the network-cost gate the clause exists for
- an end-to-end round trip: persist a probe result via
  `_save_discovered_models_to_config()`, reload it, and assert the shape
  we wrote does not read back as a user pin

The round-trip test guards the whole chain rather than one branch, so a
future change that makes the saved shape look like an intentional
allowlist fails here even if the gate logic is refactored.

* fix(model-picker): key the custom-endpoint model cache by api_mode

`cached_fetch_api_models()` fingerprints entries with `api_mode`, but no
call site in `list_authenticated_providers()` passed it, so every custom
row resolved to the `api_mode=None` fingerprint. Two rows sharing a
base_url and credential but differing by `api_mode` are deliberately
distinct picker rows — it is part of `group_key` at both sites — yet they
collapsed onto one cache entry.

That was latent while probing was the only way to fill a row: a mismatched
entry was overwritten by the row's own live fetch. Serving that entry
without a probe makes it visible, so an `anthropic_messages` row could
render the catalog an OpenAI-mode row cached against the same URL. The
wire protocols differ (`x-api-key` + `anthropic-version` vs
`Authorization: Bearer`), so those catalogs are not interchangeable.

Persist `api_mode` on the group at both grouping sites — it is already
part of `group_key`, so it is constant across the group — and pass it
into the cache read. Section 3b (bare `provider: custom`) has no
`api_mode` in scope and already reads with the empty-credential
fingerprint, so it is unchanged.

Reported by Copilot review on #81973.

---------

Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
Co-authored-by: Navlem <114683850+Navlem@users.noreply.github.com>
2026-08-08 16:07:03 -04:00
Prashant Jain fb435aae97 perf(model): disk-cache custom-provider /v1/models probes
Custom OpenAI-compatible endpoints (named custom_providers rows, bare
provider: custom, and per-endpoint-map entries) called fetch_api_models()
directly at three call sites in model_switch.py, with no disk cache — unlike
first-class providers, which go through cached_provider_model_ids(). Every
plain /model open live-probed the active custom endpoint's /v1/models,
regardless of how recently it had already been probed.

Adds cached_fetch_api_models() in hermes_cli/models.py: a TTL disk-cache
wrapper keyed on custom:<base_url> (custom endpoints have no
PROVIDER_REGISTRY slug to key on) and fingerprinted on api_key/api_mode/
headers, with the same stale-beats-nothing fallback policy as
cached_provider_model_ids(). Routes all three probe call sites through it.

Since prewarm_picker_cache_async() already calls list_authenticated_providers()
with probe_custom_providers defaulting True, this also fixes the endpoint
being warmed on boot (populating the disk cache) instead of that work being
discarded on every open — any custom endpoint (an LLM gateway, a
self-hosted vLLM/SGLang server, etc.), not just one specific provider.

Fixes #72762. Salvaged from #72810 per review feedback: extracts just the
verified custom-endpoint cache fix with real cache-contract test coverage
(hit/stale/rotation/refresh/fallback), leaving the credential-pool and
Copilot-token-exchange costs described in the issue for separate follow-up.
2026-08-07 21:02:40 +05:30
HexLab98 f66319097e fix(model-switch): treat models dict as metadata, not allowlist
hermes model saves custom_providers models: {default: {context_length}} for
local Ollama. That dict shape was treated as an explicit catalog, so no-key
endpoints skipped live /v1/models probing and Desktop/Telegram only showed
the saved default — Refresh could not help. Keep list/string shapes as
allowlists; pin dict catalogs with discover_models: false.
2026-08-04 08:52:31 -07:00
ZachariahChu de0ce24c2e perf(cli): cap /model picker custom-endpoint probe at 1.5s
The interactive /model picker probes the current custom endpoint live via
fetch_api_models(), which defaults to a 5s timeout. A slow or flaky custom
endpoint blocks the picker for up to 5s on open. The lmstudio picker probe
already uses a 1.5s timeout for exactly this reason; apply the same fast-fail
bound to the three custom-provider probe sites, gated on for_picker so the
non-picker (5s) path is unchanged.
2026-08-02 22:16:52 +05:30
Drexuxux 95eae03883 fix(gateway): offload /model context-length resolution off the event loop
resolve_display_context_length() runs two blocking chains: the route
comparison in should_clear_context_pin() and the provider probe ladder in
get_model_context_length() (blocking requests calls to Anthropic /v1/models,
Copilot, Nous, Codex, GMI, Ollama, models.dev and OpenRouter).

The gateway message path already offloads both via
get_model_context_length_async() and should_clear_context_pin_async(), but
the /model slash-command handlers (_handle_model_command, _finish_switch)
called the sync helper directly, freezing the whole event loop for the
duration of the probe ladder - no messages processed on any platform, and
the Discord heartbeat timeouts that get_model_context_length_async() was
introduced to prevent.

Add resolve_display_context_length_async(), a thin asyncio.to_thread wrapper
mirroring the two existing *_async helpers (no logic duplication), and await
it at both handlers.
2026-07-31 22:39:34 -07:00
Gille 2de1e86c16 fix(cli): stabilize custom provider identities
Use providers keys as the canonical custom-provider identity while accepting legacy bare keys, display-name slugs, bare custom fallback, and doubled custom prefixes across resolution, pickers, doctor, and runtime reverse lookup.

Co-authored-by: Bakhtier Sizhaev <bakhtiersizhaev@users.noreply.github.com>
2026-07-31 17:26:42 +05:30
teknium1 ba7da1332c refactor: single-owner model switch parsing + effective-model resolution (kills the api_server/run.py divergence class) 2026-07-29 11:54:09 -07:00
teknium1 5b751dc0ad chore: remove unused imports and dead locals (ruff F401/F841 sweep)
Cleans F401 unused imports and F841 dead local assignments across
root *.py, agent/, hermes_cli/, tools/, gateway/, cron/, tui_gateway/
(tests/, plugins/, skills/ excluded).

Intentionally KEPT (false positives / test-patch surfaces):
- agent/transports/__init__.py package re-exports
- cli.py browser_connect re-exports (DEFAULT_BROWSER_CDP_URL area,
  used by tests/cli/test_cli_browser_connect.py)
- hermes_cli/main.py _prompt_auth_credentials_choice /
  _model_flow_bedrock_api_key (accessed via main_mod attr in tests)
- gateway/run.py aliased replay_cleanup + whatsapp_identity re-exports
  and _PORT_BINDING_PLATFORM_VALUES (test-referenced)
- hermes_cli/web_server.py get_running_pid (tests monkeypatch it) and
  _OAUTH_TOKEN_URL availability probe
- hermes_cli/config.py get_process_hermes_home re-export (noqa'd F811
  chain) and yaml availability-probe import
- hermes_cli/nous_subscription.py managed_nous_tools_enabled
  (tests patch hermes_cli.nous_subscription.managed_nous_tools_enabled)
- try/except ImportError availability probes (env_loader, tts_tool,
  mcp_tool, web_server anthropic OAuth block)
- tools/web_tools.py noqa F401 re-exports
- hermes_cli/setup_whatsapp_cloud.py:263 'proceed' skipped: possible
  missing-guard bug, flagged for separate review
- unused function parameters (signature changes out of scope)

Side-effect RHS calls preserved where only the binding was dead
(e.g. web_server proc = _spawn_hermes_action -> bare call).
2026-07-29 11:53:39 -07:00
rob-maron 02d5e23085 nous portal anthropic wire 2026-07-27 11:53:48 -04:00
Kevin Haddock a75ec9278c fix(model): track explicit models: declarations in section 3 so a singular default_model doesn't suppress live discovery
A providers: entry with only a default_model/model (no explicit models:
list) is un-narrowed — the singular field is just the active selection.
Section 3 derived has_explicit_models from the merged models list, so
the lone default_model entry counted as an explicit catalog and
suppressed the /v1/models probe for no-key endpoints, leaving a
one-line /model picker menu for local llama.cpp/Ollama/vLLM servers.

Track explicit models: declarations separately at group-build time
(mirrors section 4's declaration-tracking from #40542 / PR #61928) and
gate the probe on that instead.

Salvaged from PR #68984 by @vigilancetech-com (the probe_custom_providers
gate removal in that PR is not taken — the GUI no-probe gate is
intentional).
2026-07-26 17:17:03 -07:00
ijevin 8ca4c745d0 fix(models): resolve custom provider model ids
Map picker-prefixed custom provider selections back to their configured model IDs before validation, persistence, and API requests.

Fixes #68347
2026-07-24 21:24:36 -05:00
Teknium df051c17cc fix(vertex): surface vertex in the /model picker — credential gate + curated model list
Community verification of #56688 (zmack12344321) found two follow-up gaps
that kept Vertex invisible in the /model menu even after registry
registration:

1. hermes_cli/model_switch.py: list_authenticated_providers() had a
   credential gate hard-coded to API keys (with an aws_sdk special case
   only) — add a vertex branch using has_vertex_credentials(), mirroring
   the aws_sdk shape.
2. hermes_cli/models.py: Vertex's OpenAI-compatible endpoint has no
   /models listing route, so without a curated _PROVIDER_MODELS entry the
   picker only ever showed the current model — add a Gemini curated list.

Follow-up to #56688.
2026-07-23 16:55:41 -07:00