2684412ed5f00a9d6ec00d3f5b129bc4c658f419
3 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e0c3caf3b8 |
fix(model-picker): serve cached custom-provider catalog on no-probe opens (supersedes #81665, #81556) (#81973)
* fix(model-picker): serve cached custom-provider catalog on no-probe opens #58183 stopped GUI picker opens from live-probing saved custom OpenAI-compatible endpoints so a stopped local server could not stall the picker. It gated the whole discovery block, not just the network call, so `cached_fetch_api_models()` was skipped too — and with it the catalog an earlier probe had already written to `provider_models_cache.json`. A custom endpoint that is not the current provider therefore renders only the models named in its config entry. A local server with 8 models loaded shows the 1 model that was saved when the provider was first added, on every picker open, while an explicit Refresh shows all 8. Add `cache_only` to `cached_fetch_api_models()`: answer from disk within the existing stale-serve window, never fetch, never revalidate off-thread, return None on a miss. Split the three call sites in `list_authenticated_providers()` into what the user's config permits (`discover_models`, an explicit `models:` allowlist) and how we may obtain it, so suppressing the probe now downgrades to a cached read instead of skipping discovery outright. `discover_models: false` still pins, and a cache hit no longer writes back to config since the probe that populated it already did. The latency win stands: a cold cache is a miss, so picker opens against offline endpoints still make zero network calls. * test(model-picker): pin the cached-catalog contract for no-probe opens Cover both halves of the invariant, since fixing either one alone reintroduces a bug the other guards against. `cache_only` on `cached_fetch_api_models()`: a fresh entry and an entry past its TTL but inside the stale-serve window both serve; an entry beyond that window, an empty cache, rotated credentials, `force_refresh`, and a missing base_url are all misses — and none of them fetch or spawn a background revalidation. `list_authenticated_providers()` on the GUI path: a non-current endpoint with a warm cache reports its full catalog across all three provider shapes (`custom_providers`, `providers:`, bare `provider: custom`) with no live fetch attempted. A cold cache keeps the configured list and still makes no network call, which is the #58183 guarantee. `discover_models: false` keeps pinning, and a cache hit does not write back to config. * fix: persist discovered custom-provider models in the hermes model flow The `hermes model` named-custom-provider flow (_model_flow_named_custom) probes the endpoint and shows the full catalog, but never persists it to the entry's `models:` list. No-probe surfaces (dashboard, desktop, ACP) call build_models_payload(..., probe_custom_providers=False) and only render the configured `models:` list, so a provider added via `hermes model` collapses to the single `model:` default everywhere except the CLI. OpenAI-compatible providers added via a probing picker already benefit from _save_discovered_models_to_config; the CLI flow did not. Persist the live catalog after a successful probe, mirroring the picker path in model_switch.py. A failed save is non-fatal. * fix(model-picker): stop an auto-saved catalog pinning a keyless endpoint The cached-catalog read added for no-probe picker opens still sat behind the no-key discovery gate, so it never reached the shape that motivated it: a keyless local model server. `bool(api_key) or not has_explicit_models` is a network-cost gate. It exists so Hermes does not probe an endpoint it cannot authenticate to when that endpoint already declares its catalog ( |
||
|
|
7cf71c32bb |
fix: follow-ups for salvaged PR #80740
- Give cached_fetch_api_models the same stale-while-revalidate tier as cached_provider_model_ids: TTL-expired entries within the 7d window are served instantly while a background refresh rewrites the cache — without this, every /model open an hour into the session re-blocked on the live probe (#72762's stall class, deferred). - Generalize _spawn_swr_refresh(cache_key, refresh_fn) so non-slug custom:<base_url> keys reuse the same inflight-dedupe scaffolding; slug behavior unchanged (default refresh_fn preserved). - Convert the missed sibling site: acp_adapter/server.py _named_custom_provider_catalogs() live-probed every custom_providers row's /v1/models per ACP catalog build. - Extract _cache_entry_valid() (the fp/models predicate existed 4x) and validate 'at' is numeric so hand-edited/corrupt cache JSON degrades to a live fetch instead of raising through the picker's blanket except. - Flatten the dead api_mode conditional (fetch_api_models declares api_mode=None; branch was behaviorally inert). - Tests: 4 new guards (stale-serve, stale-window cutoff, generalized SWR write-through, corrupt-at degradation) — stale-serve and corrupt-at mutation-checked; 2 existing tests updated for the new behavior. |
||
|
|
fb435aae97 |
perf(model): disk-cache custom-provider /v1/models probes
Custom OpenAI-compatible endpoints (named custom_providers rows, bare provider: custom, and per-endpoint-map entries) called fetch_api_models() directly at three call sites in model_switch.py, with no disk cache — unlike first-class providers, which go through cached_provider_model_ids(). Every plain /model open live-probed the active custom endpoint's /v1/models, regardless of how recently it had already been probed. Adds cached_fetch_api_models() in hermes_cli/models.py: a TTL disk-cache wrapper keyed on custom:<base_url> (custom endpoints have no PROVIDER_REGISTRY slug to key on) and fingerprinted on api_key/api_mode/ headers, with the same stale-beats-nothing fallback policy as cached_provider_model_ids(). Routes all three probe call sites through it. Since prewarm_picker_cache_async() already calls list_authenticated_providers() with probe_custom_providers defaulting True, this also fixes the endpoint being warmed on boot (populating the disk cache) instead of that work being discarded on every open — any custom endpoint (an LLM gateway, a self-hosted vLLM/SGLang server, etc.), not just one specific provider. Fixes #72762. Salvaged from #72810 per review feedback: extracts just the verified custom-endpoint cache fix with real cache-contract test coverage (hit/stale/rotation/refresh/fallback), leaving the credential-pool and Copilot-token-exchange costs described in the issue for separate follow-up. |