fb435aae97
Custom OpenAI-compatible endpoints (named custom_providers rows, bare provider: custom, and per-endpoint-map entries) called fetch_api_models() directly at three call sites in model_switch.py, with no disk cache — unlike first-class providers, which go through cached_provider_model_ids(). Every plain /model open live-probed the active custom endpoint's /v1/models, regardless of how recently it had already been probed. Adds cached_fetch_api_models() in hermes_cli/models.py: a TTL disk-cache wrapper keyed on custom:<base_url> (custom endpoints have no PROVIDER_REGISTRY slug to key on) and fingerprinted on api_key/api_mode/ headers, with the same stale-beats-nothing fallback policy as cached_provider_model_ids(). Routes all three probe call sites through it. Since prewarm_picker_cache_async() already calls list_authenticated_providers() with probe_custom_providers defaulting True, this also fixes the endpoint being warmed on boot (populating the disk cache) instead of that work being discarded on every open — any custom endpoint (an LLM gateway, a self-hosted vLLM/SGLang server, etc.), not just one specific provider. Fixes #72762. Salvaged from #72810 per review feedback: extracts just the verified custom-endpoint cache fix with real cache-contract test coverage (hit/stale/rotation/refresh/fallback), leaving the credential-pool and Copilot-token-exchange costs described in the issue for separate follow-up.