d7a4065568
Local-model users paid a fresh probe waterfall on EVERY CLI cold start inside AIAgent.__init__: detect_local_server_type (up to 4 HTTP GETs, 2s timeout each on a hung server) + /api/show (3s timeout). The existing caches were in-process only, so back-to-back invocations (chat -q, cron ticks, subagents) re-paid the network every time. - New 300s-TTL disk L2 at HERMES_HOME/cache/local_endpoint_probes.json for detect_local_server_type verdicts and query_ollama_num_ctx results. Only SUCCESSFUL probes persist (a down server never pins a negative verdict); stale entries pruned on write; corrupted cache degrades to a miss; atomic writes. 300s is strictly fresher than the 1h in-process TTL that already accepts server-swap staleness. - models.dev fetch timeout 15 -> (5, 10) connect/read tuple: a blackholed connect stalled the first-turn critical path 15s; now fails in 5s (matches the OpenRouter fetch convention, #46620). - _auto_detect_local_model timeout 5 -> (2, 3): runs inside _get_model_config() at startup against a LOCAL endpoint; a hung local server cost 5s before the banner. E2E (real HTTP server, two fresh subprocesses, isolated HERMES_HOME): proc1 = 2 HTTP hits, proc2 = 0 HTTP hits, identical results (ollama/131072), probe wall 74.5 -> 35.5 ms. 222 targeted tests green incl. 9 new disk-L2 contract tests.