32fe129324
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it pays agent startup on each hop. Profiling one hop showed the single largest controllable cost was a live GET /models against the provider on EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow endpoint) — the in-memory endpoint-metadata cache is per process and the Nous persistent context cache is bypassed by design so the portal stays authoritative. - model_metadata: memoize successful remote /models probes on disk (cache/endpoint_model_metadata.json) with the SAME 300s TTL as the in-memory cache, so authority semantics are unchanged (reconciliation still lands within 5 minutes) but the answer is shared across processes. Local endpoints are never memoized (LM Studio reloads). - bot_relay: the cross-machine reply waiter polls the reply file every 250ms instead of every 2s — up to 2s of dead air on every relayed reply. Nothing here changes turn ordering: DMs and group rounds stay serial. Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs): main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.