a124d16764
Four independent pre-request stalls sat on the critical path between prompt submission and the first streamed token, measured with cProfile against a live process: 1. Discord capability detection (~2.0s, worst 5s): get_tool_definitions -> _get_dynamic_schema made a BLOCKING https call to discord.com inside AIAgent.__init__ for any user with DISCORD_BOT_TOKEN set, on every platform, every cold process. Now non-blocking: memory cache -> 24h disk cache -> permissive default + one background detection that seeds the disk cache for the next process. The permissive default is pinned per-process so tool schemas never flip mid-conversation (prompt-cache safety); it mirrors the existing detection-failure fallback (all actions exposed, 403s enriched at call time). 2. Ollama /api/show probe (~0.3s): get_model_context_length step 5e POSTed to <base_url>/api/show for KNOWN providers (openrouter etc.), got a 404, and never cached the miss - so every fresh process paid a full HTTP round-trip. Known non-Ollama providers now skip the probe; local/custom/unknown endpoints keep the exact previous behavior. 3. env_probe subprocess sweep (~0.5s): the Python-toolchain probe ran 4-8 subprocess calls inside the FIRST system prompt build. Now warmed off-thread during agent init; the prompt build hits the cache (same lock, so a mid-flight warm just joins instead of recomputing). 4. tools.mcp_tool import (~0.4s): the between-turns MCP refresh in build_turn_context imported the whole mcp package even with zero MCP servers configured. MCP tools can only exist if tools.mcp_tool was already imported (discovery/reload paths), so gate the import on sys.modules membership - no behavior change for MCP users. CLI additionally pre-imports run_agent + openai off-thread during the idle banner window (same pattern as the /model picker prewarm), hiding the remaining ~1.5s of module imports while the user types. Fixes 1-4 apply to every interaction layer (CLI, gateway, TUI, desktop, cron). Measured cold first turn (submit -> request dispatched, openrouter, discord token set): 4.3s before -> 0.9s after CLI prewarm (~80%); the agent-side non-import cost drops 2.9s -> 0.36s (init) + 0.27s (turn prologue).