04603fc040
Repeat searches (same normalized query + provider) within a 20-minute TTL are served from an in-process memo, and concurrent identical queries are single-flighted so a parallel subagent fan-out pays for one vendor request instead of N. Requested limits bucket up to 10/20/50/100 so near-identical requests share an entry; callers get their requested count sliced from the bucket. Repeat extracts of the same URL are served from the existing cache/web full-text store (previously written for read_file paging but never read back), via a small JSON sidecar index. Disk-backed, so CLI, gateway, cron, and subagents share it. Cached extracts re-run the normal truncate pipeline, so per-call char_limit still works. Both caches sit after every safety gate (secret-URL, SSRF, policy, provider resolution) and directly around the paid vendor call — hits skip only the network request. Only successful responses cache; rescue-served responses are never cached (one-shot rescue must stay one-shot). Config: web.cache_enabled (default on), web.cache_ttl_minutes (default 20, clamped 1-1440). Idea credit: query coalescing + num-bucketing pattern observed in Apodex FrontierAgent (Apache-2.0).