8adef09be8
Review fixes for #94618 (all three blockers reproduced by the reviewer through the real web_extract_tool): 1. Cache lookup moved AFTER provider resolution and strict-selection validation, and gated per-URL on the website blocklist policy — a blocklist-blocked or misconfigured-backend call now behaves exactly as it would without a cache instead of serving cached content. 2. Rescue-served extract batches are never cached (mirrors the search memo's exclusion), keeping one-shot rescue one-shot. 3. Cache entries now get dedicated per-(url, format, provider) files instead of sharing the URL-keyed truncate-store file — html and markdown (or two backends') copies of one URL no longer overwrite each other, and switching extract backends within the TTL never serves the old backend's rendering. Also from review: per-process index tmp filename (cross-process writers can no longer truncate each other mid-write) and held flight locks are never evicted from the bounded lock table (eviction could have allowed a duplicate paid request). New regression tests for formats/provider keying; E2E harness extended with policy-block, strict-selection, rescue-two-call, and dual-format scenarios — 6/6 pass; original 13/13 still pass.