b2598b41e1
read_file's auto-extraction covered only the stdlib trio (.ipynb/.docx/ .xlsx). firecrawl-anydoc (MIT, Rust core, imports as `anydoc`) converts Word, PowerPoint, Excel — including legacy .doc/.ppt/.xls — OpenDocument, RTF, EPUB, and PDF to clean Markdown through one shared document model. Wiring follows the footprint ladder: no new tool, no hard dependency. - tools/read_extract.py gains an ANYDOC_EXTENSIONS set that is active only when the converter imports; the stdlib extractors remain authoritative for their three formats so behavior is identical with or without the package. - tools/lazy_deps.py adds tool.doc_extract (firecrawl-anydoc==0.1.6), installed on first read of such a file with prompt=False so read_file can never block. Lazy-only for now: the package's first release was 2026-08-04, inside uv's 14-day exclude-newer quarantine, so the mirrored pyproject extra lands after it clears. - Any anydoc ConvertError maps to ExtractionError, falling back to the existing path/binary handling instead of erroring the tool. Tests: real-binding suite skips cleanly when the wheel is absent (verified: 15 passed/3 skipped without it, 18 passed with it), plus an absent-dep contract class that pins the fallback regardless of local install state.