21cc643ac2
Scenarios target real confusion clusters in Epic's UE 5.8 catalog (StaticMesh vs SkeletalMesh set_material, three tag systems, CurveTable vs DataTable rows, Niagara Component vs System variables, four capture variants, zero-keyword phrasing). Mocks return realistic editor errors on wrong-type calls; scoring is strict (clean solve = correct tool with zero distractor calls; first-call accuracy tracked separately). Key result: first-call selection is unreliable in EVERY mode — eager with all 199K of schemas in context managed 2/10 — but clean solves stay 75-95% because agents probe (get_components, get_material_slots) before committing. The probe loop works through the 3-tool bridge at 1/4 the cost of eager ($1.60-1.69 vs $6.49/task, Opus 4.8). On Haiku the listing beats bare bridge 18/20 vs 15/20 (core-tool substitution again). Zero distractor invocations across all 50 Opus runs.