5033aedaa3
Reconstructs the 204-run benchmark battery behind #81958 (Browser Use CLI 3.0 mode) as a rerunnable eval under evals/browser-use/, following the toolperf_abeval / evals-compaction pattern. - tasks/easy.json + tasks/hard.json: the oracle-checked toscrape task batteries (5 easy, 6 hard) exactly as run for the PR - single_run.py: one cell = task x arm (base | pr | prns) x model x rep; throwaway HERMES_HOME, web-fetch creds stripped, arms pinned to separate trees via BUBENCH_BASE_TREE / BUBENCH_PR_TREE - orchestrate.py: resume-safe local-CDP battery driver - orchestrate_cloud.py: backend matrix (nous-cloud via the browser_use provider plugin, browserbase via REST) with per-cell session lifecycle - report.py: scorecard aggregation with vs-base token deltas - README.md: design, run instructions, and the recovered Aug 8-10 2026 baseline scorecards (hard battery, backend matrix, easy round 1, digest ablation) The original /tmp/bu-bench workspace was lost to a tmpfs reboot; harness and readouts were recovered verbatim from the benchmark session's tool-call history in state.db, with hardcoded paths parameterized. Smoke-verified live: report.py aggregation, and single-cell runs (pr + base arms) against a real headless Chrome CDP with sonnet-5 driving browser_exec, oracle pass.