Files
hermes-agent/evals/browser_use/tasks
Teknium 5033aedaa3 feat(evals): add the Browser Use mode A/B benchmark from PR #81958
Reconstructs the 204-run benchmark battery behind #81958 (Browser Use CLI
3.0 mode) as a rerunnable eval under evals/browser-use/, following the
toolperf_abeval / evals-compaction pattern.

- tasks/easy.json + tasks/hard.json: the oracle-checked toscrape task
  batteries (5 easy, 6 hard) exactly as run for the PR
- single_run.py: one cell = task x arm (base | pr | prns) x model x rep;
  throwaway HERMES_HOME, web-fetch creds stripped, arms pinned to separate
  trees via BUBENCH_BASE_TREE / BUBENCH_PR_TREE
- orchestrate.py: resume-safe local-CDP battery driver
- orchestrate_cloud.py: backend matrix (nous-cloud via the browser_use
  provider plugin, browserbase via REST) with per-cell session lifecycle
- report.py: scorecard aggregation with vs-base token deltas
- README.md: design, run instructions, and the recovered Aug 8-10 2026
  baseline scorecards (hard battery, backend matrix, easy round 1,
  digest ablation)

The original /tmp/bu-bench workspace was lost to a tmpfs reboot; harness
and readouts were recovered verbatim from the benchmark session's tool-call
history in state.db, with hardcoded paths parameterized. Smoke-verified
live: report.py aggregation, and single-cell runs (pr + base arms) against
a real headless Chrome CDP with sonnet-5 driving browser_exec, oracle pass.
2026-08-17 15:29:16 -07:00
..