Commit Graph

4 Commits

Author SHA1 Message Date
ouyangbo 012c9c6bed chore: anchor fresh-start history to upstream 2026-09-16 15:06:54 +08:00
teknium1 b980495847 refactor: move live-model benchmark harnesses from scripts/ into evals/
scripts/ is for repo tooling (tests runner, release, installers, CI
checks). The tool-search live tests, the toolperf A/B eval and the browser
eval benchmark are offline benchmarks that spend model budget, which is
exactly what evals/ holds; evals/browser_use already cited
scripts/toolperf_abeval as "the same pattern".

- scripts/tool_search_livetest*.py + analyze_livetest.py + LIVETEST_README
  -> evals/tool_search/ (README.md); repo-root sys.path hop adjusted for
  the extra directory level; gitignore now covers evals/tool_search/out*/
- scripts/toolperf_abeval/ -> evals/toolperf_abeval/
- scripts/benchmark_browser_eval.py -> evals/browser_use/
2026-09-13 06:06:46 -07:00
Teknium de60f789a7 simplify(compat): tools/browser_tool + browser_supervisor — drop 114 re-exports + 6 legacy aliases + PEP 562 requests/call_llm hook, repoint 21 non-test callers; siblings read sibling names directly 2026-09-03 14:16:54 -07:00
Teknium d4b26df897 perf(browser): route browser_console eval through supervisor's persistent CDP WS (180x faster) (#23226)
Adds CDPSupervisor.evaluate_runtime() and wires it into _browser_eval as a
fast path when a supervisor is alive for the current task_id. Replaces the
~180ms agent-browser subprocess fork+exec+Node-startup hop with a ~1ms
Runtime.evaluate over the supervisor's already-connected WebSocket.

Falls through to the existing agent-browser CLI path when no supervisor is
running (e.g. backends without CDP, or before the first browser_navigate
attaches one), so behaviour is unchanged where it can't apply.

JS-side exceptions surface directly without falling through to the
subprocess (the subprocess would just re-raise the same error, slower);
supervisor-side failures (loop down, no session) fall through cleanly.

Benchmark — 30 iterations of `1 + 1` against headless Chrome:
  supervisor WS              mean=  0.96ms  median=  0.91ms
  agent-browser subprocess   mean=179.35ms  median=167.73ms
  → 187x speedup mean

Tests: 14 unit tests (mocked supervisor + response-shape coverage), 5
real-Chrome e2e tests in test_browser_supervisor.py (gated on Chrome
being installed). Browser test suite: 355 passed, 1 skipped.
2026-05-10 07:37:55 -07:00