Browser agent
BU Bench
Can the AI drive a real web browser to finish tasks — clicking, filling forms, and navigating sites the way a person would? Run by Browser Use on their BU Bench task set. Higher is better.
Rankings
Higher is betterBU Bench — frequently asked questions
- What is BU Bench?
- Can the AI drive a real web browser to finish tasks — clicking, filling forms, and navigating sites the way a person would? Run by Browser Use on their BU Bench task set. Higher is better.
- Which AI model scores highest on BU Bench?
- Claude Opus 4.8 by Anthropic holds the best BU Bench result among tracked models, at 74% (released May 28 2026). Higher scores are better on this benchmark.
- What are the top 5 models on BU Bench?
- 1. Claude Opus 4.8 (Anthropic) — 74%; 2. Gemini 3.6 Flash (Google) — 68%; 3. GPT-5.6 Sol (OpenAI) — 67%; 4. Claude Sonnet 4.6 (Anthropic) — 62%; 5. Gemini 3.5 Flash (Google) — 58%.
- How many models have a published BU Bench score?
- 7 tracked models have a published BU Bench score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.