Browser agent

BU Bench

Can the AI drive a real web browser to finish tasks — clicking, filling forms, and navigating sites the way a person would? Run by Browser Use on their BU Bench task set. Higher is better.

Rankings

Higher is better

BU Bench — frequently asked questions

What is BU Bench?
Can the AI drive a real web browser to finish tasks — clicking, filling forms, and navigating sites the way a person would? Run by Browser Use on their BU Bench task set. Higher is better.
Which AI model scores highest on BU Bench?
Claude Opus 4.8 by Anthropic holds the best BU Bench result among tracked models, at 74% (released May 28 2026). Higher scores are better on this benchmark.
What are the top 5 models on BU Bench?
1. Claude Opus 4.8 (Anthropic) — 74%; 2. Gemini 3.6 Flash (Google) — 68%; 3. GPT-5.6 Sol (OpenAI) — 67%; 4. Claude Sonnet 4.6 (Anthropic) — 62%; 5. Gemini 3.5 Flash (Google) — 58%.
How many models have a published BU Bench score?
7 tracked models have a published BU Bench score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks