AI Benchmarks
Every benchmark we track across the major AI labs. Click any benchmark to see the full ranking of models, from the highest score to the lowest.
Coding
SWE-Bench Pro
long-horizon, real-world software engineering
SWE-Bench Verified
real-world software engineering
SWE-Bench Multilingual
software engineering across many languages
CursorBench v3.2
Cursor agentic coding eval (v3.2 task set)
CursorBench v3.1
Cursor agentic coding eval (harder tasks)
DeepSWE 1.1
Artificial Analysis agentic software-engineering eval (v1.1)
DeepSWE 1.0
Artificial Analysis agentic software-engineering eval
FrontierCode v1.1 (Main)
frontier-difficulty agentic coding tasks (v1.1, main split)
MLE-Bench
machine-learning engineering on real Kaggle competitions
PaperBench
reproducing the results of an ML research paper end to end
Next.js Evals
Next.js code generation & migration tasks
Supabase Evals
building and debugging real Supabase projects as a coding agent
LiveCodeBench
contamination-free code generation on freshly published problems
HumanEval
writing a single Python function from its docstring — the standard coding benchmark of 2021–2024, retired in favour of SWE-Bench-style tests on real repositories
Expert-SWE (Internal)
OpenAI's internal software-engineering eval
Terminal & CLI
Agentic & tool use
MCP Atlas
multi-step workflows using MCP
JobBench
professional workplace tool-use tasks
Toolathlon-Verified
verified real-world personal tool use
Toolathlon
real-world general tool use
BU Bench
browser agent task completion (Browser Use)
BrowseComp
agentic web browsing and research
CyberGym
cybersecurity agentic tasks
OSWorld 2.0
operating a real computer to finish tasks (v2.0)
OSWorld-Verified
operating a real computer to finish tasks
AutomationBench
end-to-end business workflow automation
Reasoning & science
Humanity's Last Exam
expert-level questions across disciplines
ARC-AGI-3
novel problem-solving in interactive environments
ARC-AGI-2
abstract reasoning puzzles
FrontierMath
research-level mathematics
BioMysteryBench
solving open biological research mysteries
GPQA Diamond
graduate-level science reasoning
MMLU
multiple-choice exam questions across 57 academic subjects
GSM8K
grade-school maths word problems needing a few steps of arithmetic reasoning