AI Benchmarks
Every benchmark we track across the major AI labs. Click any benchmark to see the full ranking of models, from the highest score to the lowest.
Coding
SWE-Bench Pro
long-horizon, real-world software engineering· 20 models
SWE-Bench Verified
real-world software engineering· 44 models
SWE-Bench Multilingual
software engineering across many languages· 10 models
DeepSWE 1.1
Artificial Analysis agentic software-engineering eval (v1.1)· 18 models
DeepSWE 1.0
Artificial Analysis agentic software-engineering eval· 6 models
FrontierCode v1.1 (Main)
frontier-difficulty agentic coding tasks (v1.1, main split)· 2 models
APEX-SWE
expert-level software-engineering tasks (AI Productivity Index)· 2 models
MLE-Bench
machine-learning engineering on real Kaggle competitions· 3 models
PaperBench
reproducing the results of an ML research paper end to end· 1 model
Next.js Evals
Next.js code generation & migration tasks· 25 models
Supabase Evals
building and debugging real Supabase projects as a coding agent· 5 models
NL2Repo-Bench
generating repository-level code from natural-language requirements· 1 model
QwenSWEBench
Qwen's in-house benchmark for software-engineering capabilities· 1 model
LiveCodeBench
contamination-free code generation on freshly published problems· 8 models
HumanEval
writing a single Python function from its docstring — the standard coding benchmark of 2021–2024, retired in favour of SWE-Bench-style tests on real repositories· 3 models
Expert-SWE (Internal)
OpenAI's internal software-engineering eval· 2 models
Terminal & CLI
Frontier-Bench v0.1
diverse, difficult agentic computer-work tasks (successor to Terminal-Bench)· 9 models
Terminal-Bench 3.0
command-line task completion (v3.0, much harder task set)· 4 models
Terminal-Bench 2.1
command-line task completion· 25 models
Terminal-Bench 2.0
command-line task completion (v2.0)· 14 models
Agentic & tool use
APEX-Agents
expert-level agentic work tasks (AI Productivity Index)· 2 models
MCP Atlas
multi-step workflows using MCP· 11 models
JobBench
professional workplace tool-use tasks· 5 models
CoWorkBench
long-horizon productivity tasks across professional domains· 1 model
Toolathlon-Verified
verified real-world personal tool use· 5 models
Toolathlon
real-world general tool use· 5 models
BU Bench
browser agent task completion (Browser Use)· 7 models
BrowseComp
agentic web browsing and research· 20 models
CyberGym
cybersecurity agentic tasks· 7 models
ExploitBench
exploit development on known V8 vulnerabilities (CMU capability-ladder eval)· 1 model
ExploitGym
end-to-end exploits built from known bugs (count of 898 instances solved)· 1 model
OSWorld 2.0
operating a real computer to finish tasks (v2.0)· 2 models
OSWorld-Verified
operating a real computer to finish tasks· 19 models
Agent's Last Exam
pass rate on multimodal desktop and OS agent tasks· 3 models
AutomationBench
end-to-end business workflow automation· 5 models
Reasoning & science
Humanity's Last Exam
expert-level questions across disciplines· 22 models
Humanity's Last Exam (Verified)
expert-level questions across disciplines, on the verified question set· 1 model
ARC-AGI-3
novel problem-solving in interactive environments· 1 model
ARC-AGI-2
abstract reasoning puzzles· 15 models
FrontierMath
research-level mathematics· 6 models
BioMysteryBench
solving open biological research mysteries· 2 models
LAB-Bench 2
real-world biology research tasks· 1 model
GPQA Diamond
graduate-level science reasoning· 50 models
IFBench
following detailed instructions and constraints· 1 model
MMLU
multiple-choice exam questions across 57 academic subjects· 4 models
GSM8K
grade-school maths word problems needing a few steps of arithmetic reasoning· 3 models
Knowledge work
AA Intelligence Index
Artificial Analysis composite intelligence index across evals· 3 models
GDPval-AA
economically valuable knowledge work· 9 models
GDPval-AA v2
economically valuable knowledge work (v2, re-based Elo)· 16 models
AA-Briefcase
Artificial Analysis agentic office-work eval (Elo)· 2 models
GDPval (win/tie rate)
GDPval win-or-tie rate vs human experts· 6 models
Finance
Legal & tax
Healthcare
Multimodal
CharXiv Reasoning
information synthesis from complex charts· 14 models
GDP.PDF
expert comprehension of PDF documents· 1 model
LVBench
understanding hour-long videos· 1 model
BabyVision
core visual reasoning· 3 models
MMMU-Pro
multimodal understanding and reasoning· 10 models
MMMU
multimodal understanding· 7 models
Blueprint-Bench 2
agentic spatial reasoning· 6 models
Long context
Robustness
Scores are published by the labs at launch, or gathered from the leaderboards that measure them — BenchLM, BullshitBench, Supabase Evals and Next.js Evals. Where the numbers come from.