AI Benchmarks
Every benchmark we track across the major AI labs. Click any benchmark to see the full ranking of models, from the highest score to the lowest.
Coding
ProgramBench
fully rebuilding programs from their executable and documentation· 17 models
SWE-Bench Pro
long-horizon, real-world software engineering· 23 models
SWE-Bench Verified
real-world software engineering· 57 models
SWE-Bench Multilingual
software engineering across many languages· 14 models
SWE-Bench Multimodal
software engineering on issues that carry screenshots and visual context· 3 models
DeepSWE 1.1
Artificial Analysis agentic software-engineering eval (v1.1)· 29 models
DeepSWE 1.0
Artificial Analysis agentic software-engineering eval· 6 models
FrontierSWE V2
34 ultra-long-horizon engineering challenges, scored mean@5 on partial credit· 2 models
SWE-Marathon v1.1
autonomous completion of 20 multi-hour software engineering tasks, all verifiers passing· 2 models
FrontierCode v1.1 (Main)
frontier-difficulty agentic coding tasks (v1.1, main split)· 5 models
APEX-SWE
expert-level software-engineering tasks (AI Productivity Index)· 2 models
MLE-Bench
machine-learning engineering on real Kaggle competitions· 3 models
PaperBench
reproducing the results of an ML research paper end to end· 1 model
Next.js Evals
Next.js code generation & migration tasks· 32 models
Supabase Evals
building and debugging real Supabase projects as a coding agent· 7 models
NL2Repo-Bench
generating repository-level code from natural-language requirements· 6 models
SWEAtlas CodeBase QnA
answering questions about an unfamiliar codebase· 1 model
QwenSWEBench
Qwen's in-house benchmark for software-engineering capabilities· 1 model
QwenSWEBench V2
Qwen's in-house benchmark for complex real-world software engineering (v2)· 2 models
AA Coding Agent Index
Artificial Analysis composite of three coding-agent evals, scored per harness-and-model pairing· 1 model
Database Migration Tasks (OpenAI Internal)
OpenAI's internal set of database migration tasks· 1 model
CADGenBench
building CAD geometry from a design prompt, graded for executability and geometric correctness· 2 models
BenchCAD
writing and editing parametric CAD programs for industrial parts· 1 model
LiveCodeBench
contamination-free code generation on freshly published problems· 12 models
HumanEval
writing a single Python function from its docstring — the standard coding benchmark of 2021–2024, retired in favour of SWE-Bench-style tests on real repositories· 3 models
Expert-SWE (Internal)
OpenAI's internal software-engineering eval· 2 models
Terminal & CLI
Frontier-Bench v0.1
diverse, difficult agentic computer-work tasks (successor to Terminal-Bench)· 9 models
Terminal-Bench 4.0
command-line task completion (v4.0, recalibrated task resources)· 18 models
Terminal-Bench 3.0
command-line task completion (v3.0, much harder task set)· 7 models
Terminal-Bench 2.1
command-line task completion· 30 models
Terminal-Bench 2.0
command-line task completion (v2.0)· 14 models
Terminal-Bench-Science 0.1
scientific computing tasks at the command line (v0.1)· 5 models
Agentic & tool use
APEX-Agents
expert-level agentic work tasks (AI Productivity Index)· 2 models
MCP Atlas
multi-step workflows using MCP· 11 models
JobBench
professional workplace tool-use tasks· 8 models
CoWorkBench
long-horizon productivity tasks across professional domains· 4 models
Toolathlon-Verified
verified real-world personal tool use· 9 models
Toolathlon
real-world general tool use· 5 models
BU Bench
browser agent task completion (Browser Use)· 7 models
BrowseComp
agentic web browsing and research· 28 models
DeepSearchQA
multi-step web search and answer synthesis· 1 model
CyberGym
cybersecurity agentic tasks· 13 models
CVE-Bench
exploiting real-world web-application CVEs in sandboxed production-like services· 3 models
CathedralBench
multi-exploit chain execution on the hard subset, run by a third-party red team· 2 models
ExploitBench
exploit development on known V8 vulnerabilities (CMU capability-ladder eval)· 2 models
ExploitGym
end-to-end exploits built from known bugs (count of 898 instances solved)· 1 model
OSWorld 2.0
operating a real computer to finish tasks (v2.0)· 8 models
OSWorld-Verified
operating a real computer to finish tasks· 28 models
Agent's Last Exam
pass rate on multimodal desktop and OS agent tasks· 8 models
SRE-Bench
Kubernetes incident response and reliability work, best of four attempts· 1 model
AutomationBench
end-to-end business workflow automation· 13 models
Reasoning & science
Humanity's Last Exam
expert-level questions across disciplines· 44 models
Humanity's Last Exam (Verified)
expert-level questions across disciplines, on the verified question set· 2 models
ARC-AGI-3
novel problem-solving in interactive environments· 2 models
ARC-AGI-2
abstract reasoning puzzles· 21 models
FrontierMath
research-level mathematics· 6 models
BioMysteryBench
solving open biological research mysteries· 3 models
LatchBio Capabilities v1.0
agentic analysis of real biological data, averaged over 11 capability benchmarks· 2 models
GeneBench-Pro
multistage statistical reasoning in genomics and translational biomedicine· 1 model
MedChemBench (Internal)
OpenAI's medicinal-chemistry eval: structure-activity reasoning, ADME and lead optimisation· 1 model
LAB-Bench 2
real-world biology research tasks· 2 models
EEBench
designing circuits and hardware, graded for physical correctness and functionality· 2 models
GPQA Diamond
graduate-level science reasoning· 59 models
IFBench
following detailed instructions and constraints· 2 models
Agentic IF Index (Internal)
Meta's internal instruction-following eval for agentic runs· 1 model
MMLU
multiple-choice exam questions across 57 academic subjects· 4 models
GSM8K
grade-school maths word problems needing a few steps of arithmetic reasoning· 3 models
Knowledge work
AA Intelligence Index
Artificial Analysis composite intelligence index across evals· 6 models
GDPval-AA
economically valuable knowledge work· 9 models
GDPval-AA v2.1
economically valuable knowledge work (v2.1, Crowd-BT Elo fit)· 4 models
GDPval-AA v2
economically valuable knowledge work (v2, re-based Elo)· 20 models
Design Tasks (OpenAI Internal)
OpenAI's internal set of professional design tasks· 1 model
Data Science Tasks (OpenAI Internal)
OpenAI's internal set of professional data-science tasks· 1 model
AA-Briefcase v1.1
Artificial Analysis agentic office-work eval (Elo, v1.1 rating fit)· 1 model
AA-Briefcase
Artificial Analysis agentic office-work eval (Elo)· 4 models
GDPval (win/tie rate)
GDPval win-or-tie rate vs human experts· 6 models
Finance
Legal & tax
Healthcare
Multimodal
CharXiv Reasoning
information synthesis from complex charts· 16 models
Chartography
chart-centred tasks, run with tools· 5 models
OfficeQA Pro
question answering over office documents· 1 model
MVBench
temporal reasoning over video clips· 1 model
MMVU
expert-level reasoning over discipline-specific video· 1 model
GDP.PDF
expert comprehension of PDF documents· 2 models
LVBench
understanding hour-long videos· 3 models
BabyVision
core visual reasoning· 4 models
OpenScore String Quartets
transcribing string-quartet sheet music, scored as 1 - OMR normalised edit distance· 1 model
MMMU-Pro
multimodal understanding and reasoning· 12 models
MMMU
multimodal understanding· 7 models
Blueprint-Bench 2
agentic spatial reasoning· 6 models
Long context
Robustness
Community preference
Scores are published by the labs at launch, or gathered from the leaderboards that measure them — BenchLM, BullshitBench, Supabase Evals, Next.js Evals and threejseval. Where the numbers come from.