AI Benchmarks

Every benchmark we track across the major AI labs. Click any benchmark to see the full ranking of models, from the highest score to the lowest.

Coding

ProgramBench
fully rebuilding programs from their executable and documentation· 17 models
SWE-Bench Pro
long-horizon, real-world software engineering· 23 models
SWE-Bench Verified
real-world software engineering· 57 models
SWE-Bench Multilingual
software engineering across many languages· 14 models
SWE-Bench Multimodal
software engineering on issues that carry screenshots and visual context· 3 models
DeepSWE 1.1
Artificial Analysis agentic software-engineering eval (v1.1)· 29 models
DeepSWE 1.0
Artificial Analysis agentic software-engineering eval· 6 models
FrontierSWE V2
34 ultra-long-horizon engineering challenges, scored mean@5 on partial credit· 2 models
SWE-Marathon v1.1
autonomous completion of 20 multi-hour software engineering tasks, all verifiers passing· 2 models
FrontierCode v1.1 (Main)
frontier-difficulty agentic coding tasks (v1.1, main split)· 5 models
APEX-SWE
expert-level software-engineering tasks (AI Productivity Index)· 2 models
MLE-Bench
machine-learning engineering on real Kaggle competitions· 3 models
PaperBench
reproducing the results of an ML research paper end to end· 1 model
Next.js Evals
Next.js code generation & migration tasks· 32 models
Supabase Evals
building and debugging real Supabase projects as a coding agent· 7 models
NL2Repo-Bench
generating repository-level code from natural-language requirements· 6 models
SWEAtlas CodeBase QnA
answering questions about an unfamiliar codebase· 1 model
QwenSWEBench
Qwen's in-house benchmark for software-engineering capabilities· 1 model
QwenSWEBench V2
Qwen's in-house benchmark for complex real-world software engineering (v2)· 2 models
AA Coding Agent Index
Artificial Analysis composite of three coding-agent evals, scored per harness-and-model pairing· 1 model
Database Migration Tasks (OpenAI Internal)
OpenAI's internal set of database migration tasks· 1 model
CADGenBench
building CAD geometry from a design prompt, graded for executability and geometric correctness· 2 models
BenchCAD
writing and editing parametric CAD programs for industrial parts· 1 model
LiveCodeBench
contamination-free code generation on freshly published problems· 12 models
HumanEval
writing a single Python function from its docstring — the standard coding benchmark of 2021–2024, retired in favour of SWE-Bench-style tests on real repositories· 3 models
Expert-SWE (Internal)
OpenAI's internal software-engineering eval· 2 models

Terminal & CLI

Agentic & tool use

APEX-Agents
expert-level agentic work tasks (AI Productivity Index)· 2 models
MCP Atlas
multi-step workflows using MCP· 11 models
JobBench
professional workplace tool-use tasks· 8 models
CoWorkBench
long-horizon productivity tasks across professional domains· 4 models
Toolathlon-Verified
verified real-world personal tool use· 9 models
Toolathlon
real-world general tool use· 5 models
BU Bench
browser agent task completion (Browser Use)· 7 models
BrowseComp
agentic web browsing and research· 28 models
DeepSearchQA
multi-step web search and answer synthesis· 1 model
CyberGym
cybersecurity agentic tasks· 13 models
CVE-Bench
exploiting real-world web-application CVEs in sandboxed production-like services· 3 models
CathedralBench
multi-exploit chain execution on the hard subset, run by a third-party red team· 2 models
ExploitBench
exploit development on known V8 vulnerabilities (CMU capability-ladder eval)· 2 models
ExploitGym
end-to-end exploits built from known bugs (count of 898 instances solved)· 1 model
OSWorld 2.0
operating a real computer to finish tasks (v2.0)· 8 models
OSWorld-Verified
operating a real computer to finish tasks· 28 models
Agent's Last Exam
pass rate on multimodal desktop and OS agent tasks· 8 models
SRE-Bench
Kubernetes incident response and reliability work, best of four attempts· 1 model
AutomationBench
end-to-end business workflow automation· 13 models

Reasoning & science

Humanity's Last Exam
expert-level questions across disciplines· 44 models
Humanity's Last Exam (Verified)
expert-level questions across disciplines, on the verified question set· 2 models
ARC-AGI-3
novel problem-solving in interactive environments· 2 models
ARC-AGI-2
abstract reasoning puzzles· 21 models
FrontierMath
research-level mathematics· 6 models
BioMysteryBench
solving open biological research mysteries· 3 models
LatchBio Capabilities v1.0
agentic analysis of real biological data, averaged over 11 capability benchmarks· 2 models
GeneBench-Pro
multistage statistical reasoning in genomics and translational biomedicine· 1 model
MedChemBench (Internal)
OpenAI's medicinal-chemistry eval: structure-activity reasoning, ADME and lead optimisation· 1 model
LAB-Bench 2
real-world biology research tasks· 2 models
EEBench
designing circuits and hardware, graded for physical correctness and functionality· 2 models
GPQA Diamond
graduate-level science reasoning· 59 models
IFBench
following detailed instructions and constraints· 2 models
Agentic IF Index (Internal)
Meta's internal instruction-following eval for agentic runs· 1 model
MMLU
multiple-choice exam questions across 57 academic subjects· 4 models
GSM8K
grade-school maths word problems needing a few steps of arithmetic reasoning· 3 models

Knowledge work

Finance

Healthcare

Multimodal

Long context

Robustness

Community preference

Scores are published by the labs at launch, or gathered from the leaderboards that measure them — BenchLM, BullshitBench, Supabase Evals, Next.js Evals and threejseval. Where the numbers come from.

Recent releasesView allGPT-6 SolGPT-6 Luna