AI Benchmarks

Every benchmark we track across the major AI labs. Click any benchmark to see the full ranking of models, from the highest score to the lowest.

Coding

SWE-Bench Pro
long-horizon, real-world software engineering· 20 models
SWE-Bench Verified
real-world software engineering· 44 models
SWE-Bench Multilingual
software engineering across many languages· 10 models
DeepSWE 1.1
Artificial Analysis agentic software-engineering eval (v1.1)· 18 models
DeepSWE 1.0
Artificial Analysis agentic software-engineering eval· 6 models
FrontierCode v1.1 (Main)
frontier-difficulty agentic coding tasks (v1.1, main split)· 2 models
APEX-SWE
expert-level software-engineering tasks (AI Productivity Index)· 2 models
MLE-Bench
machine-learning engineering on real Kaggle competitions· 3 models
PaperBench
reproducing the results of an ML research paper end to end· 1 model
Next.js Evals
Next.js code generation & migration tasks· 25 models
Supabase Evals
building and debugging real Supabase projects as a coding agent· 5 models
NL2Repo-Bench
generating repository-level code from natural-language requirements· 1 model
QwenSWEBench
Qwen's in-house benchmark for software-engineering capabilities· 1 model
LiveCodeBench
contamination-free code generation on freshly published problems· 8 models
HumanEval
writing a single Python function from its docstring — the standard coding benchmark of 2021–2024, retired in favour of SWE-Bench-style tests on real repositories· 3 models
Expert-SWE (Internal)
OpenAI's internal software-engineering eval· 2 models

Terminal & CLI

Agentic & tool use

Reasoning & science

Knowledge work

Finance

Healthcare

Multimodal

Long context

Robustness

Scores are published by the labs at launch, or gathered from the leaderboards that measure them — BenchLM, BullshitBench, Supabase Evals and Next.js Evals. Where the numbers come from.