AI Benchmarks
Every benchmark we track across the major AI labs. Click any benchmark to see the full ranking of models, from the highest score to the lowest.
Coding
SWE-Bench Pro
long-horizon, real-world software engineering
SWE-Bench Verified
real-world software engineering
SWE-Bench Multilingual
software engineering across many languages
CursorBench v3.1
Cursor agentic coding eval (harder tasks)
DeepSWE 1.1
Artificial Analysis agentic software-engineering eval (v1.1)
DeepSWE 1.0
Artificial Analysis agentic software-engineering eval
Next.js Evals
Next.js code generation & migration tasks
LiveCodeBench
contamination-free code generation on freshly published problems
Expert-SWE (Internal)
OpenAI's internal software-engineering eval
Terminal & CLI
Agentic & tool use
MCP Atlas
multi-step workflows using MCP
JobBench
professional workplace tool-use tasks
Toolathlon-Verified
verified real-world personal tool use
Toolathlon
real-world general tool use
BrowseComp
agentic web browsing and research
CyberGym
cybersecurity agentic tasks
OSWorld-Verified
operating a real computer to finish tasks