CursorBench 4.0
Cursor's own test of coding agents on ambiguous, multi-file tasks taken from real Cursor sessions — editing, refactoring, investigating a codebase, understanding what the user meant, managing jobs and following a design. Cursor runs each model at several reasoning efforts; each release here carries the score of its best listed effort. Scores aren't comparable with earlier CursorBench versions. Higher is better.
Scores from CursorBench (Cursor), which runs the benchmark and publishes the full field.
Rankings
Higher is betterScores marked “via” above are quoted with attribution from CursorBench, retrieved 28 September 2026.
CursorBench 4.0 — frequently asked questions
Cursor's own test of coding agents on ambiguous, multi-file tasks taken from real Cursor sessions — editing, refactoring, investigating a codebase, understanding what the user meant, managing jobs and following a design. Cursor runs each model at several reasoning efforts; each release here carries the score of its best listed effort. Scores aren't comparable with earlier CursorBench versions. Higher is better.