Function synthesis

HumanEval

164 small Python problems: the AI is given a function's description and has to write the working function. This was the coding benchmark of the GPT-3.5 and GPT-4 era, before the field moved to fixing real bugs in real repositories. Higher is better.

Rankings

Higher is better

HumanEval — frequently asked questions

What is HumanEval?
164 small Python problems: the AI is given a function's description and has to write the working function. This was the coding benchmark of the GPT-3.5 and GPT-4 era, before the field moved to fixing real bugs in real repositories. Higher is better.
Which AI model scores highest on HumanEval?
Claude 2 by Anthropic holds the best HumanEval result among tracked models, at 71.2% (released Jul 11 2023). Higher scores are better on this benchmark.
What are the top 3 models on HumanEval?
1. Claude 2 (Anthropic) — 71.2%; 2. GPT-4 (OpenAI) — 67%; 3. Claude Instant 1.2 (Anthropic) — 58.7%.
How many models have a published HumanEval score?
3 tracked models have a published HumanEval score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks