Function synthesis
HumanEval
164 small Python problems: the AI is given a function's description and has to write the working function. This was the coding benchmark of the GPT-3.5 and GPT-4 era, before the field moved to fixing real bugs in real repositories. Higher is better.
Rankings
Higher is betterHumanEval — frequently asked questions
- What is HumanEval?
- 164 small Python problems: the AI is given a function's description and has to write the working function. This was the coding benchmark of the GPT-3.5 and GPT-4 era, before the field moved to fixing real bugs in real repositories. Higher is better.
- Which AI model scores highest on HumanEval?
- Claude 2 by Anthropic holds the best HumanEval result among tracked models, at 71.2% (released Jul 11 2023). Higher scores are better on this benchmark.
- What are the top 3 models on HumanEval?
- 1. Claude 2 (Anthropic) — 71.2%; 2. GPT-4 (OpenAI) — 67%; 3. Claude Instant 1.2 (Anthropic) — 58.7%.
- How many models have a published HumanEval score?
- 3 tracked models have a published HumanEval score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.