coding
HumanEval
OpenAI's 164-problem Python coding benchmark. Model must complete a function given its docstring. Nearly saturated at the frontier — kept for historical continuity.
Official leaderboard →Top scores
| # | Model | Score |
|---|---|---|
| 1 | gpt-5 | 96.5% |
| 2 | claude-opus-4-8 | 95.0% |
| 3 | gemini-2-5-pro | 94.4% |
| 4 | deepseek-v3 | 91.5% |
Scores are snapshots from public leaderboards at the time of last update. Follow the source link for the live board.