📖 The AI Tool Bible
coding

HumanEval

OpenAI's 164-problem Python coding benchmark. Model must complete a function given its docstring. Nearly saturated at the frontier — kept for historical continuity.

Official leaderboard →

Top scores

#ModelScore
1gpt-596.5%
2claude-opus-4-895.0%
3gemini-2-5-pro94.4%
4deepseek-v391.5%

Scores are snapshots from public leaderboards at the time of last update. Follow the source link for the live board.