📖 The AI Tool Bible
coding

SWE-bench Verified

Real-world software-engineering benchmark — the model is given a GitHub issue and the repo, and must produce a patch that passes hidden test cases. The Verified subset filters to 500 hand-verified issues. The default benchmark for coding agents.

Official leaderboard →

Top scores

#ModelScore
1claude-opus-4-874.5%
2claude-sonnet-572.0%
3gpt-571.0%
4gemini-2-5-pro63.8%

Scores are snapshots from public leaderboards at the time of last update. Follow the source link for the live board.