coding
SWE-bench Verified
Real-world software-engineering benchmark — the model is given a GitHub issue and the repo, and must produce a patch that passes hidden test cases. The Verified subset filters to 500 hand-verified issues. The default benchmark for coding agents.
Official leaderboard →Top scores
| # | Model | Score |
|---|---|---|
| 1 | claude-opus-4-8 | 74.5% |
| 2 | claude-sonnet-5 | 72.0% |
| 3 | gpt-5 | 71.0% |
| 4 | gemini-2-5-pro | 63.8% |
Scores are snapshots from public leaderboards at the time of last update. Follow the source link for the live board.