reasoning
BIG-Bench Hard
23 tasks from BIG-Bench that were hard for models at the time. Multi-step reasoning across NLP, math, and symbolic tasks.
Official leaderboard →Top scores
| # | Model | Score |
|---|---|---|
| 1 | gpt-5 | 90.3% |
| 2 | claude-opus-4-8 | 89.8% |
| 3 | gemini-2-5-pro | 87.0% |
Scores are snapshots from public leaderboards at the time of last update. Follow the source link for the live board.