📖 The AI Tool Bible
reasoning

BIG-Bench Hard

23 tasks from BIG-Bench that were hard for models at the time. Multi-step reasoning across NLP, math, and symbolic tasks.

Official leaderboard →

Top scores

#ModelScore
1gpt-590.3%
2claude-opus-4-889.8%
3gemini-2-5-pro87.0%

Scores are snapshots from public leaderboards at the time of last update. Follow the source link for the live board.