📖 The AI Tool Bible
preference

MT-Bench

Multi-turn benchmark with 80 open-ended questions across 8 categories, judged by GPT-4. A quick chat-quality proxy before Arena Elo became dominant.

Official leaderboard →

Top scores

#ModelScore
1gpt-59.4 / 10
2claude-opus-4-89.3 / 10
3gemini-2-5-pro9.2 / 10

Scores are snapshots from public leaderboards at the time of last update. Follow the source link for the live board.