preference
MT-Bench
Multi-turn benchmark with 80 open-ended questions across 8 categories, judged by GPT-4. A quick chat-quality proxy before Arena Elo became dominant.
Official leaderboard →Top scores
| # | Model | Score |
|---|---|---|
| 1 | gpt-5 | 9.4 / 10 |
| 2 | claude-opus-4-8 | 9.3 / 10 |
| 3 | gemini-2-5-pro | 9.2 / 10 |
Scores are snapshots from public leaderboards at the time of last update. Follow the source link for the live board.