OlympicArena alternatives
6 evaluation tools in the same lane as OlympicArena, ranked by editorial score.
OE
OpenAI Evals
Evaluation · OpenAI GPT models (extensible)
8.1
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.
Free· Free (MIT); you pay OpenAI API costs for eval runsllm-benchmarkingregression-testing
VI
VisualWebArena
Evaluation · Model-agnostic (GPT-4V, Gemini, Claude, open VLMs)
7.1
Open benchmark for evaluating multimodal web agents on realistic visual browsing tasks.
Free· Free and open source (MIT-style research release)multimodal-agent-evalweb-browsing-benchmark
AA
Arena AI
Evaluation · Multi-model
6.8
Head-to-head LLM battle arena with a public leaderboard for ranking AI models.
Free· Free to use; no public paid tier listedllm-benchmarkingmodel-comparison
MA
MathEval
Evaluation · GPT-4 grader / DeepSeek-LLM-7B verifier
7.3
Holistic benchmark suite for evaluating mathematical reasoning in large language models.
Free· Free; open-source benchmark with leaderboard submissions via matheval.aillm-math-benchmarkingmodel-leaderboards
LI
LiveBench
Evaluation · Multi-model
8.2
Contamination-free LLM benchmark that refreshes its questions monthly to keep frontier models honest.
Free· Free and open source; self-hosted evaluation runnerllm-benchmarkingmodel-selection
LL
LLMEval
Evaluation · Multi-model
7.2
Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.
Free· Free; open-source academic benchmarksllm-benchmarkingacademic-evaluation