AI eval dataset builder
Editorial picks for "llm eval datasets".
32 tools
Braintrust
FeaturedEval, monitor, and improve AI products end-to-end.
LangSmith
LangChain's eval + observability platform.
Weights & Biases
The ML experiment tracker, now with LLM eval features.
Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.
Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.
Great Expectations
Open-source data quality framework for validating the datasets that feed your ML and analytics pipelines.
Humanloop
Prompt management + evals for collaborative AI teams.
LiveBench
Contamination-free LLM benchmark that refreshes its questions monthly to keep frontier models honest.
Athina AI
Collaborative LLM evaluation and observability platform for teams shipping AI features to production.
Berkeley Function-Calling Leaderboard
Open benchmark from UC Berkeley that ranks LLMs on real-world tool-use and function-calling accuracy.
OpenAI Evals
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.
TruLens
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.
CAMEL-AI
Open-source Python framework for building multi-agent systems and synthetic data pipelines.
PromptHub
Git-style prompt management, testing, and deployment platform for teams running multiple LLMs in production.
Patronus
Automated LLM evaluation for hallucinations, safety, and quality.
MathEval
Holistic benchmark suite for evaluating mathematical reasoning in large language models.
Inspect AI
Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.
Kiln AI
Open-source workbench for building, evaluating, and fine-tuning AI agents across 190+ models.
LLMEval
Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.
Promptfoo
Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.
Weco AI
Autoresearch engine that iteratively rewrites code to optimize against a numeric evaluation metric.
AlpacaEval
Automatic LLM evaluator and leaderboard that benchmarks instruction-following with length-controlled win rates.
Maxim AI
End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.
SEAL Leaderboard
Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.
VisualWebArena
Open benchmark for evaluating multimodal web agents on realistic visual browsing tasks.
OlympicArena
Olympiad-level multi-discipline benchmark for stress-testing reasoning in LLMs and multimodal models.
Agenta
Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.
CompassRank
Public leaderboard from the OpenCompass project ranking open and closed LLMs across 100+ benchmarks.
InfiBench
Stack Overflow-derived benchmark for evaluating code LLMs on real-world programming questions.
MixEval
Dynamic LLM benchmark that mixes web queries with existing datasets to mirror Chatbot Arena rankings at a fraction of the cost.
Parea AI
LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.
ModelBias
100 models, 100 prompts, 30,000 answers — an interactive look at AI defaults