Best AI tools for evals datasets
48 tools in the Evaluation category, filtered to evals datasets.
Braintrust
FeaturedEval, monitor, and improve AI products end-to-end.
LangSmith
LangChain's eval + observability platform.
Weights & Biases
The ML experiment tracker, now with LLM eval features.
Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.
Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.
Great Expectations
Open-source data quality framework for validating the datasets that feed your ML and analytics pipelines.
Humanloop
Prompt management + evals for collaborative AI teams.
LiveBench
Contamination-free LLM benchmark that refreshes its questions monthly to keep frontier models honest.
Athina AI
Collaborative LLM evaluation and observability platform for teams shipping AI features to production.
Berkeley Function-Calling Leaderboard
Open benchmark from UC Berkeley that ranks LLMs on real-world tool-use and function-calling accuracy.
HoneyHive
OpenTelemetry-native observability and evaluation platform for LLM agents in production.
MLflow
Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.
OpenAI Evals
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.
TruLens
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.
W&B Weave
Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.
PromptHub
Git-style prompt management, testing, and deployment platform for teams running multiple LLMs in production.
LLM Stats
Live leaderboard and side-by-side comparison hub for 300+ frontier LLMs across reasoning, coding, and multimodal benchmarks.
Patronus
Automated LLM evaluation for hallucinations, safety, and quality.
Langfuse
Open-source LLM observability, prompt management, and evaluation in one platform.
MathEval
Holistic benchmark suite for evaluating mathematical reasoning in large language models.
Opik
Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.
Inspect AI
Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.
Kiln AI
Open-source workbench for building, evaluating, and fine-tuning AI agents across 190+ models.
LLMEval
Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.
Promptfoo
Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.
Weco AI
Autoresearch engine that iteratively rewrites code to optimize against a numeric evaluation metric.
AlpacaEval
Automatic LLM evaluator and leaderboard that benchmarks instruction-following with length-controlled win rates.
Maxim AI
End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.
Prompt Foundry
Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.
Respan (formerly Keywords AI)
LLM engineering platform combining a multi-model gateway with tracing, evals, and prompt management.
SEAL Leaderboard
Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.
VisualWebArena
Open benchmark for evaluating multimodal web agents on realistic visual browsing tasks.
llmfit
Terminal tool that scores hundreds of open LLMs against your actual CPU, RAM, and GPU and tells you which ones will run well.
OlympicArena
Olympiad-level multi-discipline benchmark for stress-testing reasoning in LLMs and multimodal models.
Phoenix
Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.
Agenta
Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.
CompassRank
Public leaderboard from the OpenCompass project ranking open and closed LLMs across 100+ benchmarks.
InfiBench
Stack Overflow-derived benchmark for evaluating code LLMs on real-world programming questions.
MixEval
Dynamic LLM benchmark that mixes web queries with existing datasets to mirror Chatbot Arena rankings at a fraction of the cost.
Arena AI
Head-to-head LLM battle arena with a public leaderboard for ranking AI models.
Artificial Analysis
Independent benchmarking platform comparing AI models and inference providers across intelligence, speed, and cost.
Parea AI
LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.
AI World Bakeoff
Ten AI coding models, three identical briefs, thirty explorable 3D worlds
Lagotto Meter
Measure the gap between what your site claims and what an AI agent actually finds
LangWatch
Simulation-based testing, evaluation, and observability for LLM agents
LLM GPU Checker (KO)
Match LLMs to GPUs and plan multi-model AI stacks by VRAM, bandwidth and precision.
ModelBias
100 models, 100 prompts, 30,000 answers — an interactive look at AI defaults
ModelFuzz
Open-source red-teaming and execution-layer defense for AI agents against prompt injection.