Skip to main content
📖 The AI Tool Bible

Best AI tools for evals datasets

48 tools in the Evaluation category, filtered to evals datasets.

All Evaluation →
BR

Braintrust

Featured
Evaluation · Platform (any LLM)
8.9

Eval, monitor, and improve AI products end-to-end.

Freemium· Starter: $0 · Pro: $249 · Enterprise: Custom pricingevalsmonitoring
LA

LangSmith

Evaluation · Platform (any LLM)
8.7

LangChain's eval + observability platform.

Freemium· Developer: $0 · Plus: $39 · Enterprise: Custom pricingLLM tracingevals
WB

Weights & Biases

Evaluation · Platform (any LLM)
8.4

The ML experiment tracker, now with LLM eval features.

Freemium· Free: $0/mo · Pro: Starts at $60/month · Enterprise: Custom plans · Personal: $0/mo · Advanced Enterprise: Custom planML experimentsLLM eval
AA

Arize AI

Evaluation · Multi-model
8.2

Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-observabilityagent-evaluation
GI

Giskard

Evaluation · Multi-model
8.2

Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.

Freemium· Open-source free tier; Giskard Hub enterprise pricing on requestllm-red-teamingagent-security-testing
GE

Great Expectations

Evaluation
8.2

Open-source data quality framework for validating the datasets that feed your ML and analytics pipelines.

Freemium· Developer: Free · Team: Custom · Enterprise: Contact Salesdata-validationpipeline-testing
HU

Humanloop

Evaluation · Platform (any LLM)
8.2

Prompt management + evals for collaborative AI teams.

Paid· From $200/mo teamprompt managementteam collab
LI

LiveBench

Evaluation · Multi-model
8.2

Contamination-free LLM benchmark that refreshes its questions monthly to keep frontier models honest.

Free· Free and open source; self-hosted evaluation runnerllm-benchmarkingmodel-selection
AA

Athina AI

Evaluation · Multi-model
8.1

Collaborative LLM evaluation and observability platform for teams shipping AI features to production.

Freemium· Starter free (10k logs/mo); Pro & Enterprise customllm-evaluationprompt-management
BF

Berkeley Function-Calling Leaderboard

Evaluation · Multi-model
8.1

Open benchmark from UC Berkeley that ranks LLMs on real-world tool-use and function-calling accuracy.

Free· Free and open source; you pay only for inference when reproducing runs.function-calling evaltool-use benchmarking
HO

HoneyHive

Evaluation · Multi-model
8.1

OpenTelemetry-native observability and evaluation platform for LLM agents in production.

Freemium· Free tier available; paid/enterprise tiers via salesagent-observabilityllm-evaluation
ML

MLflow

Evaluation · Multi-model
8.1

Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.

Free· Free and open source (Apache 2.0); managed offering via Databricksllm-evaluationexperiment-tracking
OE

OpenAI Evals

Evaluation · OpenAI GPT models (extensible)
8.1

OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.

Free· Free (MIT); you pay OpenAI API costs for eval runsllm-benchmarkingregression-testing
TR

TruLens

Evaluation · Multi-model (LLM-as-judge)
8.1

Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.

Free· Free, open source (Apache-licensed Python package)llm-evaluationrag-evaluation
WB

W&B Weave

Evaluation · Multi-model
8.1

Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

Freemium· Free tier available; paid and enterprise plans via W&Bllm-tracingagent-observability
PR

PromptHub

Writing · Multi-model (OpenAI, Anthropic, Google, Meta, Mistral, Bedrock, Azure)
8.0

Git-style prompt management, testing, and deployment platform for teams running multiple LLMs in production.

Freemium· Free signup; paid team plans (contact sales / in-app)prompt-managementprompt-versioning
LS

LLM Stats

Evaluation · Multi-model
7.9

Live leaderboard and side-by-side comparison hub for 300+ frontier LLMs across reasoning, coding, and multimodal benchmarks.

Free· Free to browse; underlying model usage billed by each providermodel-comparisonbenchmark-tracking
PA

Patronus

Evaluation · Platform (any LLM)
7.8

Automated LLM evaluation for hallucinations, safety, and quality.

Paid· Individual: Free · Base: $25 · Enterprise: Contact us for Pricinghallucination detectionsafety
LA

Langfuse

Evaluation · Model-agnostic
7.3

Open-source LLM observability, prompt management, and evaluation in one platform.

Freemium· Free self-host & Hobby tier; Core $29/mo, Pro $199/mo, Enterprise $2,499/mollm-observabilityprompt-management
MA

MathEval

Evaluation · GPT-4 grader / DeepSeek-LLM-7B verifier
7.3

Holistic benchmark suite for evaluating mathematical reasoning in large language models.

Free· Free; open-source benchmark with leaderboard submissions via matheval.aillm-math-benchmarkingmodel-leaderboards
OP

Opik

Evaluation · Multi-model
7.3

Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Freemium· Free open-source self-host; free Cloud tier (no card); Enterprise contact salesllm-tracingagent-evaluation
IA

Inspect AI

Evaluation · Multi-model
7.2

Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.

Free· Free and open source (MIT-style license); you pay only for underlying model API usage.llm-benchmarkingagent-evaluation
KA

Kiln AI

Evaluation · Multi-model
7.2

Open-source workbench for building, evaluating, and fine-tuning AI agents across 190+ models.

Freemium· Free Individual tier; Team (request access); Enterprise (custom)llm-evaluationfine-tuning
LL

LLMEval

Evaluation · Multi-model
7.2

Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.

Free· Free; open-source academic benchmarksllm-benchmarkingacademic-evaluation
PR

Promptfoo

Evaluation · Multi-model
7.2

Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.

Freemium· Community: Free · Enterprise: Custom · On-Premise: Customllm-evalsred-teaming
WA

Weco AI

Evaluation · Multi-model (LLM + AIDE tree search)
7.2

Autoresearch engine that iteratively rewrites code to optimize against a numeric evaluation metric.

Freemium· Open-source CLI; hosted/commercial pricing not publishedcode-optimizationgpu-kernel-tuning
AL

AlpacaEval

Evaluation · GPT-4 Preview (Nov 2024) as annotator
7.1

Automatic LLM evaluator and leaderboard that benchmarks instruction-following with length-controlled win rates.

Free· Free and open-source; pay only for the underlying OpenAI annotator API callsllm-benchmarkinginstruction-following eval
MA

Maxim AI

Evaluation · Multi-model
7.1

End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.

Freemium· Developer: Free · Professional: $29 /seat /month · Business: $49 /seat /month · Enterprise: Customagent-evaluationllm-observability
PF

Prompt Foundry

Evaluation · OpenAI + Anthropic (multi-model)
7.1

Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.

Freemium· Free tier (10 prompts, 500 evals/mo); Pro $15/user/mo; Enterprise customprompt-managementmodel-comparison
RF

Respan (formerly Keywords AI)

Evaluation · Multi-model (500+ via gateway)
7.1

LLM engineering platform combining a multi-model gateway with tracing, evals, and prompt management.

Freemium· Free tier; paid plans (pricing not public); enterprise on requestllm-observabilityprompt-management
SL

SEAL Leaderboard

Evaluation · Multi-model (GPT, Claude, Gemini, Llama, etc.)
7.1

Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.

Free· Free to view; paid custom evals via Scale enterprise salesmodel-selectionbenchmark-tracking
VI

VisualWebArena

Evaluation · Model-agnostic (GPT-4V, Gemini, Claude, open VLMs)
7.1

Open benchmark for evaluating multimodal web agents on realistic visual browsing tasks.

Free· Free and open source (MIT-style research release)multimodal-agent-evalweb-browsing-benchmark
LL

llmfit

Evaluation · Multi-model
7.0

Terminal tool that scores hundreds of open LLMs against your actual CPU, RAM, and GPU and tells you which ones will run well.

Free· Free, MIT-licensedlocal-llm-selectionhardware-benchmarking
OL

OlympicArena

Evaluation · GPT-4o, Claude-3.5-Sonnet, Doubao-Pro-32k, DeepSeek-Coder-V2, Qwen2-72B-Instruct
7.0

Olympiad-level multi-discipline benchmark for stress-testing reasoning in LLMs and multimodal models.

Free· Free, open-source research benchmarkllm-evaluationmultimodal-eval
PH

Phoenix

Evaluation · Multi-model
7.0

Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-tracingagent-debugging
AG

Agenta

Evaluation · Multi-model
6.9

Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.

Freemium· Hobby: $0 forever · Pro: $29 /month · Business: $299 /month · Enterprise: Custom · Open source: Free foreverprompt-engineeringllm-evaluation
CO

CompassRank

Evaluation · Multi-model
6.9

Public leaderboard from the OpenCompass project ranking open and closed LLMs across 100+ benchmarks.

Free· Free leaderboard; OpenCompass toolkit is Apache 2.0 open sourcellm-benchmarkingmodel-selection
IN

InfiBench

Evaluation
6.9

Stack Overflow-derived benchmark for evaluating code LLMs on real-world programming questions.

Free· Free and open source (CC BY-SA 4.0)code-llm-evalmodel-benchmarking
MI

MixEval

Evaluation · GPT-3.5-Turbo-0125, GPT-4o-2024-05-13, Claude 3.5 Sonnet, MixEval, MixEval-Hard
6.9

Dynamic LLM benchmark that mixes web queries with existing datasets to mirror Chatbot Arena rankings at a fraction of the cost.

Free· Free and open sourcellm-benchmarkingmodel-ranking
AA

Arena AI

Evaluation · Multi-model
6.8

Head-to-head LLM battle arena with a public leaderboard for ranking AI models.

Free· Free to use; no public paid tier listedllm-benchmarkingmodel-comparison
AA

Artificial Analysis

Evaluation · Multi-model
6.8

Independent benchmarking platform comparing AI models and inference providers across intelligence, speed, and cost.

Freemium· Pro: $417/month per seat · Enterprise: Custom pricingmodel-benchmarkingprovider-comparison
PA

Parea AI

Evaluation · Multi-model
6.8

LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.

Freemium· Free (2 seats, 3k logs/mo); Team $150/mo; Enterprise customllm-evaluationprompt-management
AW

AI World Bakeoff

Evaluation · Ten frontier coding models (specific list not published on landing page; includes at least one Claude Opus generation referenced as 'Opus 5')

Ten AI coding models, three identical briefs, thirty explorable 3D worlds

Free· Free to view. No paid tiers, sign-up, or accounts.One-shot AI coding model comparison3D generative code benchmarking
LM

Lagotto Meter

Evaluation

Measure the gap between what your site claims and what an AI agent actually finds

Free· Free (currently in public beta)AI-visibility audit of a landing pageFact-vs-claim gap analysis for marketing copy
LA

LangWatch

Evaluation · Model-agnostic; supports OpenAI, Anthropic, AWS Bedrock, Azure OpenAI, Vertex AI, and any OpenTelemetry-instrumented LLM

Simulation-based testing, evaluation, and observability for LLM agents

Freemium· Developer: €0 · Growth: €29/ core-seat / month · Enterprise: CustomLLM agent regression testing in CIRAG answer-quality evaluation
LG

LLM GPU Checker (KO)

Evaluation · Catalog covers open models on Hugging Face (Llama, Qwen, Mistral, Gemma, etc.)

Match LLMs to GPUs and plan multi-model AI stacks by VRAM, bandwidth and precision.

Free· Free (open web tool hosted on GitHub Pages).GPU sizing for self-hosted LLMsMulti-GPU RAG stack planning
MO

ModelBias

Evaluation · 100 models across Anthropic, OpenAI, Google, DeepSeek, Meta, xAI, Mistral, Qwen and others (via OpenRouter)

100 models, 100 prompts, 30,000 answers — an interactive look at AI defaults

Free· Free to browse and download the full dataset from GitHub.Comparing default model preferences across vendorsIllustrating RLHF homogenisation in talks and articles
MO

ModelFuzz

Evaluation · Model-agnostic; works with any OpenAI-compatible endpoint (Qwen 2.5 used in official examples).

Open-source red-teaming and execution-layer defense for AI agents against prompt injection.

Freemium· Free / open-source (MIT) via pip. Hosted dashboard with centralized policies, audit logs and continuous scanning coming soon via waitlist (pricing not yet public).Red-teaming OpenAI-compatible agent endpointsBlocking indirect prompt injection via retrieved documents