Skip to main content
📖 The AI Tool Bible

Evaluation

Observability, prompt testing, and quality scoring.

59 tools

Why it matters

Evaluation is the discipline most underinvested in by AI product teams. Choosing an eval tool early is much cheaper than retrofitting one when an LLM regression hits production.

What's in here

Spans full eval + observability platforms (Braintrust, LangSmith), prompt management (Humanloop, PromptLayer), ML-broad tracking with LLM features (Weights & Biases), and proxy-based observability (Helicone).

How to pick

Pick Braintrust or LangSmith for full eval + observability. Pick Humanloop if PMs need to edit prompts. Pick Helicone for a one-line install on existing OpenAI/Claude code. Pick Patronus for automated hallucination/safety evals at scale.

Braintrust preview image
Braintrust logo

Braintrust

Featured
Evaluation · Platform (any LLM)
8.9

Eval, monitor, and improve AI products end-to-end.

Freemium· Starter: $0 · Pro: $249 · Enterprise: Custom pricingevalsmonitoring
LangSmith preview image
LangSmith logo

LangSmith

Evaluation · Platform (any LLM)
8.7

LangChain's eval + observability platform.

Freemium· Developer: $0 · Plus: $39 · Enterprise: Custom pricingLLM tracingevals
Weights & Biases preview image
Weights & Biases logo

Weights & Biases

Evaluation · Platform (any LLM)
8.4

The ML experiment tracker, now with LLM eval features.

Freemium· Free: $0/mo · Pro: Starts at $60/month · Enterprise: Custom plans · Personal: $0/mo · Advanced Enterprise: Custom planML experimentsLLM eval
Helicone preview image
Helicone logo

Helicone

Evaluation · Platform (any LLM)
8.3

Open-source LLM observability — one-line proxy install.

Freemium· Free 100k req/mo; Pro from $25/moobservabilitycost tracking
Arize AI preview image
Arize AI logo

Arize AI

Evaluation · Multi-model
8.2

Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-observabilityagent-evaluation
Giskard preview image
Giskard logo

Giskard

Evaluation · Multi-model
8.2

Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.

Freemium· Open-source free tier; Giskard Hub enterprise pricing on requestllm-red-teamingagent-security-testing
Great Expectations preview image
Great Expectations logo

Great Expectations

Evaluation
8.2

Open-source data quality framework for validating the datasets that feed your ML and analytics pipelines.

Freemium· Developer: Free · Team: Custom · Enterprise: Contact Salesdata-validationpipeline-testing
Humanloop preview image
Humanloop logo

Humanloop

Evaluation · Platform (any LLM)
8.2

Prompt management + evals for collaborative AI teams.

Paid· From $200/mo teamprompt managementteam collab
LiveBench preview image
LiveBench logo

LiveBench

Evaluation · Multi-model
8.2

Contamination-free LLM benchmark that refreshes its questions monthly to keep frontier models honest.

Free· Free and open source; self-hosted evaluation runnerllm-benchmarkingmodel-selection
Athina AI preview image
Athina AI logo

Athina AI

Evaluation · Multi-model
8.1

Collaborative LLM evaluation and observability platform for teams shipping AI features to production.

Freemium· Starter free (10k logs/mo); Pro & Enterprise customllm-evaluationprompt-management
Berkeley Function-Calling Leaderboard preview image
Berkeley Function-Calling Leaderboard logo

Berkeley Function-Calling Leaderboard

Evaluation · Multi-model
8.1

Open benchmark from UC Berkeley that ranks LLMs on real-world tool-use and function-calling accuracy.

Free· Free and open source; you pay only for inference when reproducing runs.function-calling evaltool-use benchmarking
HoneyHive preview image
HoneyHive logo

HoneyHive

Evaluation · Multi-model
8.1

OpenTelemetry-native observability and evaluation platform for LLM agents in production.

Freemium· Free tier available; paid/enterprise tiers via salesagent-observabilityllm-evaluation
MLflow preview image
MLflow logo

MLflow

Evaluation · Multi-model
8.1

Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.

Free· Free and open source (Apache 2.0); managed offering via Databricksllm-evaluationexperiment-tracking
OpenAI Evals preview image
OpenAI Evals logo

OpenAI Evals

Evaluation · OpenAI GPT models (extensible)
8.1

OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.

Free· Free (MIT); you pay OpenAI API costs for eval runsllm-benchmarkingregression-testing
TruLens preview image
TruLens logo

TruLens

Evaluation · Multi-model (LLM-as-judge)
8.1

Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.

Free· Free, open source (Apache-licensed Python package)llm-evaluationrag-evaluation
W&B Weave preview image
W&B Weave logo

W&B Weave

Evaluation · Multi-model
8.1

Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

Freemium· Free tier available; paid and enterprise plans via W&Bllm-tracingagent-observability
LLM Stats preview image
LLM Stats logo

LLM Stats

Evaluation · Multi-model
7.9

Live leaderboard and side-by-side comparison hub for 300+ frontier LLMs across reasoning, coding, and multimodal benchmarks.

Free· Free to browse; underlying model usage billed by each providermodel-comparisonbenchmark-tracking
PromptLayer preview image
PromptLayer logo

PromptLayer

Evaluation · Platform (any LLM)
7.9

Lightweight prompt logging + management for OpenAI/Claude apps.

Freemium· Free: $0 · Pro: $49 · Team: $500 · Enterprise: Customprompt loggingversioning
Patronus preview image
Patronus logo

Patronus

Evaluation · Platform (any LLM)
7.8

Automated LLM evaluation for hallucinations, safety, and quality.

Paid· Individual: Free · Base: $25 · Enterprise: Contact us for Pricinghallucination detectionsafety
Langfuse preview image
Langfuse logo

Langfuse

Evaluation · Model-agnostic
7.3

Open-source LLM observability, prompt management, and evaluation in one platform.

Freemium· Free self-host & Hobby tier; Core $29/mo, Pro $199/mo, Enterprise $2,499/mollm-observabilityprompt-management
MathEval preview image
MathEval logo

MathEval

Evaluation · GPT-4 grader / DeepSeek-LLM-7B verifier
7.3

Holistic benchmark suite for evaluating mathematical reasoning in large language models.

Free· Free; open-source benchmark with leaderboard submissions via matheval.aillm-math-benchmarkingmodel-leaderboards
Opik preview image
Opik logo

Opik

Evaluation · Multi-model
7.3

Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Freemium· Free open-source self-host; free Cloud tier (no card); Enterprise contact salesllm-tracingagent-evaluation
Inspect AI preview image
Inspect AI logo

Inspect AI

Evaluation · Multi-model
7.2

Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.

Free· Free and open source (MIT-style license); you pay only for underlying model API usage.llm-benchmarkingagent-evaluation
Kiln AI preview image
Kiln AI logo

Kiln AI

Evaluation · Multi-model
7.2

Open-source workbench for building, evaluating, and fine-tuning AI agents across 190+ models.

Freemium· Free Individual tier; Team (request access); Enterprise (custom)llm-evaluationfine-tuning