Skip to main content
📖 The AI Tool Bible

Evaluation

Observability, prompt testing, and quality scoring.

59 tools

Why it matters

Evaluation is the discipline most underinvested in by AI product teams. Choosing an eval tool early is much cheaper than retrofitting one when an LLM regression hits production.

What's in here

Spans full eval + observability platforms (Braintrust, LangSmith), prompt management (Humanloop, PromptLayer), ML-broad tracking with LLM features (Weights & Biases), and proxy-based observability (Helicone).

How to pick

Pick Braintrust or LangSmith for full eval + observability. Pick Humanloop if PMs need to edit prompts. Pick Helicone for a one-line install on existing OpenAI/Claude code. Pick Patronus for automated hallucination/safety evals at scale.

LLMEval preview image
LLMEval logo

LLMEval

Evaluation · Multi-model
7.2

Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.

Free· Free; open-source academic benchmarksllm-benchmarkingacademic-evaluation
Promptfoo preview image
Promptfoo logo

Promptfoo

Evaluation · Multi-model
7.2

Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.

Freemium· Community: Free · Enterprise: Custom · On-Premise: Customllm-evalsred-teaming
Weco AI preview image
Weco AI logo

Weco AI

Evaluation · Multi-model (LLM + AIDE tree search)
7.2

Autoresearch engine that iteratively rewrites code to optimize against a numeric evaluation metric.

Freemium· Open-source CLI; hosted/commercial pricing not publishedcode-optimizationgpu-kernel-tuning
AlpacaEval preview image
AlpacaEval logo

AlpacaEval

Evaluation · GPT-4 Preview (Nov 2024) as annotator
7.1

Automatic LLM evaluator and leaderboard that benchmarks instruction-following with length-controlled win rates.

Free· Free and open-source; pay only for the underlying OpenAI annotator API callsllm-benchmarkinginstruction-following eval
Arthur preview image
Arthur logo

Arthur

Evaluation · Multi-model
7.1

Open-source toolkit for testing, tracing, and monitoring production AI agents.

Freemium· Free: $0/mo · Premium: $60/mo · Enterprise: Customagent-evaluationprompt-management
Fiddler AI preview image
Fiddler AI logo

Fiddler AI

Evaluation · Fiddler Centor (proprietary evaluators)
7.1

Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.

Enterprise· Free: Free · Developer: $0.002 per trace · Enterprise: Contact salesllm-observabilityagent-monitoring
Maxim AI preview image
Maxim AI logo

Maxim AI

Evaluation · Multi-model
7.1

End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.

Freemium· Developer: Free · Professional: $29 /seat /month · Business: $49 /seat /month · Enterprise: Customagent-evaluationllm-observability
Prompt Foundry preview image
Prompt Foundry logo

Prompt Foundry

Evaluation · OpenAI + Anthropic (multi-model)
7.1

Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.

Freemium· Free tier (10 prompts, 500 evals/mo); Pro $15/user/mo; Enterprise customprompt-managementmodel-comparison
Respan (formerly Keywords AI) preview image
Respan (formerly Keywords AI) logo

Respan (formerly Keywords AI)

Evaluation · Multi-model (500+ via gateway)
7.1

LLM engineering platform combining a multi-model gateway with tracing, evals, and prompt management.

Freemium· Free tier; paid plans (pricing not public); enterprise on requestllm-observabilityprompt-management
SEAL Leaderboard preview image
SEAL Leaderboard logo

SEAL Leaderboard

Evaluation · Multi-model (GPT, Claude, Gemini, Llama, etc.)
7.1

Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.

Free· Free to view; paid custom evals via Scale enterprise salesmodel-selectionbenchmark-tracking
VisualWebArena preview image
VisualWebArena logo

VisualWebArena

Evaluation · Model-agnostic (GPT-4V, Gemini, Claude, open VLMs)
7.1

Open benchmark for evaluating multimodal web agents on realistic visual browsing tasks.

Free· Free and open source (MIT-style research release)multimodal-agent-evalweb-browsing-benchmark
LangFast preview image
LangFast logo

LangFast

Evaluation · Multi-model
7.0

No-signup LLM playground for testing, comparing, and versioning prompts against your own API keys.

Paid· One-time lifetime ~$60-$120; 14-day money-backprompt-testingprompt-versioning
llmfit preview image
llmfit logo

llmfit

Evaluation · Multi-model
7.0

Terminal tool that scores hundreds of open LLMs against your actual CPU, RAM, and GPU and tells you which ones will run well.

Free· Free, MIT-licensedlocal-llm-selectionhardware-benchmarking
OlympicArena preview image
OlympicArena logo

OlympicArena

Evaluation · GPT-4o, Claude-3.5-Sonnet, Doubao-Pro-32k, DeepSeek-Coder-V2, Qwen2-72B-Instruct
7.0

Olympiad-level multi-discipline benchmark for stress-testing reasoning in LLMs and multimodal models.

Free· Free, open-source research benchmarkllm-evaluationmultimodal-eval
Phoenix preview image
Phoenix logo

Phoenix

Evaluation · Multi-model
7.0

Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-tracingagent-debugging
Superwise preview image
Superwise logo

Superwise

Evaluation · Multi-model
7.0

Agentic management platform for runtime guardrails, policy enforcement, and observability across LLM agents.

Freemium· Starter: Free · Pro+: ? · Enterprise: ?llm-guardrailsai-governance
Agenta preview image
Agenta logo

Agenta

Evaluation · Multi-model
6.9

Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.

Freemium· Hobby: $0 forever · Pro: $29 /month · Business: $299 /month · Enterprise: Custom · Open source: Free foreverprompt-engineeringllm-evaluation
CompassRank preview image
CompassRank logo

CompassRank

Evaluation · Multi-model
6.9

Public leaderboard from the OpenCompass project ranking open and closed LLMs across 100+ benchmarks.

Free· Free leaderboard; OpenCompass toolkit is Apache 2.0 open sourcellm-benchmarkingmodel-selection
InfiBench preview image
InfiBench logo

InfiBench

Evaluation
6.9

Stack Overflow-derived benchmark for evaluating code LLMs on real-world programming questions.

Free· Free and open source (CC BY-SA 4.0)code-llm-evalmodel-benchmarking
MixEval preview image
MixEval logo

MixEval

Evaluation · GPT-3.5-Turbo-0125, GPT-4o-2024-05-13, Claude 3.5 Sonnet, MixEval, MixEval-Hard
6.9

Dynamic LLM benchmark that mixes web queries with existing datasets to mirror Chatbot Arena rankings at a fraction of the cost.

Free· Free and open sourcellm-benchmarkingmodel-ranking
Arena AI preview image
Arena AI logo

Arena AI

Evaluation · Multi-model
6.8

Head-to-head LLM battle arena with a public leaderboard for ranking AI models.

Free· Free to use; no public paid tier listedllm-benchmarkingmodel-comparison
Artificial Analysis preview image
Artificial Analysis logo

Artificial Analysis

Evaluation · Multi-model
6.8

Independent benchmarking platform comparing AI models and inference providers across intelligence, speed, and cost.

Freemium· Pro: $417/month per seat · Enterprise: Custom pricingmodel-benchmarkingprovider-comparison
Cleanlab TLM preview image
Cleanlab TLM logo

Cleanlab TLM

Evaluation · Multi-model (wraps any LLM)
6.8

Trustworthiness scoring layer that flags LLM hallucinations in real time.

Freemium· Free tier for evaluation; usage-based API pricing; enterprise/private deployment via saleshallucination-detectionrag-evaluation
Parea AI preview image
Parea AI logo

Parea AI

Evaluation · Multi-model
6.8

LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.

Freemium· Free (2 seats, 3k logs/mo); Team $150/mo; Enterprise customllm-evaluationprompt-management