Skip to main content
📖 The AI Tool Bible

Evaluation

Observability, prompt testing, and quality scoring.

59 tools

Why it matters

Evaluation is the discipline most underinvested in by AI product teams. Choosing an eval tool early is much cheaper than retrofitting one when an LLM regression hits production.

What's in here

Spans full eval + observability platforms (Braintrust, LangSmith), prompt management (Humanloop, PromptLayer), ML-broad tracking with LLM features (Weights & Biases), and proxy-based observability (Helicone).

How to pick

Pick Braintrust or LangSmith for full eval + observability. Pick Humanloop if PMs need to edit prompts. Pick Helicone for a one-line install on existing OpenAI/Claude code. Pick Patronus for automated hallucination/safety evals at scale.

AI Meter preview image
AI Meter logo

AI Meter

Evaluation

Local usage meter that turns AI coding-agent tokens into estimated electricity and water consumption.

Free· Free for individuals and companies; open source under a public GitHub repo.Tracking daily token usage across Claude Code and CursorEstimating electricity draw of an AI-assisted coding session
AI World Bakeoff preview image
AI World Bakeoff logo

AI World Bakeoff

Evaluation · Ten frontier coding models (specific list not published on landing page; includes at least one Claude Opus generation referenced as 'Opus 5')

Ten AI coding models, three identical briefs, thirty explorable 3D worlds

Free· Free to view. No paid tiers, sign-up, or accounts.One-shot AI coding model comparison3D generative code benchmarking
Habibi preview image
Habibi logo

Habibi

Evaluation · GPT-4o, Perplexity Sonar, Gemini (bring-your-own API keys)

Self-hosted generative engine optimization for tracking brand mentions across ChatGPT, Perplexity, and Gemini

Freemium· Opsily Server: €40AI answer engine visibility trackingChatGPT brand mention monitoring
Lagotto Meter preview image
Lagotto Meter logo

Lagotto Meter

Evaluation

Measure the gap between what your site claims and what an AI agent actually finds

Free· Free (currently in public beta)AI-visibility audit of a landing pageFact-vs-claim gap analysis for marketing copy
Lakera preview image
Lakera logo

Lakera

Evaluation · Proprietary in-house classifiers; model-agnostic (works in front of GPT-4o, Claude, Gemini, Llama, and custom LLMs)

Runtime security and guardrails for GenAI apps, agents, and RAG systems.

Freemium· Free community/developer tier at platform.lakera.ai; paid Enterprise plans (custom pricing, contact sales). No public price list.Prompt injection defense for chatbotsRAG guardrails against indirect injection
LangWatch preview image
LangWatch logo

LangWatch

Evaluation · Model-agnostic; supports OpenAI, Anthropic, AWS Bedrock, Azure OpenAI, Vertex AI, and any OpenTelemetry-instrumented LLM

Simulation-based testing, evaluation, and observability for LLM agents

Freemium· Developer: €0 · Growth: €29/ core-seat / month · Enterprise: CustomLLM agent regression testing in CIRAG answer-quality evaluation
LLM GPU Checker (KO) preview image
LLM GPU Checker (KO) logo

LLM GPU Checker (KO)

Evaluation · Catalog covers open models on Hugging Face (Llama, Qwen, Mistral, Gemma, etc.)

Match LLMs to GPUs and plan multi-model AI stacks by VRAM, bandwidth and precision.

Free· Free (open web tool hosted on GitHub Pages).GPU sizing for self-hosted LLMsMulti-GPU RAG stack planning
ModelBias preview image
ModelBias logo

ModelBias

Evaluation · 100 models across Anthropic, OpenAI, Google, DeepSeek, Meta, xAI, Mistral, Qwen and others (via OpenRouter)

100 models, 100 prompts, 30,000 answers — an interactive look at AI defaults

Free· Free to browse and download the full dataset from GitHub.Comparing default model preferences across vendorsIllustrating RLHF homogenisation in talks and articles
ModelFuzz preview image
ModelFuzz logo

ModelFuzz

Evaluation · Model-agnostic; works with any OpenAI-compatible endpoint (Qwen 2.5 used in official examples).

Open-source red-teaming and execution-layer defense for AI agents against prompt injection.

Freemium· Free / open-source (MIT) via pip. Hosted dashboard with centralized policies, audit logs and continuous scanning coming soon via waitlist (pricing not yet public).Red-teaming OpenAI-compatible agent endpointsBlocking indirect prompt injection via retrieved documents
Netron preview image
Netron logo

Netron

Evaluation

Visualizer for neural network, deep learning, and machine learning models

Free· Free and open-source (MIT license). Available as a web app, desktop app (macOS/Linux/Windows), and Python package at no cost.ONNX model inspectionPyTorch checkpoint debugging
QuantProbe preview image
QuantProbe logo

QuantProbe

Evaluation

Physics-based calculator that predicts LLM decode speed, memory fit, and quantization quality on any hardware.

Free· Free and open source; install with `pip install quantprobe` or use the hosted web calculator.GPU and workstation sizing for local LLM inferenceQuantization tier selection (Q4/Q5/Q8) under a latency budget