Evaluation
Observability, prompt testing, and quality scoring.
59 tools
Evaluation is the discipline most underinvested in by AI product teams. Choosing an eval tool early is much cheaper than retrofitting one when an LLM regression hits production.
Spans full eval + observability platforms (Braintrust, LangSmith), prompt management (Humanloop, PromptLayer), ML-broad tracking with LLM features (Weights & Biases), and proxy-based observability (Helicone).
Pick Braintrust or LangSmith for full eval + observability. Pick Humanloop if PMs need to edit prompts. Pick Helicone for a one-line install on existing OpenAI/Claude code. Pick Patronus for automated hallucination/safety evals at scale.

AI Meter
Local usage meter that turns AI coding-agent tokens into estimated electricity and water consumption.

AI World Bakeoff
Ten AI coding models, three identical briefs, thirty explorable 3D worlds

Habibi
Self-hosted generative engine optimization for tracking brand mentions across ChatGPT, Perplexity, and Gemini

Lagotto Meter
Measure the gap between what your site claims and what an AI agent actually finds

Lakera
Runtime security and guardrails for GenAI apps, agents, and RAG systems.

LangWatch
Simulation-based testing, evaluation, and observability for LLM agents

LLM GPU Checker (KO)
Match LLMs to GPUs and plan multi-model AI stacks by VRAM, bandwidth and precision.

ModelBias
100 models, 100 prompts, 30,000 answers — an interactive look at AI defaults

ModelFuzz
Open-source red-teaming and execution-layer defense for AI agents against prompt injection.

Netron
Visualizer for neural network, deep learning, and machine learning models

QuantProbe
Physics-based calculator that predicts LLM decode speed, memory fit, and quantization quality on any hardware.