Best AI tools for llm observability
42 tools in the Evaluation category, filtered to llm observability.

LangSmith
LangChain's eval + observability platform.

Weights & Biases
The ML experiment tracker, now with LLM eval features.
Helicone
Open-source LLM observability — one-line proxy install.

Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.

LiveBench
Contamination-free LLM benchmark that refreshes its questions monthly to keep frontier models honest.

Athina AI
Collaborative LLM evaluation and observability platform for teams shipping AI features to production.

HoneyHive
OpenTelemetry-native observability and evaluation platform for LLM agents in production.

MLflow
Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.

OpenAI Evals
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.

TruLens
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.

W&B Weave
Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

Langfuse
Open-source LLM observability, prompt management, and evaluation in one platform.

MathEval
Holistic benchmark suite for evaluating mathematical reasoning in large language models.

Opik
Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Inspect AI
Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.

Kiln AI
Open-source workbench for building, evaluating, and fine-tuning AI agents across 190+ models.

LLMEval
Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.

Promptfoo
Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.

AlpacaEval
Automatic LLM evaluator and leaderboard that benchmarks instruction-following with length-controlled win rates.

Arthur
Open-source toolkit for testing, tracing, and monitoring production AI agents.

Fiddler AI
Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.

Maxim AI
End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.

Respan (formerly Keywords AI)
LLM engineering platform combining a multi-model gateway with tracing, evals, and prompt management.

llmfit
Terminal tool that scores hundreds of open LLMs against your actual CPU, RAM, and GPU and tells you which ones will run well.

OlympicArena
Olympiad-level multi-discipline benchmark for stress-testing reasoning in LLMs and multimodal models.

Phoenix
Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.

Superwise
Agentic management platform for runtime guardrails, policy enforcement, and observability across LLM agents.

Agenta
Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.

CompassRank
Public leaderboard from the OpenCompass project ranking open and closed LLMs across 100+ benchmarks.

InfiBench
Stack Overflow-derived benchmark for evaluating code LLMs on real-world programming questions.

MixEval
Dynamic LLM benchmark that mixes web queries with existing datasets to mirror Chatbot Arena rankings at a fraction of the cost.

Arena AI
Head-to-head LLM battle arena with a public leaderboard for ranking AI models.

Parea AI
LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.

AI Meter
Local usage meter that turns AI coding-agent tokens into estimated electricity and water consumption.

Lagotto Meter
Measure the gap between what your site claims and what an AI agent actually finds

Lakera
Runtime security and guardrails for GenAI apps, agents, and RAG systems.

LangWatch
Simulation-based testing, evaluation, and observability for LLM agents

LLM GPU Checker (KO)
Match LLMs to GPUs and plan multi-model AI stacks by VRAM, bandwidth and precision.

ModelFuzz
Open-source red-teaming and execution-layer defense for AI agents against prompt injection.

Netron
Visualizer for neural network, deep learning, and machine learning models

QuantProbe
Physics-based calculator that predicts LLM decode speed, memory fit, and quantization quality on any hardware.