Best AI tools for observability
19 tools in the Evaluation category, filtered to observability.
Braintrust
FeaturedEval, monitor, and improve AI products end-to-end.
LangSmith
LangChain's eval + observability platform.
Helicone
Open-source LLM observability — one-line proxy install.
AgentOps
Observability and debugging platform purpose-built for AI agents, with time-travel replay and cost tracking across 400+ LLMs.
Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.
HoneyHive
OpenTelemetry-native observability and evaluation platform for LLM agents in production.
TruLens
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.
W&B Weave
Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.
Langfuse
Open-source LLM observability, prompt management, and evaluation in one platform.
Opik
Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.
Arthur
Open-source toolkit for testing, tracing, and monitoring production AI agents.
Fiddler AI
Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.
Phoenix
Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.
Superwise
Agentic management platform for runtime guardrails, policy enforcement, and observability across LLM agents.
Agenta
Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.
ClickHouse
The open-source columnar database powering real-time analytics — and, increasingly, LLM observability and RAG backends.
Lagotto Meter
Measure the gap between what your site claims and what an AI agent actually finds
LangWatch
Simulation-based testing, evaluation, and observability for LLM agents
Netron
Visualizer for neural network, deep learning, and machine learning models