LLM observability
Editorial picks for "best llm observability".
48 tools
Braintrust
FeaturedEval, monitor, and improve AI products end-to-end.
LangSmith
LangChain's eval + observability platform.
Weights & Biases
The ML experiment tracker, now with LLM eval features.
Helicone
Open-source LLM observability — one-line proxy install.
AgentOps
Observability and debugging platform purpose-built for AI agents, with time-travel replay and cost tracking across 400+ LLMs.
Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.
Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.
Humanloop
Prompt management + evals for collaborative AI teams.
Athina AI
Collaborative LLM evaluation and observability platform for teams shipping AI features to production.
HoneyHive
OpenTelemetry-native observability and evaluation platform for LLM agents in production.
MLflow
Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.
OpenAI Evals
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.
TruLens
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.
W&B Weave
Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.
PromptHub
Git-style prompt management, testing, and deployment platform for teams running multiple LLMs in production.
PromptLayer
Lightweight prompt logging + management for OpenAI/Claude apps.
Patronus
Automated LLM evaluation for hallucinations, safety, and quality.
Kong AI Gateway
Enterprise API gateway extended to route, govern, and observe LLM and agent traffic across providers.
Langfuse
Open-source LLM observability, prompt management, and evaluation in one platform.
Opik
Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.
Seldon
Kubernetes-native MLOps platform for deploying and orchestrating ML and generative AI models in production.
Inspect AI
Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.
Plano
Envoy-based data plane for AI agents that handles routing, guardrails, and observability outside your app code.
Portkey AI Gateway
Open-source AI gateway that routes a single API call across 1,600+ LLMs with caching, fallbacks, and observability.
Promptfoo
Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.
Wallaroo.AI
Production AI inference platform for deploying and monitoring models across cloud, on-prem, and edge.
Arthur
Open-source toolkit for testing, tracing, and monitoring production AI agents.
Fiddler AI
Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.
Maxim AI
End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.
Portkey
Production LLM gateway with observability, guardrails, and prompt management for teams shipping AI in anger.
Puzzlet AI
Git-native prompt management and observability platform for teams shipping LLM applications.
Respan (formerly Keywords AI)
LLM engineering platform combining a multi-model gateway with tracing, evals, and prompt management.
Phoenix
Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.
Superwise
Agentic management platform for runtime guardrails, policy enforcement, and observability across LLM agents.
SystemPrompt
Self-hosted AI governance gateway that audits, gates, and logs every LLM call before it leaves your network.
Agenta
Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.
Guild AI
Control plane for deploying, governing, and auditing AI agents in production.
TrueFoundry
Enterprise control plane for deploying, governing, and scaling agentic AI on your own infrastructure.
Cleanlab TLM
Trustworthiness scoring layer that flags LLM hallucinations in real time.
Parea AI
LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.
AI Meter
Local usage meter that turns AI coding-agent tokens into estimated electricity and water consumption.
ClickHouse
The open-source columnar database powering real-time analytics — and, increasingly, LLM observability and RAG backends.
Habibi
Self-hosted generative engine optimization for tracking brand mentions across ChatGPT, Perplexity, and Gemini
Hydra
Local-first trust control plane that routes AI tasks to the cheapest model that clears your confidence bar.
Lagotto Meter
Measure the gap between what your site claims and what an AI agent actually finds
Lakera
Runtime security and guardrails for GenAI apps, agents, and RAG systems.
LangWatch
Simulation-based testing, evaluation, and observability for LLM agents
TokenPath
Token-level citation and attribution API for AI-generated answers