Skip to main content
📖 The AI Tool Bible

LLM observability

Editorial picks for "best llm observability".

48 tools

All Evaluation →
BR

Braintrust

Featured
Evaluation · Platform (any LLM)
8.9

Eval, monitor, and improve AI products end-to-end.

Freemium· Starter: $0 · Pro: $249 · Enterprise: Custom pricingevalsmonitoring
LA

LangSmith

Evaluation · Platform (any LLM)
8.7

LangChain's eval + observability platform.

Freemium· Developer: $0 · Plus: $39 · Enterprise: Custom pricingLLM tracingevals
WB

Weights & Biases

Evaluation · Platform (any LLM)
8.4

The ML experiment tracker, now with LLM eval features.

Freemium· Free: $0/mo · Pro: Starts at $60/month · Enterprise: Custom plans · Personal: $0/mo · Advanced Enterprise: Custom planML experimentsLLM eval
HE

Helicone

Evaluation · Platform (any LLM)
8.3

Open-source LLM observability — one-line proxy install.

Freemium· Free 100k req/mo; Pro from $25/moobservabilitycost tracking
AG

AgentOps

Agents · Multi-model
8.2

Observability and debugging platform purpose-built for AI agents, with time-travel replay and cost tracking across 400+ LLMs.

Freemium· Free up to 5,000 events; Pro from $40/mo; Enterprise customagent-observabilityllm-tracing
AA

Arize AI

Evaluation · Multi-model
8.2

Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-observabilityagent-evaluation
GI

Giskard

Evaluation · Multi-model
8.2

Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.

Freemium· Open-source free tier; Giskard Hub enterprise pricing on requestllm-red-teamingagent-security-testing
HU

Humanloop

Evaluation · Platform (any LLM)
8.2

Prompt management + evals for collaborative AI teams.

Paid· From $200/mo teamprompt managementteam collab
AA

Athina AI

Evaluation · Multi-model
8.1

Collaborative LLM evaluation and observability platform for teams shipping AI features to production.

Freemium· Starter free (10k logs/mo); Pro & Enterprise customllm-evaluationprompt-management
HO

HoneyHive

Evaluation · Multi-model
8.1

OpenTelemetry-native observability and evaluation platform for LLM agents in production.

Freemium· Free tier available; paid/enterprise tiers via salesagent-observabilityllm-evaluation
ML

MLflow

Evaluation · Multi-model
8.1

Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.

Free· Free and open source (Apache 2.0); managed offering via Databricksllm-evaluationexperiment-tracking
OE

OpenAI Evals

Evaluation · OpenAI GPT models (extensible)
8.1

OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.

Free· Free (MIT); you pay OpenAI API costs for eval runsllm-benchmarkingregression-testing
TR

TruLens

Evaluation · Multi-model (LLM-as-judge)
8.1

Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.

Free· Free, open source (Apache-licensed Python package)llm-evaluationrag-evaluation
WB

W&B Weave

Evaluation · Multi-model
8.1

Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

Freemium· Free tier available; paid and enterprise plans via W&Bllm-tracingagent-observability
PR

PromptHub

Writing · Multi-model (OpenAI, Anthropic, Google, Meta, Mistral, Bedrock, Azure)
8.0

Git-style prompt management, testing, and deployment platform for teams running multiple LLMs in production.

Freemium· Free signup; paid team plans (contact sales / in-app)prompt-managementprompt-versioning
PR

PromptLayer

Evaluation · Platform (any LLM)
7.9

Lightweight prompt logging + management for OpenAI/Claude apps.

Freemium· Free: $0 · Pro: $49 · Team: $500 · Enterprise: Customprompt loggingversioning
PA

Patronus

Evaluation · Platform (any LLM)
7.8

Automated LLM evaluation for hallucinations, safety, and quality.

Paid· Individual: Free · Base: $25 · Enterprise: Contact us for Pricinghallucination detectionsafety
KA

Kong AI Gateway

Agents · Multi-model
7.3

Enterprise API gateway extended to route, govern, and observe LLM and agent traffic across providers.

Freemium· Free trial: $0 · Plus: Charged per Gateway per month · Enterprise: Custom pricing · Essentials: $0 · Pro: $12llm-gatewaymulti-llm-routing
LA

Langfuse

Evaluation · Model-agnostic
7.3

Open-source LLM observability, prompt management, and evaluation in one platform.

Freemium· Free self-host & Hobby tier; Core $29/mo, Pro $199/mo, Enterprise $2,499/mollm-observabilityprompt-management
OP

Opik

Evaluation · Multi-model
7.3

Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Freemium· Free open-source self-host; free Cloud tier (no card); Enterprise contact salesllm-tracingagent-evaluation
SE

Seldon

Agents · Multi-model (bring your own)
7.3

Kubernetes-native MLOps platform for deploying and orchestrating ML and generative AI models in production.

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesmodel-servinginference-pipelines
IA

Inspect AI

Evaluation · Multi-model
7.2

Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.

Free· Free and open source (MIT-style license); you pay only for underlying model API usage.llm-benchmarkingagent-evaluation
PL

Plano

Agents · Multi-model
7.2

Envoy-based data plane for AI agents that handles routing, guardrails, and observability outside your app code.

Free· Open source; commercial pricing not disclosedagent-orchestrationllm-routing
PA

Portkey AI Gateway

Agents · Multi-model (1,600+ LLMs)
7.2

Open-source AI gateway that routes a single API call across 1,600+ LLMs with caching, fallbacks, and observability.

Freemium· Developer: Free Forever · Production: $49/month · Enterprise: Custom Pricingllm-routingfallbacks-and-retries
PR

Promptfoo

Evaluation · Multi-model
7.2

Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.

Freemium· Community: Free · Enterprise: Custom · On-Premise: Customllm-evalsred-teaming
WA

Wallaroo.AI

Agents · Multi-model
7.2

Production AI inference platform for deploying and monitoring models across cloud, on-prem, and edge.

Enterprise· Starter: $500 · Wallaroo Community Edition: Free · Ampere Community Edition: Freemodel-deploymentmlops
AR

Arthur

Evaluation · Multi-model
7.1

Open-source toolkit for testing, tracing, and monitoring production AI agents.

Freemium· Free: $0/mo · Premium: $60/mo · Enterprise: Customagent-evaluationprompt-management
FA

Fiddler AI

Evaluation · Fiddler Centor (proprietary evaluators)
7.1

Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.

Enterprise· Free: Free · Developer: $0.002 per trace · Enterprise: Contact salesllm-observabilityagent-monitoring
MA

Maxim AI

Evaluation · Multi-model
7.1

End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.

Freemium· Developer: Free · Professional: $29 /seat /month · Business: $49 /seat /month · Enterprise: Customagent-evaluationllm-observability
PO

Portkey

Agents · Multi-model
7.1

Production LLM gateway with observability, guardrails, and prompt management for teams shipping AI in anger.

Freemium· Developer: Free Forever · Production: $49/month · Enterprise: Custom Pricingllm-gatewayobservability
PA

Puzzlet AI

Agents · Multi-model
7.1

Git-native prompt management and observability platform for teams shipping LLM applications.

Freemium· Basic: $20 · Pro: $50 · Enterprise: Contact salesprompt-managementllm-observability
RF

Respan (formerly Keywords AI)

Evaluation · Multi-model (500+ via gateway)
7.1

LLM engineering platform combining a multi-model gateway with tracing, evals, and prompt management.

Freemium· Free tier; paid plans (pricing not public); enterprise on requestllm-observabilityprompt-management
PH

Phoenix

Evaluation · Multi-model
7.0

Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-tracingagent-debugging
SU

Superwise

Evaluation · Multi-model
7.0

Agentic management platform for runtime guardrails, policy enforcement, and observability across LLM agents.

Freemium· Starter: Free · Pro+: ? · Enterprise: ?llm-guardrailsai-governance
SY

SystemPrompt

Agents · Multi-model
7.0

Self-hosted AI governance gateway that audits, gates, and logs every LLM call before it leaves your network.

Freemium· Free self-hosted tier; commercial licensing on requestai-governancellm-gateway
AG

Agenta

Evaluation · Multi-model
6.9

Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.

Freemium· Hobby: $0 forever · Pro: $29 /month · Business: $299 /month · Enterprise: Custom · Open source: Free foreverprompt-engineeringllm-evaluation
GA

Guild AI

Agents · Multi-model (bring your own)
6.9

Control plane for deploying, governing, and auditing AI agents in production.

Freemium· Free: $0 · Individual: $20 · Team: $200 · Enterprise: Contact usagent deploymentagent governance
TR

TrueFoundry

Agents · Multi-model
6.9

Enterprise control plane for deploying, governing, and scaling agentic AI on your own infrastructure.

Enterprise· Developer: $0 · Pro*: $499 · Pro Plus: $2999 · Enterprise: Customagent-deploymentllm-serving
CT

Cleanlab TLM

Evaluation · Multi-model (wraps any LLM)
6.8

Trustworthiness scoring layer that flags LLM hallucinations in real time.

Freemium· Free tier for evaluation; usage-based API pricing; enterprise/private deployment via saleshallucination-detectionrag-evaluation
PA

Parea AI

Evaluation · Multi-model
6.8

LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.

Freemium· Free (2 seats, 3k logs/mo); Team $150/mo; Enterprise customllm-evaluationprompt-management
AM

AI Meter

Evaluation

Local usage meter that turns AI coding-agent tokens into estimated electricity and water consumption.

Free· Free for individuals and companies; open source under a public GitHub repo.Tracking daily token usage across Claude Code and CursorEstimating electricity draw of an AI-assisted coding session
CL

ClickHouse

RAG

The open-source columnar database powering real-time analytics — and, increasingly, LLM observability and RAG backends.

Freemium· Open-source self-managed: free. ClickHouse Cloud: from $50/month (usage-based on compute + storage, AWS/GCP/Azure). Enterprise tier available with dedicated support and BYOC options.LLM trace and cost analyticsRAG retrieval with hybrid vector + metadata filters
HA

Habibi

Evaluation · GPT-4o, Perplexity Sonar, Gemini (bring-your-own API keys)

Self-hosted generative engine optimization for tracking brand mentions across ChatGPT, Perplexity, and Gemini

Freemium· Opsily Server: €40AI answer engine visibility trackingChatGPT brand mention monitoring
HY

Hydra

Agents · Multi-provider: routes across Claude, GPT, Gemini Flash, OpenRouter-hosted models, and local Qwen via Ollama / LM Studio

Local-first trust control plane that routes AI tasks to the cheapest model that clears your confidence bar.

Free· Free and open-source under the MIT license; no hosted tier or paid plan. You still pay whatever the underlying providers (Anthropic, OpenAI, OpenRouter, etc.) charge for tokens Hydra dispatches to them.multi-model CLI routingcost-optimized code generation
LM

Lagotto Meter

Evaluation

Measure the gap between what your site claims and what an AI agent actually finds

Free· Free (currently in public beta)AI-visibility audit of a landing pageFact-vs-claim gap analysis for marketing copy
LA

Lakera

Evaluation · Proprietary in-house classifiers; model-agnostic (works in front of GPT-4o, Claude, Gemini, Llama, and custom LLMs)

Runtime security and guardrails for GenAI apps, agents, and RAG systems.

Freemium· Free community/developer tier at platform.lakera.ai; paid Enterprise plans (custom pricing, contact sales). No public price list.Prompt injection defense for chatbotsRAG guardrails against indirect injection
LA

LangWatch

Evaluation · Model-agnostic; supports OpenAI, Anthropic, AWS Bedrock, Azure OpenAI, Vertex AI, and any OpenTelemetry-instrumented LLM

Simulation-based testing, evaluation, and observability for LLM agents

Freemium· Developer: €0 · Growth: €29/ core-seat / month · Enterprise: CustomLLM agent regression testing in CIRAG answer-quality evaluation
TO

TokenPath

RAG · model-agnostic (works with any LLM output; uses in-house attribution model)

Token-level citation and attribution API for AI-generated answers

Freemium· 10M tokens free to start (no card), then $1 per 1M tokens pay-as-you-goRAG chatbot citationcontract and policy Q&A