Skip to main content
📖 The AI Tool Bible

LLM monitoring

Editorial picks for "best llm monitoring".

38 tools

All Evaluation →
BR

Braintrust

Featured
Evaluation · Platform (any LLM)
8.9

Eval, monitor, and improve AI products end-to-end.

Freemium· Starter: $0 · Pro: $249 · Enterprise: Custom pricingevalsmonitoring
LA

LangSmith

Evaluation · Platform (any LLM)
8.7

LangChain's eval + observability platform.

Freemium· Developer: $0 · Plus: $39 · Enterprise: Custom pricingLLM tracingevals
WB

Weights & Biases

Evaluation · Platform (any LLM)
8.4

The ML experiment tracker, now with LLM eval features.

Freemium· Free: $0/mo · Pro: Starts at $60/month · Enterprise: Custom plans · Personal: $0/mo · Advanced Enterprise: Custom planML experimentsLLM eval
HE

Helicone

Evaluation · Platform (any LLM)
8.3

Open-source LLM observability — one-line proxy install.

Freemium· Free 100k req/mo; Pro from $25/moobservabilitycost tracking
AG

AgentOps

Agents · Multi-model
8.2

Observability and debugging platform purpose-built for AI agents, with time-travel replay and cost tracking across 400+ LLMs.

Freemium· Free up to 5,000 events; Pro from $40/mo; Enterprise customagent-observabilityllm-tracing
AA

Arize AI

Evaluation · Multi-model
8.2

Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-observabilityagent-evaluation
AA

Athina AI

Evaluation · Multi-model
8.1

Collaborative LLM evaluation and observability platform for teams shipping AI features to production.

Freemium· Starter free (10k logs/mo); Pro & Enterprise customllm-evaluationprompt-management
HO

HoneyHive

Evaluation · Multi-model
8.1

OpenTelemetry-native observability and evaluation platform for LLM agents in production.

Freemium· Free tier available; paid/enterprise tiers via salesagent-observabilityllm-evaluation
ML

MLflow

Evaluation · Multi-model
8.1

Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.

Free· Free and open source (Apache 2.0); managed offering via Databricksllm-evaluationexperiment-tracking
TR

TruLens

Evaluation · Multi-model (LLM-as-judge)
8.1

Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.

Free· Free, open source (Apache-licensed Python package)llm-evaluationrag-evaluation
WB

W&B Weave

Evaluation · Multi-model
8.1

Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

Freemium· Free tier available; paid and enterprise plans via W&Bllm-tracingagent-observability
PR

PromptLayer

Evaluation · Platform (any LLM)
7.9

Lightweight prompt logging + management for OpenAI/Claude apps.

Freemium· Free: $0 · Pro: $49 · Team: $500 · Enterprise: Customprompt loggingversioning
KA

Kong AI Gateway

Agents · Multi-model
7.3

Enterprise API gateway extended to route, govern, and observe LLM and agent traffic across providers.

Freemium· Free trial: $0 · Plus: Charged per Gateway per month · Enterprise: Custom pricing · Essentials: $0 · Pro: $12llm-gatewaymulti-llm-routing
LA

Langfuse

Evaluation · Model-agnostic
7.3

Open-source LLM observability, prompt management, and evaluation in one platform.

Freemium· Free self-host & Hobby tier; Core $29/mo, Pro $199/mo, Enterprise $2,499/mollm-observabilityprompt-management
OP

Opik

Evaluation · Multi-model
7.3

Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Freemium· Free open-source self-host; free Cloud tier (no card); Enterprise contact salesllm-tracingagent-evaluation
PL

Plano

Agents · Multi-model
7.2

Envoy-based data plane for AI agents that handles routing, guardrails, and observability outside your app code.

Free· Open source; commercial pricing not disclosedagent-orchestrationllm-routing
PA

Portkey AI Gateway

Agents · Multi-model (1,600+ LLMs)
7.2

Open-source AI gateway that routes a single API call across 1,600+ LLMs with caching, fallbacks, and observability.

Freemium· Developer: Free Forever · Production: $49/month · Enterprise: Custom Pricingllm-routingfallbacks-and-retries
WA

Wallaroo.AI

Agents · Multi-model
7.2

Production AI inference platform for deploying and monitoring models across cloud, on-prem, and edge.

Enterprise· Starter: $500 · Wallaroo Community Edition: Free · Ampere Community Edition: Freemodel-deploymentmlops
AR

Arthur

Evaluation · Multi-model
7.1

Open-source toolkit for testing, tracing, and monitoring production AI agents.

Freemium· Free: $0/mo · Premium: $60/mo · Enterprise: Customagent-evaluationprompt-management
FA

Fiddler AI

Evaluation · Fiddler Centor (proprietary evaluators)
7.1

Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.

Enterprise· Free: Free · Developer: $0.002 per trace · Enterprise: Contact salesllm-observabilityagent-monitoring
MA

Maxim AI

Evaluation · Multi-model
7.1

End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.

Freemium· Developer: Free · Professional: $29 /seat /month · Business: $49 /seat /month · Enterprise: Customagent-evaluationllm-observability
PO

Portkey

Agents · Multi-model
7.1

Production LLM gateway with observability, guardrails, and prompt management for teams shipping AI in anger.

Freemium· Developer: Free Forever · Production: $49/month · Enterprise: Custom Pricingllm-gatewayobservability
PA

Puzzlet AI

Agents · Multi-model
7.1

Git-native prompt management and observability platform for teams shipping LLM applications.

Freemium· Basic: $20 · Pro: $50 · Enterprise: Contact salesprompt-managementllm-observability
RF

Respan (formerly Keywords AI)

Evaluation · Multi-model (500+ via gateway)
7.1

LLM engineering platform combining a multi-model gateway with tracing, evals, and prompt management.

Freemium· Free tier; paid plans (pricing not public); enterprise on requestllm-observabilityprompt-management
PH

Phoenix

Evaluation · Multi-model
7.0

Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-tracingagent-debugging
SU

Superwise

Evaluation · Multi-model
7.0

Agentic management platform for runtime guardrails, policy enforcement, and observability across LLM agents.

Freemium· Starter: Free · Pro+: ? · Enterprise: ?llm-guardrailsai-governance
SY

SystemPrompt

Agents · Multi-model
7.0

Self-hosted AI governance gateway that audits, gates, and logs every LLM call before it leaves your network.

Freemium· Free self-hosted tier; commercial licensing on requestai-governancellm-gateway
AG

Agenta

Evaluation · Multi-model
6.9

Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.

Freemium· Hobby: $0 forever · Pro: $29 /month · Business: $299 /month · Enterprise: Custom · Open source: Free foreverprompt-engineeringllm-evaluation
GA

Guild AI

Agents · Multi-model (bring your own)
6.9

Control plane for deploying, governing, and auditing AI agents in production.

Freemium· Free: $0 · Individual: $20 · Team: $200 · Enterprise: Contact usagent deploymentagent governance
CT

Cleanlab TLM

Evaluation · Multi-model (wraps any LLM)
6.8

Trustworthiness scoring layer that flags LLM hallucinations in real time.

Freemium· Free tier for evaluation; usage-based API pricing; enterprise/private deployment via saleshallucination-detectionrag-evaluation
PA

Parea AI

Evaluation · Multi-model
6.8

LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.

Freemium· Free (2 seats, 3k logs/mo); Team $150/mo; Enterprise customllm-evaluationprompt-management
AM

AI Meter

Evaluation

Local usage meter that turns AI coding-agent tokens into estimated electricity and water consumption.

Free· Free for individuals and companies; open source under a public GitHub repo.Tracking daily token usage across Claude Code and CursorEstimating electricity draw of an AI-assisted coding session
CL

ClickHouse

RAG

The open-source columnar database powering real-time analytics — and, increasingly, LLM observability and RAG backends.

Freemium· Open-source self-managed: free. ClickHouse Cloud: from $50/month (usage-based on compute + storage, AWS/GCP/Azure). Enterprise tier available with dedicated support and BYOC options.LLM trace and cost analyticsRAG retrieval with hybrid vector + metadata filters
GM

Grafana MCP

MCP Servers

Official Grafana Labs MCP server — dashboards, Prometheus, Loki, alerts and incidents in your LLM client

Free· Free and open source (Apache 2.0). You still need a Grafana instance — OSS Grafana is free; Grafana Cloud has a free tier plus paid Pro/Advanced/Enterprise plans.Incident investigation from Claude DesktopConversational PromQL and LogQL querying
HA

Habibi

Evaluation · GPT-4o, Perplexity Sonar, Gemini (bring-your-own API keys)

Self-hosted generative engine optimization for tracking brand mentions across ChatGPT, Perplexity, and Gemini

Freemium· Opsily Server: €40AI answer engine visibility trackingChatGPT brand mention monitoring
LM

Lagotto Meter

Evaluation

Measure the gap between what your site claims and what an AI agent actually finds

Free· Free (currently in public beta)AI-visibility audit of a landing pageFact-vs-claim gap analysis for marketing copy
LA

Lakera

Evaluation · Proprietary in-house classifiers; model-agnostic (works in front of GPT-4o, Claude, Gemini, Llama, and custom LLMs)

Runtime security and guardrails for GenAI apps, agents, and RAG systems.

Freemium· Free community/developer tier at platform.lakera.ai; paid Enterprise plans (custom pricing, contact sales). No public price list.Prompt injection defense for chatbotsRAG guardrails against indirect injection
LA

LangWatch

Evaluation · Model-agnostic; supports OpenAI, Anthropic, AWS Bedrock, Azure OpenAI, Vertex AI, and any OpenTelemetry-instrumented LLM

Simulation-based testing, evaluation, and observability for LLM agents

Freemium· Developer: €0 · Growth: €29/ core-seat / month · Enterprise: CustomLLM agent regression testing in CIRAG answer-quality evaluation