Skip to main content
📖 The AI Tool Bible

AI cost tracking

Editorial picks for "track llm costs".

31 tools

All Evaluation →
LA

LangSmith

Evaluation · Platform (any LLM)
8.7

LangChain's eval + observability platform.

Freemium· Developer: $0 · Plus: $39 · Enterprise: Custom pricingLLM tracingevals
HE

Helicone

Evaluation · Platform (any LLM)
8.3

Open-source LLM observability — one-line proxy install.

Freemium· Free 100k req/mo; Pro from $25/moobservabilitycost tracking
AG

AgentOps

Agents · Multi-model
8.2

Observability and debugging platform purpose-built for AI agents, with time-travel replay and cost tracking across 400+ LLMs.

Freemium· Free up to 5,000 events; Pro from $40/mo; Enterprise customagent-observabilityllm-tracing
AA

Arize AI

Evaluation · Multi-model
8.2

Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-observabilityagent-evaluation
AA

Athina AI

Evaluation · Multi-model
8.1

Collaborative LLM evaluation and observability platform for teams shipping AI features to production.

Freemium· Starter free (10k logs/mo); Pro & Enterprise customllm-evaluationprompt-management
HO

HoneyHive

Evaluation · Multi-model
8.1

OpenTelemetry-native observability and evaluation platform for LLM agents in production.

Freemium· Free tier available; paid/enterprise tiers via salesagent-observabilityllm-evaluation
WB

W&B Weave

Evaluation · Multi-model
8.1

Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

Freemium· Free tier available; paid and enterprise plans via W&Bllm-tracingagent-observability
HE

Headroom

Agents · Model-agnostic (Anthropic, OpenAI, Vertex, Bedrock, Azure, 100+ via LiteLLM)
7.4

Open-source context compression layer that strips 70-95% of boilerplate before it hits your LLM.

Free· Apache 2.0 open source; free for commercial usetoken-compressionagent-context
BI

Bifrost

Agents · Multi-model (OpenAI, Anthropic, Bedrock, Vertex, 1000+ via providers)
7.3

Open-source AI gateway that unifies 1000+ models behind one OpenAI-compatible endpoint with failover, budgets, and MCP routing.

Freemium· Free Forever: Free · Enterprise: Contact salesllm-gatewaymulti-provider-routing
KA

Kong AI Gateway

Agents · Multi-model
7.3

Enterprise API gateway extended to route, govern, and observe LLM and agent traffic across providers.

Freemium· Free trial: $0 · Plus: Charged per Gateway per month · Enterprise: Custom pricing · Essentials: $0 · Pro: $12llm-gatewaymulti-llm-routing
LA

Langfuse

Evaluation · Model-agnostic
7.3

Open-source LLM observability, prompt management, and evaluation in one platform.

Freemium· Free self-host & Hobby tier; Core $29/mo, Pro $199/mo, Enterprise $2,499/mollm-observabilityprompt-management
OP

Opik

Evaluation · Multi-model
7.3

Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Freemium· Free open-source self-host; free Cloud tier (no card); Enterprise contact salesllm-tracingagent-evaluation
PL

Plano

Agents · Multi-model
7.2

Envoy-based data plane for AI agents that handles routing, guardrails, and observability outside your app code.

Free· Open source; commercial pricing not disclosedagent-orchestrationllm-routing
PA

Portkey AI Gateway

Agents · Multi-model (1,600+ LLMs)
7.2

Open-source AI gateway that routes a single API call across 1,600+ LLMs with caching, fallbacks, and observability.

Freemium· Developer: Free Forever · Production: $49/month · Enterprise: Custom Pricingllm-routingfallbacks-and-retries
AR

Arthur

Evaluation · Multi-model
7.1

Open-source toolkit for testing, tracing, and monitoring production AI agents.

Freemium· Free: $0/mo · Premium: $60/mo · Enterprise: Customagent-evaluationprompt-management
MA

Manifest

Agents · Multi-model
7.1

Open-source LLM router that fans your agent traffic across providers and your existing AI subscriptions.

Freemium· Free: $0 /month · Pro: $19 /month · Enterprise: Let's Talkllm-routingcost-control
PO

Portkey

Agents · Multi-model
7.1

Production LLM gateway with observability, guardrails, and prompt management for teams shipping AI in anger.

Freemium· Developer: Free Forever · Production: $49/month · Enterprise: Custom Pricingllm-gatewayobservability
PA

Puzzlet AI

Agents · Multi-model
7.1

Git-native prompt management and observability platform for teams shipping LLM applications.

Freemium· Basic: $20 · Pro: $50 · Enterprise: Contact salesprompt-managementllm-observability
RF

Respan (formerly Keywords AI)

Evaluation · Multi-model (500+ via gateway)
7.1

LLM engineering platform combining a multi-model gateway with tracing, evals, and prompt management.

Freemium· Free tier; paid plans (pricing not public); enterprise on requestllm-observabilityprompt-management
AI

AICamp

Writing · Multi-model (GPT-5.2, Claude 4, Gemini)
7.0

Team workspace that puts GPT, Claude, and Gemini behind one admin console with shared prompts, agents, and usage controls.

Freemium· Starter free (3 users, $5 credits); Business $10/user/mo; Enterprise customteam-ai-workspacemulti-model-chat
SY

SystemPrompt

Agents · Multi-model
7.0

Self-hosted AI governance gateway that audits, gates, and logs every LLM call before it leaves your network.

Freemium· Free self-hosted tier; commercial licensing on requestai-governancellm-gateway
TE

TeamoRouter

Agents · Multi-model
7.0

Unified LLM gateway that brokers Claude, GPT, and Gemini through one API key with usage-based discounts.

Paid· Pay-as-you-go; rates from ~10% of model list price, avg ~19% hourly discountllm-routingmulti-model-access
GA

Guild AI

Agents · Multi-model (bring your own)
6.9

Control plane for deploying, governing, and auditing AI agents in production.

Freemium· Free: $0 · Individual: $20 · Team: $200 · Enterprise: Contact usagent deploymentagent governance
PP

Price Per Token

Coding · Multi-model
6.9

Daily-updated LLM API pricing comparison across 300+ models, with calculators, leaderboards, and a free MCP server.

Free· Free; ad-supported, no signupllm-pricing-comparisontoken-cost-calculator
AA

Artificial Analysis

Evaluation · Multi-model
6.8

Independent benchmarking platform comparing AI models and inference providers across intelligence, speed, and cost.

Freemium· Pro: $417/month per seat · Enterprise: Custom pricingmodel-benchmarkingprovider-comparison
PA

Parea AI

Evaluation · Multi-model
6.8

LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.

Freemium· Free (2 seats, 3k logs/mo); Team $150/mo; Enterprise customllm-evaluationprompt-management
CL

ClevAgent

Agents
6.6

Middleware that supervises AI coding agents in real time, catching wasteful actions and enforcing safety guardrails.

Freemium· Free tier available; paid tiers not disclosedagent-supervisioncost-control
AM

AI Meter

Evaluation

Local usage meter that turns AI coding-agent tokens into estimated electricity and water consumption.

Free· Free for individuals and companies; open source under a public GitHub repo.Tracking daily token usage across Claude Code and CursorEstimating electricity draw of an AI-assisted coding session
CL

ClickHouse

RAG

The open-source columnar database powering real-time analytics — and, increasingly, LLM observability and RAG backends.

Freemium· Open-source self-managed: free. ClickHouse Cloud: from $50/month (usage-based on compute + storage, AWS/GCP/Azure). Enterprise tier available with dedicated support and BYOC options.LLM trace and cost analyticsRAG retrieval with hybrid vector + metadata filters
HY

Hydra

Agents · Multi-provider: routes across Claude, GPT, Gemini Flash, OpenRouter-hosted models, and local Qwen via Ollama / LM Studio

Local-first trust control plane that routes AI tasks to the cheapest model that clears your confidence bar.

Free· Free and open-source under the MIT license; no hosted tier or paid plan. You still pay whatever the underlying providers (Anthropic, OpenAI, OpenRouter, etc.) charge for tokens Hydra dispatches to them.multi-model CLI routingcost-optimized code generation
LA

LangWatch

Evaluation · Model-agnostic; supports OpenAI, Anthropic, AWS Bedrock, Azure OpenAI, Vertex AI, and any OpenTelemetry-instrumented LLM

Simulation-based testing, evaluation, and observability for LLM agents

Freemium· Developer: €0 · Growth: €29/ core-seat / month · Enterprise: CustomLLM agent regression testing in CIRAG answer-quality evaluation