Skip to main content
📖 The AI Tool Bible

AI cost tracking

Editorial picks for "track llm costs".

48 tools

All Evaluation
Helicone preview image
Helicone logo

Helicone

Evaluation · Platform (any LLM)
8.3

Open-source LLM observability — one-line proxy install.

Freemium· Free 100k req/mo; Pro from $25/moobservabilitycost tracking
OpenAI Evals preview image
OpenAI Evals logo

OpenAI Evals

Evaluation · OpenAI GPT models (extensible)
8.1

OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.

Free· Free (MIT); you pay OpenAI API costs for eval runsllm-benchmarkingregression-testing
Domino Data Lab preview image
Domino Data Lab logo

Domino Data Lab

Agents · Multi-model
7.2

Enterprise AI platform for building, deploying, and governing models and agents at scale.

Enterprise· Domino Cloud: Contact sales · Premium: Contact sales · Enterprise: Contact salesenterprise mlopsagentic ai
AgentOps preview image
AgentOps logo

AgentOps

Agents · Multi-model
8.2

Observability and debugging platform purpose-built for AI agents, with time-travel replay and cost tracking across 400+ LLMs.

Freemium· Free up to 5,000 events; Pro from $40/mo; Enterprise customagent-observabilityllm-tracing
Sim preview image
Sim logo

Sim

Agents · Multi-model (Claude, OpenAI, Google)
8.2

Open-source visual workspace for building, deploying, and monitoring AI agents.

Freemium· Free: $0 · Pro: $25 · Max: $100 · Enterprise: Customagent-workflowsinternal-automation
LLM Stats preview image
LLM Stats logo

LLM Stats

Evaluation · Multi-model
7.9

Live leaderboard and side-by-side comparison hub for 300+ frontier LLMs across reasoning, coding, and multimodal benchmarks.

Free· Free to browse; underlying model usage billed by each providermodel-comparisonbenchmark-tracking
Langfuse preview image
Langfuse logo

Langfuse

Evaluation · Model-agnostic
7.3

Open-source LLM observability, prompt management, and evaluation in one platform.

Freemium· Free self-host & Hobby tier; Core $29/mo, Pro $199/mo, Enterprise $2,499/mollm-observabilityprompt-management
Opik preview image
Opik logo

Opik

Evaluation · Multi-model
7.3

Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Freemium· Free open-source self-host; free Cloud tier (no card); Enterprise contact salesllm-tracingagent-evaluation
Valohai preview image
Valohai logo

Valohai

Agents · Multi-model
7.3

MLOps platform for versioned pipelines, distributed training, and LLM evaluation across any cloud.

Enterprise· Free trial; contact sales for pricingmlopsllm-evaluation
Portkey AI Gateway preview image
Portkey AI Gateway logo

Portkey AI Gateway

Agents · Multi-model (1,600+ LLMs)
7.2

Open-source AI gateway that routes a single API call across 1,600+ LLMs with caching, fallbacks, and observability.

Freemium· Developer: Free Forever · Production: $49/month · Enterprise: Custom Pricingllm-routingfallbacks-and-retries
Portkey preview image
Portkey logo

Portkey

Agents · Multi-model
7.1

Production LLM gateway with observability, guardrails, and prompt management for teams shipping AI in anger.

Freemium· Developer: Free Forever · Production: $49/month · Enterprise: Custom Pricingllm-gatewayobservability
Fynk preview image
Fynk logo

Fynk

Writing
6.9

AI-assisted contract lifecycle management with eIDAS-grade e-signatures and risk flagging built in.

Paid· Free trial; pricing on requestcontract-managemente-signature
LangWatch preview image
LangWatch logo

LangWatch

Evaluation · Model-agnostic; supports OpenAI, Anthropic, AWS Bedrock, Azure OpenAI, Vertex AI, and any OpenTelemetry-instrumented LLM

Simulation-based testing, evaluation, and observability for LLM agents

Freemium· Developer: €0 · Growth: €29/ core-seat / month · Enterprise: CustomLLM agent regression testing in CIRAG answer-quality evaluation
LLM Gateway preview image
LLM Gateway logo

LLM Gateway

Agents · Routes to GPT-4o, Claude 3.5 Sonnet, Gemini 1.5, Llama 3.1, Mistral, and 200+ others

One API, 200+ models, transparent pricing, and no vendor lock-in.

Freemium· Free: $0 · Enterprise: CustomMulti-provider LLM routingAutomatic failover between model vendors
Aider preview image
Aider logo

Aider

Coding · BYO (Claude / GPT-4 / Gemini / DeepSeek)
8.4

Terminal-based AI pair programmer that writes commits.

Free· Free / open-source; you pay the underlying LLM API costsCLIgit workflow
Weights & Biases preview image
Weights & Biases logo

Weights & Biases

Evaluation · Platform (any LLM)
8.4

The ML experiment tracker, now with LLM eval features.

Freemium· Free: $0/mo · Pro: Starts at $60/month · Enterprise: Custom plans · Personal: $0/mo · Advanced Enterprise: Custom planML experimentsLLM eval
Yi (01.AI) preview image
Yi (01.AI) logo

Yi (01.AI)

Agents · Yi-Lightning (MoE), Yi-Large, Yi-1.5 (6B/9B/34B), Yi-VL, Yi-Coder — in-house 01.AI foundation models
8.4

Foundation models from 01.AI — open-weight Yi family plus frontier Yi-Lightning and Yi-Large

Freemium· Open-source Yi models free under permissive license; hosted API via platform.lingyiwanwu.com with pay-per-token pricing (Yi-Lightning positioned as a low-cost frontier tier; Yi-Large priced higher; exact per-token rates on the platform dashboard). Enterprise custom-training and consulting on quote.Self-hosted coding assistantBilingual English-Chinese chatbot
ClearML preview image
ClearML logo

ClearML

Agents · Model-agnostic
8.3

End-to-end MLOps and GenAI platform with open-source experiment tracking and enterprise GPU orchestration.

Freemium· Community: $0 · Pro: $15 Per User/Month + Usage · Scale: Custom Quote · Enterprise: Request a Quoteexperiment-trackinggpu-orchestration
LiveBench preview image
LiveBench logo

LiveBench

Evaluation · Multi-model
8.2

Contamination-free LLM benchmark that refreshes its questions monthly to keep frontier models honest.

Free· Free and open source; self-hosted evaluation runnerllm-benchmarkingmodel-selection
Berkeley Function-Calling Leaderboard preview image
Berkeley Function-Calling Leaderboard logo

Berkeley Function-Calling Leaderboard

Evaluation · Multi-model
8.1

Open benchmark from UC Berkeley that ranks LLMs on real-world tool-use and function-calling accuracy.

Free· Free and open source; you pay only for inference when reproducing runs.function-calling evaltool-use benchmarking
Hugging Face AutoTrain preview image
Hugging Face AutoTrain logo

Hugging Face AutoTrain

Fine-tuning · Multi-model (Hugging Face Hub)
8.1

No-code fine-tuning and training pipeline that spins up state-of-the-art models on the Hugging Face Hub.

Paid· Per-minute billing based on hardware tier; self-hosted OSS version is freellm-fine-tuningtext-classification
MLflow preview image
MLflow logo

MLflow

Evaluation · Multi-model
8.1

Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.

Free· Free and open source (Apache 2.0); managed offering via Databricksllm-evaluationexperiment-tracking
TruLens preview image
TruLens logo

TruLens

Evaluation · Multi-model (LLM-as-judge)
8.1

Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.

Free· Free, open source (Apache-licensed Python package)llm-evaluationrag-evaluation
W&B Weave preview image
W&B Weave logo

W&B Weave

Evaluation · Multi-model
8.1

Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

Freemium· Free tier available; paid and enterprise plans via W&Bllm-tracingagent-observability
Firecrawl preview image
Firecrawl logo

Firecrawl

RAG · Claude, Cursor, Windsurf, OpenAI, Gemini
8.0

Web scraping and crawling API that returns LLM-ready markdown, JSON, or structured data from any URL.

Freemium· Free Plan: $0 · Hobby: $16 · Standard: $83 · Growth: $333 · Scale: $599/monthlyweb-scrapingrag-ingestion
OpenHands preview image
OpenHands logo

OpenHands

Agents · Multi-model (Claude, Gemini, GPT/Codex)
8.0

Open-source autonomous coding agent that plans, edits, and ships changes across real codebases.

Freemium· Open Source: Free · Individual: Free · Enterprise: Custom pricingautonomous-codingpr-review
Sematic preview image
Sematic logo

Sematic

Agents
7.3

Open-source Python-first orchestrator for ML training pipelines from laptop to cloud.

Freemium· Open-source free; managed/enterprise tier on requestml-pipelinestraining-orchestration
Cherry Studio preview image
Cherry Studio logo

Cherry Studio

Agents · Multi-model
7.2

Open-source desktop AI client that wires 300+ LLMs into one chat, knowledge-base, and agent workspace.

Free· Free and open source; bring your own API keysmulti-model chatlocal knowledge base
FinChat (Fiscal.ai) preview image
FinChat (Fiscal.ai) logo

FinChat (Fiscal.ai)

RAG · Multi-model (proprietary finance-tuned copilot)
7.2

AI copilot for equity research that reads filings, transcripts, and KPI tables across 100,000+ public companies.

Freemium· Free; Pro $39/mo (annual) or $49/mo; Max and Enterprise API tiers aboveequity-researchearnings-call-analysis
Graphiti preview image
Graphiti logo

Graphiti

RAG · Multi-model
7.2

Open-source temporal knowledge graph framework for building agent memory that updates in real time.

Freemium· Basic: $10 · Pro: $20 · Enterprise: $50agent-memorytemporal-knowledge-graphs
Inspect AI preview image
Inspect AI logo

Inspect AI

Evaluation · Multi-model
7.2

Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.

Free· Free and open source (MIT-style license); you pay only for underlying model API usage.llm-benchmarkingagent-evaluation
LLaMA Factory preview image
LLaMA Factory logo

LLaMA Factory

Fine-tuning · Multi-model (LLaMA, Mistral, Qwen, Gemma, Phi, LLaVA, ChatGLM, Yi)
7.2

Open-source, no-code WebUI for fine-tuning 100+ open LLMs with LoRA, QLoRA, DPO, and PPO.

Free· Free, open-source (Apache-2.0); self-hostedlora-fine-tuningqlora
Perplexity AI preview image
Perplexity AI logo

Perplexity AI

RAG · Multi-model (Sonar, GPT-4 class, Claude, Gemini)
7.2

Conversational answer engine that cites its sources by default.

Freemium· Basic: $10 · Pro: $30 · Enterprise: Contact salesai-searchresearch
GPT for Sheets preview image
GPT for Sheets logo

GPT for Sheets

Writing · Multi-model (GPT, Claude, Gemini, Grok, DeepSeek, Mistral, OpenRouter)
7.1

Bulk AI prompts inside Google Sheets, Docs, Slides, Forms and Gmail with 50+ frontier models.

Freemium· Starter: €3.40 · Standard: €6.80 · Plus: €10.55 · Enterprise: Contact salesbulk-promptingspreadsheet-automation
SEAL Leaderboard preview image
SEAL Leaderboard logo

SEAL Leaderboard

Evaluation · Multi-model (GPT, Claude, Gemini, Llama, etc.)
7.1

Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.

Free· Free to view; paid custom evals via Scale enterprise salesmodel-selectionbenchmark-tracking
VisualWebArena preview image
VisualWebArena logo

VisualWebArena

Evaluation · Model-agnostic (GPT-4V, Gemini, Claude, open VLMs)
7.1

Open benchmark for evaluating multimodal web agents on realistic visual browsing tasks.

Free· Free and open source (MIT-style research release)multimodal-agent-evalweb-browsing-benchmark
W&B Sweeps preview image
W&B Sweeps logo

W&B Sweeps

Fine-tuning · Multi-model (Llama, DeepSeek, Qwen, Kimi)
7.1

Hyperparameter optimization from Weights & Biases with Bayesian search and Hyperband early stopping.

Freemium· Free: $0/mo · Pro: $60/month, billed monthly · Enterprise: Custom plans · Personal: $0/mo · Advanced Enterprise: Custom planhyperparameter-tuningbayesian-optimization
Kotaemon preview image
Kotaemon logo

Kotaemon

RAG · Multi-model (OpenAI, LlamaCPP, any OpenAI-compatible endpoint)
7.0

Open-source RAG UI for chatting with your own documents, locally or self-hosted.

Free· Free, open-source (MIT-style); self-hosted infrastructure costs onlydocument-qaprivate-rag
RAGs by LlamaIndex preview image
RAGs by LlamaIndex logo

RAGs by LlamaIndex

RAG · Multi-model (OpenAI, Anthropic, Replicate, HuggingFace)
7.0

Open-source Streamlit app that builds a custom RAG pipeline from a natural-language brief.

Free· Free, MIT-licensed; bring your own model/API keysnatural-language-rag-builderdocument-qa
SWE-agent preview image
SWE-agent logo

SWE-agent

Agents · Multi-model (GPT-4o, Claude Sonnet, DeepSeek, local via LiteLLM)
7.0

Open-source autonomous agent framework that lets LLMs fix GitHub issues and find security vulnerabilities by using a custom agent-computer interface.

Free· Free and open-source; you pay your own LLM API costsgithub-issue-fixingautonomous-coding
Copysmith preview image
Copysmith logo

Copysmith

Writing · Multi-model
6.9

Umbrella content platform bundling Frase, Describely, and Rytr for SEO-aware AI writing at enterprise scale.

Freemium· Per-product pricing; Rytr has free tier, Frase/Describely paidseo-contentproduct-descriptions
Huntr AI Resume Builder preview image
Huntr AI Resume Builder logo

Huntr AI Resume Builder

Writing
6.9

AI resume builder bolted onto a full job-search CRM, tuned to tailor bullets to specific job descriptions.

Freemium· Free: Free · Pro: $40/monthresume-tailoringcover-letter-generation
Katonic AI preview image
Katonic AI logo

Katonic AI

Agents · Multi-model (2,600+ via AI Gateway)
6.9

Sovereign enterprise platform for building, deploying, and governing AI agents on your own infrastructure.

Enterprise· Contact sales for quote; no public pricingenterprise-agentson-prem-llm
MindPal preview image
MindPal logo

MindPal

Agents · Multi-model
6.9

No-code platform for building and orchestrating teams of AI agents that automate business workflows.

Freemium· Free tier; paid plans for higher usage and team seatsmulti-agent-workflowsbusiness-automation
Price Per Token preview image
Price Per Token logo

Price Per Token

Coding · Multi-model
6.9

Daily-updated LLM API pricing comparison across 300+ models, with calculators, leaderboards, and a free MCP server.

Free· Free; ad-supported, no signupllm-pricing-comparisontoken-cost-calculator
ReBillion.ai preview image
ReBillion.ai logo

ReBillion.ai

Agents
6.9

AI transaction coordinator that extracts dates, deadlines, and compliance rules from real estate contracts.

Paid· AI Toolkit + Customization: $199 · Human Assistant + AI + Customization: $499transaction-coordinationcontract-analysis
ShyEditor preview image
ShyEditor logo

ShyEditor

Writing · Multi-model (GPT-4, Grok)
6.9

AI-native writing environment that bolts an assistant, citation manager, and knowledge base onto a distraction-free markdown editor.

Freemium· Free Basic (20 AI credits/mo, 100MB); Pro $10/mo (500 credits, 5GB)long-form-writingacademic-writing
Ask AI preview image
Ask AI logo

Ask AI

Writing · GPT-4o mini
6.8

ChatGPT on your wrist: a watchOS-only AI assistant from Sindre Sorhus with no subscription.

Paid· One-time App Store purchase; no subscriptionquick-qatranslation