AI cost tracking
Editorial picks for "track llm costs".
48 tools
Helicone
Open-source LLM observability — one-line proxy install.

OpenAI Evals
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.

Domino Data Lab
Enterprise AI platform for building, deploying, and governing models and agents at scale.

AgentOps
Observability and debugging platform purpose-built for AI agents, with time-travel replay and cost tracking across 400+ LLMs.

Sim
Open-source visual workspace for building, deploying, and monitoring AI agents.

LLM Stats
Live leaderboard and side-by-side comparison hub for 300+ frontier LLMs across reasoning, coding, and multimodal benchmarks.

Langfuse
Open-source LLM observability, prompt management, and evaluation in one platform.

Opik
Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Valohai
MLOps platform for versioned pipelines, distributed training, and LLM evaluation across any cloud.

Portkey AI Gateway
Open-source AI gateway that routes a single API call across 1,600+ LLMs with caching, fallbacks, and observability.

Portkey
Production LLM gateway with observability, guardrails, and prompt management for teams shipping AI in anger.

Fynk
AI-assisted contract lifecycle management with eIDAS-grade e-signatures and risk flagging built in.

LangWatch
Simulation-based testing, evaluation, and observability for LLM agents

LLM Gateway
One API, 200+ models, transparent pricing, and no vendor lock-in.

Aider
Terminal-based AI pair programmer that writes commits.

Weights & Biases
The ML experiment tracker, now with LLM eval features.

Yi (01.AI)
Foundation models from 01.AI — open-weight Yi family plus frontier Yi-Lightning and Yi-Large

ClearML
End-to-end MLOps and GenAI platform with open-source experiment tracking and enterprise GPU orchestration.

LiveBench
Contamination-free LLM benchmark that refreshes its questions monthly to keep frontier models honest.

Berkeley Function-Calling Leaderboard
Open benchmark from UC Berkeley that ranks LLMs on real-world tool-use and function-calling accuracy.

Hugging Face AutoTrain
No-code fine-tuning and training pipeline that spins up state-of-the-art models on the Hugging Face Hub.

MLflow
Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.

TruLens
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.

W&B Weave
Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

Firecrawl
Web scraping and crawling API that returns LLM-ready markdown, JSON, or structured data from any URL.

OpenHands
Open-source autonomous coding agent that plans, edits, and ships changes across real codebases.

Sematic
Open-source Python-first orchestrator for ML training pipelines from laptop to cloud.

Cherry Studio
Open-source desktop AI client that wires 300+ LLMs into one chat, knowledge-base, and agent workspace.

FinChat (Fiscal.ai)
AI copilot for equity research that reads filings, transcripts, and KPI tables across 100,000+ public companies.

Graphiti
Open-source temporal knowledge graph framework for building agent memory that updates in real time.

Inspect AI
Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.

LLaMA Factory
Open-source, no-code WebUI for fine-tuning 100+ open LLMs with LoRA, QLoRA, DPO, and PPO.

Perplexity AI
Conversational answer engine that cites its sources by default.

GPT for Sheets
Bulk AI prompts inside Google Sheets, Docs, Slides, Forms and Gmail with 50+ frontier models.

SEAL Leaderboard
Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.

VisualWebArena
Open benchmark for evaluating multimodal web agents on realistic visual browsing tasks.

W&B Sweeps
Hyperparameter optimization from Weights & Biases with Bayesian search and Hyperband early stopping.

Kotaemon
Open-source RAG UI for chatting with your own documents, locally or self-hosted.

RAGs by LlamaIndex
Open-source Streamlit app that builds a custom RAG pipeline from a natural-language brief.

SWE-agent
Open-source autonomous agent framework that lets LLMs fix GitHub issues and find security vulnerabilities by using a custom agent-computer interface.

Copysmith
Umbrella content platform bundling Frase, Describely, and Rytr for SEO-aware AI writing at enterprise scale.

Huntr AI Resume Builder
AI resume builder bolted onto a full job-search CRM, tuned to tailor bullets to specific job descriptions.

Katonic AI
Sovereign enterprise platform for building, deploying, and governing AI agents on your own infrastructure.

MindPal
No-code platform for building and orchestrating teams of AI agents that automate business workflows.

Price Per Token
Daily-updated LLM API pricing comparison across 300+ models, with calculators, leaderboards, and a free MCP server.

ReBillion.ai
AI transaction coordinator that extracts dates, deadlines, and compliance rules from real estate contracts.

ShyEditor
AI-native writing environment that bolts an assistant, citation manager, and knowledge base onto a distraction-free markdown editor.

Ask AI
ChatGPT on your wrist: a watchOS-only AI assistant from Sindre Sorhus with no subscription.