All AI tools
59 tools
Braintrust
FeaturedEval, monitor, and improve AI products end-to-end.
LangSmith
LangChain's eval + observability platform.
Weights & Biases
The ML experiment tracker, now with LLM eval features.
Helicone
Open-source LLM observability — one-line proxy install.
Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.
Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.
Great Expectations
Open-source data quality framework for validating the datasets that feed your ML and analytics pipelines.
Humanloop
Prompt management + evals for collaborative AI teams.
LiveBench
Contamination-free LLM benchmark that refreshes its questions monthly to keep frontier models honest.
Athina AI
Collaborative LLM evaluation and observability platform for teams shipping AI features to production.
Berkeley Function-Calling Leaderboard
Open benchmark from UC Berkeley that ranks LLMs on real-world tool-use and function-calling accuracy.
HoneyHive
OpenTelemetry-native observability and evaluation platform for LLM agents in production.
MLflow
Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.
OpenAI Evals
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.
TruLens
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.
W&B Weave
Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.
LLM Stats
Live leaderboard and side-by-side comparison hub for 300+ frontier LLMs across reasoning, coding, and multimodal benchmarks.
PromptLayer
Lightweight prompt logging + management for OpenAI/Claude apps.
Patronus
Automated LLM evaluation for hallucinations, safety, and quality.
Langfuse
Open-source LLM observability, prompt management, and evaluation in one platform.
MathEval
Holistic benchmark suite for evaluating mathematical reasoning in large language models.
Opik
Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.
Inspect AI
Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.
Kiln AI
Open-source workbench for building, evaluating, and fine-tuning AI agents across 190+ models.