AI prompt evaluation
Editorial picks for "prompt evaluation tool".
48 tools
Braintrust
FeaturedEval, monitor, and improve AI products end-to-end.
LangSmith
LangChain's eval + observability platform.
Weights & Biases
The ML experiment tracker, now with LLM eval features.
Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.
Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.
Humanloop
Prompt management + evals for collaborative AI teams.
LiveBench
Contamination-free LLM benchmark that refreshes its questions monthly to keep frontier models honest.
OpenAI Playground
OpenAI's official browser sandbox for prompting, tuning, and testing every model on the platform before you ship API code.
Athina AI
Collaborative LLM evaluation and observability platform for teams shipping AI features to production.
Berkeley Function-Calling Leaderboard
Open benchmark from UC Berkeley that ranks LLMs on real-world tool-use and function-calling accuracy.
HoneyHive
OpenTelemetry-native observability and evaluation platform for LLM agents in production.
MLflow
Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.
OpenAI Evals
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.
TruLens
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.
W&B Weave
Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.
PromptHub
Git-style prompt management, testing, and deployment platform for teams running multiple LLMs in production.
Rivet
Open-source visual IDE for building and debugging LLM agent graphs.
LLM Stats
Live leaderboard and side-by-side comparison hub for 300+ frontier LLMs across reasoning, coding, and multimodal benchmarks.
PromptLayer
Lightweight prompt logging + management for OpenAI/Claude apps.
Patronus
Automated LLM evaluation for hallucinations, safety, and quality.
Langfuse
Open-source LLM observability, prompt management, and evaluation in one platform.
Opik
Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.
Inspect AI
Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.
Kiln AI
Open-source workbench for building, evaluating, and fine-tuning AI agents across 190+ models.
Promptfoo
Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.
Weco AI
Autoresearch engine that iteratively rewrites code to optimize against a numeric evaluation metric.
AlpacaEval
Automatic LLM evaluator and leaderboard that benchmarks instruction-following with length-controlled win rates.
Maxim AI
End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.
Prompt Foundry
Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.
Puzzlet AI
Git-native prompt management and observability platform for teams shipping LLM applications.
Respan (formerly Keywords AI)
LLM engineering platform combining a multi-model gateway with tracing, evals, and prompt management.
SEAL Leaderboard
Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.
VisualWebArena
Open benchmark for evaluating multimodal web agents on realistic visual browsing tasks.
LangFast
No-signup LLM playground for testing, comparing, and versioning prompts against your own API keys.
Phoenix
Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.
Prompteams
Git-style version control and testing for LLM prompts, with auto-generated APIs that ship updates without redeploys.
Agenta
Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.
CompassRank
Public leaderboard from the OpenCompass project ranking open and closed LLMs across 100+ benchmarks.
Izlo
Prompt management platform with version control, collaboration, and an API for production deployment.
MixEval
Dynamic LLM benchmark that mixes web queries with existing datasets to mirror Chatbot Arena rankings at a fraction of the cost.
Arena AI
Head-to-head LLM battle arena with a public leaderboard for ranking AI models.
Parea AI
LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.
PromptPerfect
Prompt optimizer from Jina AI that rewrites and stress-tests prompts across major LLMs.
Magic Potion
Visual drag-and-drop prompt editor for crafting, organizing, and reusing prompts across OpenAI, Anthropic, and Google models.
AI World Bakeoff
Ten AI coding models, three identical briefs, thirty explorable 3D worlds
Lagotto Meter
Measure the gap between what your site claims and what an AI agent actually finds
LangWatch
Simulation-based testing, evaluation, and observability for LLM agents
ModelBias
100 models, 100 prompts, 30,000 answers — an interactive look at AI defaults