Best AI tools for prompt management
30 tools in the Evaluation category, filtered to prompt management.

Braintrust
FeaturedEval, monitor, and improve AI products end-to-end.

Weights & Biases
The ML experiment tracker, now with LLM eval features.

Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.

Humanloop
Prompt management + evals for collaborative AI teams.

Athina AI
Collaborative LLM evaluation and observability platform for teams shipping AI features to production.

MLflow
Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.

OpenAI Evals
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.

W&B Weave
Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

PromptLayer
Lightweight prompt logging + management for OpenAI/Claude apps.

Langfuse
Open-source LLM observability, prompt management, and evaluation in one platform.

Opik
Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Kiln AI
Open-source workbench for building, evaluating, and fine-tuning AI agents across 190+ models.

Promptfoo
Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.

Weco AI
Autoresearch engine that iteratively rewrites code to optimize against a numeric evaluation metric.

Arthur
Open-source toolkit for testing, tracing, and monitoring production AI agents.

Maxim AI
End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.

Prompt Foundry
Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.

Respan (formerly Keywords AI)
LLM engineering platform combining a multi-model gateway with tracing, evals, and prompt management.

LangFast
No-signup LLM playground for testing, comparing, and versioning prompts against your own API keys.

llmfit
Terminal tool that scores hundreds of open LLMs against your actual CPU, RAM, and GPU and tells you which ones will run well.

Phoenix
Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.

Agenta
Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.

Parea AI
LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.

AI World Bakeoff
Ten AI coding models, three identical briefs, thirty explorable 3D worlds

Habibi
Self-hosted generative engine optimization for tracking brand mentions across ChatGPT, Perplexity, and Gemini

Lakera
Runtime security and guardrails for GenAI apps, agents, and RAG systems.

LangWatch
Simulation-based testing, evaluation, and observability for LLM agents

ModelBias
100 models, 100 prompts, 30,000 answers — an interactive look at AI defaults

ModelFuzz
Open-source red-teaming and execution-layer defense for AI agents against prompt injection.