AI prompt evaluation
Editorial picks for "prompt evaluation tool".
48 tools

Kittl
AI-first design platform combining multi-model image and vector generation with a full browser editor

Dataiku
Enterprise AI platform unifying data, ML, LLMs, and agents under one governed workflow.

Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.

OpenPipe
Fine-tuning and reinforcement learning platform for turning expensive prompts into cheap, fast, task-specific models.

Athina AI
Collaborative LLM evaluation and observability platform for teams shipping AI features to production.

Berkeley Function-Calling Leaderboard
Open benchmark from UC Berkeley that ranks LLMs on real-world tool-use and function-calling accuracy.

HoneyHive
OpenTelemetry-native observability and evaluation platform for LLM agents in production.

MLflow
Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.

W&B Weave
Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

Langfuse
Open-source LLM observability, prompt management, and evaluation in one platform.

Opik
Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Kiln AI
Open-source workbench for building, evaluating, and fine-tuning AI agents across 190+ models.

LLM by Datasette
A CLI and Python library for running prompts against any LLM provider and logging everything to SQLite.

Arthur
Open-source toolkit for testing, tracing, and monitoring production AI agents.

Fiddler AI
Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.

Prompt Foundry
Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.

SEAL Leaderboard
Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.

Phoenix
Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.

Agenta
Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.

TreeScale
No-code platform that wraps LLM prompt chains into deployable, integration-ready APIs.

Dify
Open-source LLMOps platform for building agentic workflows, RAG pipelines, and AI applications

Google Agent Development Kit (ADK)
Google's open-source framework for building, evaluating, and deploying production AI agents

Lakera
Runtime security and guardrails for GenAI apps, agents, and RAG systems.

LangWatch
Simulation-based testing, evaluation, and observability for LLM agents

LDBD Prediction Leaderboard
Public leaderboard where AI bots and humans forecast markets and get auto-scored against real outcomes.

Magic.dev
Frontier code models with ultra-long context aimed at automating software engineering

Octomind
Homebrew for AI agents: install specialized, budget-capped AI specialists with one command.

Open WebUI
Self-hosted, extensible AI chat platform that runs on your infrastructure

Relevance AI
Build and deploy an AI workforce of specialized agents across your business tools

The Email Game
An arena for autonomous email agents.

Suno
FeaturedText-to-song AI — full vocal tracks from a prompt.

Flux
FeaturedBlack Forest Labs' open-weights image model — rivals Midjourney quality.

Runway
FeaturedPro-grade AI video editor and Gen-4 generation.

Braintrust
FeaturedEval, monitor, and improve AI products end-to-end.

Replit Agent
FeaturedBuild & deploy a full app from a single prompt.

Stable Diffusion
Open-source image generation — run anywhere, fine-tune anything.

Ernie Bot
Baidu's Mandarin-first ChatGPT rival, powered by the ERNIE model family

LangSmith
LangChain's eval + observability platform.

Nano Banana (Gemini Image)
Google DeepMind's Gemini-powered image generation and conversational editing model family

Snowflake Cortex
Generative AI and RAG built into the Snowflake data cloud

AWS Bedrock
Build and scale generative AI applications with foundation models

IBM watsonx
Enterprise AI platform for building, deploying, and governing models and agents

MongoDB Atlas Vector Search
Vector search built into the operational database you're already using.

Canva Magic Studio
Canva's all-in-one AI creative suite for design, image, video, copy, and presentations

Glean
Work AI platform that unifies enterprise knowledge, search, and agents

Amp (Sourcegraph)
Frontier coding agent from Sourcegraph with pass-through model pricing

ScrapeGraphAI
LLM-driven web scraping API that turns natural-language prompts into structured JSON.

Semantic Kernel
Microsoft's open-source SDK for wiring LLMs, plugins, and agents into enterprise .NET, Python, and Java apps.