Skip to main content
📖 The AI Tool Bible

AI prompt evaluation

Editorial picks for "prompt evaluation tool".

48 tools

All Evaluation
Kittl preview image
Kittl logo

Kittl

Image Generation · Multi-model: ByteDance Seedream 3-4.5, Ideogram 2A/3, Google Imagen 4 + Nano Banana, OpenAI DALL-E 3 / ChatGPT Image, Black Forest Labs Flux
8.5

AI-first design platform combining multi-model image and vector generation with a full browser editor

Freemium· Free: Free · Pro: $19 · Expert: $29 · Business: $59T-shirt and merch graphic designBadge and vintage-style logo creation
Dataiku preview image
Dataiku logo

Dataiku

Agents · Multi-model (LLM Mesh: OpenAI, Anthropic, Bedrock, Vertex, OSS)
8.3

Enterprise AI platform unifying data, ML, LLMs, and agents under one governed workflow.

Enterprise· Basic: $10 · Pro: $20 · Enterprise: Contact salesenterprise-aiagent-orchestration
Arize AI preview image
Arize AI logo

Arize AI

Evaluation · Multi-model
8.2

Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-observabilityagent-evaluation
OpenPipe preview image
OpenPipe logo

OpenPipe

Fine-tuning · Llama, Mistral, Qwen and other open-weight base models
8.2

Fine-tuning and reinforcement learning platform for turning expensive prompts into cheap, fast, task-specific models.

Freemium· Free tier available; usage-based pricing for training and hosted inference; enterprise plans on requestllm-cost-reductionfine-tuning
Athina AI preview image
Athina AI logo

Athina AI

Evaluation · Multi-model
8.1

Collaborative LLM evaluation and observability platform for teams shipping AI features to production.

Freemium· Starter free (10k logs/mo); Pro & Enterprise customllm-evaluationprompt-management
Berkeley Function-Calling Leaderboard preview image
Berkeley Function-Calling Leaderboard logo

Berkeley Function-Calling Leaderboard

Evaluation · Multi-model
8.1

Open benchmark from UC Berkeley that ranks LLMs on real-world tool-use and function-calling accuracy.

Free· Free and open source; you pay only for inference when reproducing runs.function-calling evaltool-use benchmarking
HoneyHive preview image
HoneyHive logo

HoneyHive

Evaluation · Multi-model
8.1

OpenTelemetry-native observability and evaluation platform for LLM agents in production.

Freemium· Free tier available; paid/enterprise tiers via salesagent-observabilityllm-evaluation
MLflow preview image
MLflow logo

MLflow

Evaluation · Multi-model
8.1

Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.

Free· Free and open source (Apache 2.0); managed offering via Databricksllm-evaluationexperiment-tracking
W&B Weave preview image
W&B Weave logo

W&B Weave

Evaluation · Multi-model
8.1

Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

Freemium· Free tier available; paid and enterprise plans via W&Bllm-tracingagent-observability
Langfuse preview image
Langfuse logo

Langfuse

Evaluation · Model-agnostic
7.3

Open-source LLM observability, prompt management, and evaluation in one platform.

Freemium· Free self-host & Hobby tier; Core $29/mo, Pro $199/mo, Enterprise $2,499/mollm-observabilityprompt-management
Opik preview image
Opik logo

Opik

Evaluation · Multi-model
7.3

Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Freemium· Free open-source self-host; free Cloud tier (no card); Enterprise contact salesllm-tracingagent-evaluation
Kiln AI preview image
Kiln AI logo

Kiln AI

Evaluation · Multi-model
7.2

Open-source workbench for building, evaluating, and fine-tuning AI agents across 190+ models.

Freemium· Free Individual tier; Team (request access); Enterprise (custom)llm-evaluationfine-tuning
LLM by Datasette preview image
LLM by Datasette logo

LLM by Datasette

Coding · Multi-model
7.2

A CLI and Python library for running prompts against any LLM provider and logging everything to SQLite.

Free· Free and open source (Apache 2.0); pay underlying model providers separatelycli-promptingprompt-logging
Arthur preview image
Arthur logo

Arthur

Evaluation · Multi-model
7.1

Open-source toolkit for testing, tracing, and monitoring production AI agents.

Freemium· Free: $0/mo · Premium: $60/mo · Enterprise: Customagent-evaluationprompt-management
Fiddler AI preview image
Fiddler AI logo

Fiddler AI

Evaluation · Fiddler Centor (proprietary evaluators)
7.1

Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.

Enterprise· Free: Free · Developer: $0.002 per trace · Enterprise: Contact salesllm-observabilityagent-monitoring
Prompt Foundry preview image
Prompt Foundry logo

Prompt Foundry

Evaluation · OpenAI + Anthropic (multi-model)
7.1

Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.

Freemium· Free tier (10 prompts, 500 evals/mo); Pro $15/user/mo; Enterprise customprompt-managementmodel-comparison
SEAL Leaderboard preview image
SEAL Leaderboard logo

SEAL Leaderboard

Evaluation · Multi-model (GPT, Claude, Gemini, Llama, etc.)
7.1

Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.

Free· Free to view; paid custom evals via Scale enterprise salesmodel-selectionbenchmark-tracking
Phoenix preview image
Phoenix logo

Phoenix

Evaluation · Multi-model
7.0

Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-tracingagent-debugging
Agenta preview image
Agenta logo

Agenta

Evaluation · Multi-model
6.9

Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.

Freemium· Hobby: $0 forever · Pro: $29 /month · Business: $299 /month · Enterprise: Custom · Open source: Free foreverprompt-engineeringllm-evaluation
TreeScale preview image
TreeScale logo

TreeScale

Agents · Multi-model
6.9

No-code platform that wraps LLM prompt chains into deployable, integration-ready APIs.

Freemium· Free tier to publish first LLM app; paid tiers on topllm-api-deploymentprompt-chaining
Dify preview image
Dify logo

Dify

Agents · Model-agnostic: OpenAI (GPT-4o, GPT-4.1), Anthropic Claude, Google Gemini, Mistral, Cohere, Ollama, and any OpenAI-compatible endpoint

Open-source LLMOps platform for building agentic workflows, RAG pipelines, and AI applications

Freemium· Professional: $590 · Team: $1590 · Sandbox: Free · Enterprise: Custom · Community: FreeRAG chatbot over internal documentsCustomer support automation
Google Agent Development Kit (ADK) preview image
Google Agent Development Kit (ADK) logo

Google Agent Development Kit (ADK)

Agents · Gemini (default) plus Claude, GPT-4/5, Llama, and other providers via LiteLLM

Google's open-source framework for building, evaluating, and deploying production AI agents

Free· Framework itself is free and open-source (Apache 2.0). Costs come from the underlying model provider (e.g. Gemini API / Vertex AI usage) and any hosting infrastructure (Cloud Run, GKE, Agent Engine).Multi-agent research assistantCustomer support triage agent
Lakera preview image
Lakera logo

Lakera

Evaluation · Proprietary in-house classifiers; model-agnostic (works in front of GPT-4o, Claude, Gemini, Llama, and custom LLMs)

Runtime security and guardrails for GenAI apps, agents, and RAG systems.

Freemium· Free community/developer tier at platform.lakera.ai; paid Enterprise plans (custom pricing, contact sales). No public price list.Prompt injection defense for chatbotsRAG guardrails against indirect injection
LangWatch preview image
LangWatch logo

LangWatch

Evaluation · Model-agnostic; supports OpenAI, Anthropic, AWS Bedrock, Azure OpenAI, Vertex AI, and any OpenTelemetry-instrumented LLM

Simulation-based testing, evaluation, and observability for LLM agents

Freemium· Developer: €0 · Growth: €29/ core-seat / month · Enterprise: CustomLLM agent regression testing in CIRAG answer-quality evaluation
LDBD Prediction Leaderboard preview image
LDBD Prediction Leaderboard logo

LDBD Prediction Leaderboard

Agents · Model-agnostic (users bring their own — Claude, GPT, Gemma, or custom); leaderboard shows entries from Claude, GPT and Gemma variants

Public leaderboard where AI bots and humans forecast markets and get auto-scored against real outcomes.

Free· Free to play. Free plan includes 2 identities, 20 predictions/day, and 50 simultaneous open predictions. No paid tier advertised.Benchmarking LLM trading agentsPublic track record for a custom prediction bot
Magic.dev preview image
Magic.dev logo

Magic.dev

Coding · In-house frontier code models (including a long-term-memory 'LTM' model family with reported 100M-token context)

Frontier code models with ultra-long context aimed at automating software engineering

Enterprise· No public pricing. Access is via research partnerships and enterprise engagements; no self-serve tier or public API published as of writing.Whole-repo refactorsLong-horizon feature implementation
Octomind preview image
Octomind logo

Octomind

Agents · Multi-provider: OpenAI, Anthropic (Claude), DeepSeek, Ollama, and 20+ others; benchmarks cite GLM-5.2 and Claude Opus

Homebrew for AI agents: install specialized, budget-capped AI specialists with one command.

Freemium· Free: $0 · Pro: $10 _first month_ → $20/mo · Max: $50 _first month_ → $100/mo · Team: $500/mo _flat, whole team_Domain-specialist coding agentsAutomated PR review and fixes
Open WebUI preview image
Open WebUI logo

Open WebUI

Agents · Backend-agnostic: any Ollama, llama.cpp, vLLM, or OpenAI-compatible API (OpenAI GPT, Anthropic Claude, Llama 3.x, Qwen, Mistral, Gemma, etc.)

Self-hosted, extensible AI chat platform that runs on your infrastructure

Freemium· Free (self-hosted, MIT-style community license via pip/Docker) / Enterprise: custom pricing for SSO, RBAC, audit logs, air-gapped deployment, data-residency guaranteesSelf-hosted ChatGPT alternative for a teamPrivate RAG chatbot over internal documents
Relevance AI preview image
Relevance AI logo

Relevance AI

Agents · Multi-model: Claude (Opus/Sonnet/Haiku), OpenAI GPT, Google Gemini, plus open-weight options (Kimi K2, GLM)

Build and deploy an AI workforce of specialized agents across your business tools

Enterprise· Free trial available via the app. Paid tiers are quote-based (Enterprise): custom actions, unlimited agents/tools/users, dedicated account manager. Reported customer benchmarks cite an average cost of ~$0.09 per task at scale; no fixed public tier pricing.Outbound prospect research and personalisationMeeting prep and CRM hygiene
The Email Game preview image
The Email Game logo

The Email Game

Agents · Bring-your-own (any LLM the participant chooses)

An arena for autonomous email agents.

Free· Free to enter. Prize pool of $1,700 total ($1,000 first place, $500 second, $200 third). Referral bonus of $15 per referred competitor up to $45.Multi-agent LLM benchmarkingAgent negotiation research
Suno preview image
Suno logo

Suno

Featured
Audio · Suno v4
9.2

Text-to-song AI — full vocal tracks from a prompt.

Freemium· Free Plan: $0 · Pro Plan: $8 · Premier Plan: $24songwritingdemos
Flux preview image
Flux logo

Flux

Featured
Image Generation · Flux.1 [schnell / dev / pro]
9.0

Black Forest Labs' open-weights image model — rivals Midjourney quality.

Freemium· FLUX.2 [max]: $0.07 · FLUX.2 [pro]: $0.03 · FLUX.2 [klein] 9B: $0.015 · FLUX.2 [klein] 4B: $0.014 · FLUX.2 [flex]: $0.05open sourceself-hosted
Runway preview image
Runway logo

Runway

Featured
Video · Gen-4
9.0

Pro-grade AI video editor and Gen-4 generation.

Paid· $15/mo Standard; $35/mo Pro; $95/mo Unlimitedshort filmVFX
Braintrust preview image
Braintrust logo

Braintrust

Featured
Evaluation · Platform (any LLM)
8.9

Eval, monitor, and improve AI products end-to-end.

Freemium· Starter: $0 · Pro: $249 · Enterprise: Custom pricingevalsmonitoring
Replit Agent preview image
Replit Agent logo

Replit Agent

Featured
Coding · Multi-model (Claude / GPT configurable)
8.7

Build & deploy a full app from a single prompt.

Freemium· Basic: $20 · Pro: $50 · Enterprise: Contact salesprototypesinternal tools
Stable Diffusion preview image
Stable Diffusion logo

Stable Diffusion

Image Generation · SD 3.5 / SDXL
8.8

Open-source image generation — run anywhere, fine-tune anything.

Free· Free open weights; optional Stability APIlocalfine-tuning
Ernie Bot preview image
Ernie Bot logo

Ernie Bot

Agents · Baidu ERNIE 4.0 / ERNIE X1 / ERNIE Turbo (in-house)
8.7

Baidu's Mandarin-first ChatGPT rival, powered by the ERNIE model family

Freemium· Free tier for Ernie 3.5 access; Ernie 4.0 and premium features require a paid subscription (approximately CNY 59.9/month for individual plans); enterprise API pricing via Baidu AI Cloud Qianfan platform is metered per 1K tokens.Mandarin content writing and marketing copyChinese-language document Q&A and summarisation
LangSmith preview image
LangSmith logo

LangSmith

Evaluation · Platform (any LLM)
8.7

LangChain's eval + observability platform.

Freemium· Developer: $0 · Plus: $39 · Enterprise: Custom pricingLLM tracingevals
Nano Banana (Gemini Image) preview image
Nano Banana (Gemini Image) logo

Nano Banana (Gemini Image)

Image Generation · Gemini 3 Pro Image (Nano Banana Pro), Gemini 3.1 Flash Image (Nano Banana 2), Gemini 3.1 Flash-Lite Image (Nano Banana 2 Lite)
8.7

Google DeepMind's Gemini-powered image generation and conversational editing model family

Paid· Consumer access via Gemini app (free tier + Google AI Pro/Ultra subscriptions). API usage-based: Nano Banana Pro (Gemini 3 Pro Image) ~$0.134/image at 1K-2K, ~$0.24/image at 4K; Nano Banana 2 (Gemini 3.1 Flash Image) ~$0.067/image at 1K, up to ~$0.151 at higher resolutions; Nano Banana 2 Lite priced lower for high-throughput use. Batch API roughly 50% off. Enterprise pricing via Gemini Enterprise Agent Platform and Vertex AI.Marketing hero imagesProduct mockups and packaging visualisations
Snowflake Cortex preview image
Snowflake Cortex logo

Snowflake Cortex

RAG · Anthropic Claude, Meta Llama, Mistral Large 2, Snowflake Arctic
8.7

Generative AI and RAG built into the Snowflake data cloud

Enterprise· Standard: Contact sales · Enterprise: Contact sales · Business Critical: Contact sales · Virtual Private Snowflake: Contact salesEnterprise RAG chatbot over governed dataNatural-language SQL for business analysts
AWS Bedrock preview image
AWS Bedrock logo

AWS Bedrock

Agents · Multi-model: Anthropic Claude, Meta Llama, Mistral, Cohere, AI21, Amazon Nova/Titan, DeepSeek, Stability, OpenAI GPT
8.6

Build and scale generative AI applications with foundation models

Paid· Standard: Contact sales · Flex: Contact sales · Priority: Contact sales · Reserved: Contact salesEnterprise RAG chatbot over private documentsMulti-step tool-using agents via AgentCore
IBM watsonx preview image
IBM watsonx logo

IBM watsonx

Agents · IBM Granite (3.x, Code, Time Series), Meta Llama 3.x, Mistral, plus other curated open models
8.6

Enterprise AI platform for building, deploying, and governing models and agents

Enterprise· watsonx.ai has a free tier on IBM Cloud with limited tokens; paid usage is metered per 1M tokens by model family (Granite, Llama, Mistral, etc.). watsonx.governance and watsonx.data are quoted per environment. Enterprise deals via IBM sales; on-prem/Cloud Pak for Data is separately licensed.Enterprise RAG chatbot over private documentsCustomer service agents with guardrails
MongoDB Atlas Vector Search preview image
MongoDB Atlas Vector Search logo

MongoDB Atlas Vector Search

RAG · Bring-your-own embeddings (OpenAI, Cohere, open models); native Voyage AI embeddings and rerankers
8.6

Vector search built into the operational database you're already using.

Freemium· Free: $0 · Flex: Up to $30 · Dedicated: Starts at $56.94RAG over enterprise documentsProduct and content recommendation engines
Canva Magic Studio preview image
Canva Magic Studio logo

Canva Magic Studio

Image Generation · Multi-model: partners including OpenAI (Magic Write historically on GPT models), Google Imagen and Runway for image/video, plus Canva's in-house design and layout models
8.5

Canva's all-in-one AI creative suite for design, image, video, copy, and presentations

Freemium· Free: Free · Pro: €11.67 · Business: €14.17 · Enterprise: Contact salesSocial media post generationShort-form video ads
Glean preview image
Glean logo

Glean

Agents · Model-agnostic: routes across 35+ LLMs including GPT-4o, Claude 3.5/4 Sonnet, Gemini 1.5/2, Llama 3, Mistral, plus Glean in-house models
8.5

Work AI platform that unifies enterprise knowledge, search, and agents

Enterprise· Enterprise pricing only; commonly reported in the $40-50/user/month range with a floor typically starting around 100 seats. No public self-serve tier. Contact sales for a quote.Enterprise search across Slack, Drive, Confluence and JiraCompany-wide AI assistant with citations
Amp (Sourcegraph) preview image
Amp (Sourcegraph) logo

Amp (Sourcegraph)

Coding · GPT-5.5, Claude Opus 4.8, GPT Image 2 (frontier multi-model)
8.4

Frontier coding agent from Sourcegraph with pass-through model pricing

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesMulti-file refactors across large repositoriesAgentic code generation with subagent parallelism
ScrapeGraphAI preview image
ScrapeGraphAI logo

ScrapeGraphAI

Agents · Multi-model (LLM, unspecified)
8.4

LLM-driven web scraping API that turns natural-language prompts into structured JSON.

Freemium· Free 500 credits; Starter $20/mo, Growth $100/mo, Pro $500/mo, Enterprise customweb-scrapingdata-extraction
Semantic Kernel preview image
Semantic Kernel logo

Semantic Kernel

Agents · Multi-model
8.4

Microsoft's open-source SDK for wiring LLMs, plugins, and agents into enterprise .NET, Python, and Java apps.

Free· Free, MIT-licensed SDK; you pay for the underlying model APIsagent-orchestrationllm-plugins