Skip to main content
📖 The AI Tool Bible

AI safety testing

Editorial picks for "ai safety eval".

39 tools

All Evaluation
OpenAI Evals preview image
OpenAI Evals logo

OpenAI Evals

Evaluation · OpenAI GPT models (extensible)
8.1

OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.

Free· Free (MIT); you pay OpenAI API costs for eval runsllm-benchmarkingregression-testing
Inspect AI preview image
Inspect AI logo

Inspect AI

Evaluation · Multi-model
7.2

Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.

Free· Free and open source (MIT-style license); you pay only for underlying model API usage.llm-benchmarkingagent-evaluation
Google Agent Development Kit (ADK) preview image
Google Agent Development Kit (ADK) logo

Google Agent Development Kit (ADK)

Agents · Gemini (default) plus Claude, GPT-4/5, Llama, and other providers via LiteLLM

Google's open-source framework for building, evaluating, and deploying production AI agents

Free· Framework itself is free and open-source (Apache 2.0). Costs come from the underlying model provider (e.g. Gemini API / Vertex AI usage) and any hosting infrastructure (Cloud Run, GKE, Agent Engine).Multi-agent research assistantCustomer support triage agent
LangWatch preview image
LangWatch logo

LangWatch

Evaluation · Model-agnostic; supports OpenAI, Anthropic, AWS Bedrock, Azure OpenAI, Vertex AI, and any OpenTelemetry-instrumented LLM

Simulation-based testing, evaluation, and observability for LLM agents

Freemium· Developer: €0 · Growth: €29/ core-seat / month · Enterprise: CustomLLM agent regression testing in CIRAG answer-quality evaluation
Arize AI preview image
Arize AI logo

Arize AI

Evaluation · Multi-model
8.2

Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-observabilityagent-evaluation
Giskard preview image
Giskard logo

Giskard

Evaluation · Multi-model
8.2

Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.

Freemium· Open-source free tier; Giskard Hub enterprise pricing on requestllm-red-teamingagent-security-testing
Google AI Studio preview image
Google AI Studio logo

Google AI Studio

Coding · Gemini 2.5 Pro / Flash, Imagen, Veo
8.2

Browser-based playground and API console for prototyping with Google's Gemini models.

Freemium· Free tier with rate limits; paid via Gemini API usage-based pricingprompt-prototypinggemini-api-keys
Great Expectations preview image
Great Expectations logo

Great Expectations

Evaluation
8.2

Open-source data quality framework for validating the datasets that feed your ML and analytics pipelines.

Freemium· Developer: Free · Team: Custom · Enterprise: Contact Salesdata-validationpipeline-testing
OpenAI Playground preview image
OpenAI Playground logo

OpenAI Playground

Writing · Multi-model (GPT-4o, GPT-4.1, o-series, DALL-E, Whisper, TTS)
8.2

OpenAI's official browser sandbox for prompting, tuning, and testing every model on the platform before you ship API code.

Paid· Basic: $10 · Pro: $20 · Enterprise: Contact salesprompt-engineeringmodel-comparison
HoneyHive preview image
HoneyHive logo

HoneyHive

Evaluation · Multi-model
8.1

OpenTelemetry-native observability and evaluation platform for LLM agents in production.

Freemium· Free tier available; paid/enterprise tiers via salesagent-observabilityllm-evaluation
MLflow preview image
MLflow logo

MLflow

Evaluation · Multi-model
8.1

Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.

Free· Free and open source (Apache 2.0); managed offering via Databricksllm-evaluationexperiment-tracking
TruLens preview image
TruLens logo

TruLens

Evaluation · Multi-model (LLM-as-judge)
8.1

Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.

Free· Free, open source (Apache-licensed Python package)llm-evaluationrag-evaluation
W&B Weave preview image
W&B Weave logo

W&B Weave

Evaluation · Multi-model
8.1

Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

Freemium· Free tier available; paid and enterprise plans via W&Bllm-tracingagent-observability
PromptHub preview image
PromptHub logo

PromptHub

Writing · Multi-model (OpenAI, Anthropic, Google, Meta, Mistral, Bedrock, Azure)
8.0

Git-style prompt management, testing, and deployment platform for teams running multiple LLMs in production.

Freemium· Free signup; paid team plans (contact sales / in-app)prompt-managementprompt-versioning
Patronus preview image
Patronus logo

Patronus

Evaluation · Platform (any LLM)
7.8

Automated LLM evaluation for hallucinations, safety, and quality.

Paid· Individual: Free · Base: $25 · Enterprise: Contact us for Pricinghallucination detectionsafety
Gorilla preview image
Gorilla logo

Gorilla

Agents · gorilla-openfunctions-v2 (6.91B)
7.3

Open-source LLM purpose-built for function calling and API invocation across thousands of tools.

Free· Free and Apache 2.0; self-hostedfunction-callingtool-use
Opik preview image
Opik logo

Opik

Evaluation · Multi-model
7.3

Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Freemium· Free open-source self-host; free Cloud tier (no card); Enterprise contact salesllm-tracingagent-evaluation
LLMEval preview image
LLMEval logo

LLMEval

Evaluation · Multi-model
7.2

Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.

Free· Free; open-source academic benchmarksllm-benchmarkingacademic-evaluation
Promptfoo preview image
Promptfoo logo

Promptfoo

Evaluation · Multi-model
7.2

Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.

Freemium· Community: Free · Enterprise: Custom · On-Premise: Customllm-evalsred-teaming
SAS Viya preview image
SAS Viya logo

SAS Viya

Agents · Multi-model
7.2

Enterprise-grade data and AI analytics platform with built-in governance, MCP server, and a Copilot for regulated industries.

Enterprise· Contact sales; 14-day free trialenterprise-analyticsai-governance
Wallaroo.AI preview image
Wallaroo.AI logo

Wallaroo.AI

Agents · Multi-model
7.2

Production AI inference platform for deploying and monitoring models across cloud, on-prem, and edge.

Enterprise· Starter: $500 · Wallaroo Community Edition: Free · Ampere Community Edition: Freemodel-deploymentmlops
AlpacaEval preview image
AlpacaEval logo

AlpacaEval

Evaluation · GPT-4 Preview (Nov 2024) as annotator
7.1

Automatic LLM evaluator and leaderboard that benchmarks instruction-following with length-controlled win rates.

Free· Free and open-source; pay only for the underlying OpenAI annotator API callsllm-benchmarkinginstruction-following eval
Arthur preview image
Arthur logo

Arthur

Evaluation · Multi-model
7.1

Open-source toolkit for testing, tracing, and monitoring production AI agents.

Freemium· Free: $0/mo · Premium: $60/mo · Enterprise: Customagent-evaluationprompt-management
Prompt Foundry preview image
Prompt Foundry logo

Prompt Foundry

Evaluation · OpenAI + Anthropic (multi-model)
7.1

Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.

Freemium· Free tier (10 prompts, 500 evals/mo); Pro $15/user/mo; Enterprise customprompt-managementmodel-comparison
Respan (formerly Keywords AI) preview image
Respan (formerly Keywords AI) logo

Respan (formerly Keywords AI)

Evaluation · Multi-model (500+ via gateway)
7.1

LLM engineering platform combining a multi-model gateway with tracing, evals, and prompt management.

Freemium· Free tier; paid plans (pricing not public); enterprise on requestllm-observabilityprompt-management
SEAL Leaderboard preview image
SEAL Leaderboard logo

SEAL Leaderboard

Evaluation · Multi-model (GPT, Claude, Gemini, Llama, etc.)
7.1

Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.

Free· Free to view; paid custom evals via Scale enterprise salesmodel-selectionbenchmark-tracking
OlympicArena preview image
OlympicArena logo

OlympicArena

Evaluation · GPT-4o, Claude-3.5-Sonnet, Doubao-Pro-32k, DeepSeek-Coder-V2, Qwen2-72B-Instruct
7.0

Olympiad-level multi-discipline benchmark for stress-testing reasoning in LLMs and multimodal models.

Free· Free, open-source research benchmarkllm-evaluationmultimodal-eval
PySpur preview image
PySpur logo

PySpur

Agents · Multi-model
7.0

Open-source agent builder with a drag-and-drop canvas, Python escape hatch, and a built-in test harness.

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesagent-orchestrationagent-evaluation
InfiBench preview image
InfiBench logo

InfiBench

Evaluation
6.9

Stack Overflow-derived benchmark for evaluating code LLMs on real-world programming questions.

Free· Free and open source (CC BY-SA 4.0)code-llm-evalmodel-benchmarking
Parea AI preview image
Parea AI logo

Parea AI

Evaluation · Multi-model
6.8

LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.

Freemium· Free (2 seats, 3k logs/mo); Team $150/mo; Enterprise customllm-evaluationprompt-management
Ailin preview image
Ailin logo

Ailin

Agents · Multi-model — routes to OpenAI, Anthropic Claude, Google Gemini, and customer-supplied custom models

One prompt, every AI model — plus a no-code builder for agents and workflows

Paid· Basic: $20 · Pro: $50 · Enterprise: Contact salesinternal knowledge-base chatbotmulti-model prompt A/B testing
Dify preview image
Dify logo

Dify

Agents · Model-agnostic: OpenAI (GPT-4o, GPT-4.1), Anthropic Claude, Google Gemini, Mistral, Cohere, Ollama, and any OpenAI-compatible endpoint

Open-source LLMOps platform for building agentic workflows, RAG pipelines, and AI applications

Freemium· Professional: $590 · Team: $1590 · Sandbox: Free · Enterprise: Custom · Community: FreeRAG chatbot over internal documentsCustomer support automation
Greptile preview image
Greptile logo

Greptile

Coding

AI code review that understands the whole codebase, not just the diff

Paid· Starter: Free · Pro: $30/seat/month · Enterprise: Custom pricingAutomated pull request reviewCross-file bug detection
Nomic Atlas preview image
Nomic Atlas logo

Nomic Atlas

RAG · nomic-embed-text-v1.5, nomic-embed-vision-v1.5 (in-house open-weights); optional integrations with OpenAI, Cohere, and other embedding providers

Interactive maps and embeddings for unstructured text, image, and multimodal data.

Freemium· Starter: Free · Plus: $10/month · Business: $125/seat/month · Enterprise: Custom solutions for security-first organizationsRAG corpus exploration and debuggingEmbedding quality auditing
Relevance AI preview image
Relevance AI logo

Relevance AI

Agents · Multi-model: Claude (Opus/Sonnet/Haiku), OpenAI GPT, Google Gemini, plus open-weight options (Kimi K2, GLM)

Build and deploy an AI workforce of specialized agents across your business tools

Enterprise· Free trial available via the app. Paid tiers are quote-based (Enterprise): custom actions, unlimited agents/tools/users, dedicated account manager. Reported customer benchmarks cite an average cost of ~$0.09 per task at scale; no fixed public tier pricing.Outbound prospect research and personalisationMeeting prep and CRM hygiene
Retell AI preview image
Retell AI logo

Retell AI

Audio · LLM-agnostic (OpenAI, Anthropic, Google and others selectable per agent); proprietary turn-taking model and voice pipeline in-house

Build, test and deploy production-grade AI voice agents for inbound and outbound calls.

Freemium· Pay-as-you-go: $0.07-$0.31/minute · Enterprise: Custom PricingAI inbound call answeringOutbound lead qualification
Same.new preview image
Same.new logo

Same.new

Coding · Undisclosed premium frontier models (paid tiers)

Prompt-first AI web app builder that turns a URL or description into a live, hosted full-stack app.

Freemium· Free (500K tokens/mo) / Basic $10/mo (2M tokens) / Pro $25/mo (5M tokens) / Max $50/mo (10M tokens) / Ultra $100/mo (20M tokens, $10 per additional 2M overage). Tokens do not roll over.Cloning a reference site's UI as a starting templatePrototyping SaaS onboarding flows
The Email Game preview image
The Email Game logo

The Email Game

Agents · Bring-your-own (any LLM the participant chooses)

An arena for autonomous email agents.

Free· Free to enter. Prize pool of $1,700 total ($1,000 first place, $500 second, $200 third). Referral bonus of $15 per referred competitor up to $45.Multi-agent LLM benchmarkingAgent negotiation research
Vectara preview image
Vectara logo

Vectara

RAG · In-house Boomerang (retrieval) and Mockingbird (generation) plus BYOM for GPT, Claude, Gemini, and open-weight LLMs

Enterprise agent platform with built-in retrieval, grounding, and hallucination controls

Enterprise· SaaS: $100K/ year · VPC: $250K/ year · On-prem: $500K/ yearEnterprise knowledge-base searchGrounded customer-support chatbots