Skip to main content
📖 The AI Tool Bible

AI safety testing

Editorial picks for "ai safety eval".

22 tools

All Evaluation →
AA

Arize AI

Evaluation · Multi-model
8.2

Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-observabilityagent-evaluation
GI

Giskard

Evaluation · Multi-model
8.2

Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.

Freemium· Open-source free tier; Giskard Hub enterprise pricing on requestllm-red-teamingagent-security-testing
OE

OpenAI Evals

Evaluation · OpenAI GPT models (extensible)
8.1

OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.

Free· Free (MIT); you pay OpenAI API costs for eval runsllm-benchmarkingregression-testing
TR

TruLens

Evaluation · Multi-model (LLM-as-judge)
8.1

Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.

Free· Free, open source (Apache-licensed Python package)llm-evaluationrag-evaluation
PA

Patronus

Evaluation · Platform (any LLM)
7.8

Automated LLM evaluation for hallucinations, safety, and quality.

Paid· Individual: Free · Base: $25 · Enterprise: Contact us for Pricinghallucination detectionsafety
IA

Inspect AI

Evaluation · Multi-model
7.2

Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.

Free· Free and open source (MIT-style license); you pay only for underlying model API usage.llm-benchmarkingagent-evaluation
PR

Promptfoo

Evaluation · Multi-model
7.2

Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.

Freemium· Community: Free · Enterprise: Custom · On-Premise: Customllm-evalsred-teaming
AR

Arthur

Evaluation · Multi-model
7.1

Open-source toolkit for testing, tracing, and monitoring production AI agents.

Freemium· Free: $0/mo · Premium: $60/mo · Enterprise: Customagent-evaluationprompt-management
FA

Fiddler AI

Evaluation · Fiddler Centor (proprietary evaluators)
7.1

Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.

Enterprise· Free: Free · Developer: $0.002 per trace · Enterprise: Contact salesllm-observabilityagent-monitoring
MA

Maxim AI

Evaluation · Multi-model
7.1

End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.

Freemium· Developer: Free · Professional: $29 /seat /month · Business: $49 /seat /month · Enterprise: Customagent-evaluationllm-observability
SL

SEAL Leaderboard

Evaluation · Multi-model (GPT, Claude, Gemini, Llama, etc.)
7.1

Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.

Free· Free to view; paid custom evals via Scale enterprise salesmodel-selectionbenchmark-tracking
VI

VisualWebArena

Evaluation · Model-agnostic (GPT-4V, Gemini, Claude, open VLMs)
7.1

Open benchmark for evaluating multimodal web agents on realistic visual browsing tasks.

Free· Free and open source (MIT-style research release)multimodal-agent-evalweb-browsing-benchmark
SU

Superwise

Evaluation · Multi-model
7.0

Agentic management platform for runtime guardrails, policy enforcement, and observability across LLM agents.

Freemium· Starter: Free · Pro+: ? · Enterprise: ?llm-guardrailsai-governance
GA

Guild AI

Agents · Multi-model (bring your own)
6.9

Control plane for deploying, governing, and auditing AI agents in production.

Freemium· Free: $0 · Individual: $20 · Team: $200 · Enterprise: Contact usagent deploymentagent governance
CT

Cleanlab TLM

Evaluation · Multi-model (wraps any LLM)
6.8

Trustworthiness scoring layer that flags LLM hallucinations in real time.

Freemium· Free tier for evaluation; usage-based API pricing; enterprise/private deployment via saleshallucination-detectionrag-evaluation
PG

Prediction Guard

Agents · Multi-model
6.6

Self-hosted AI control plane that lets regulated enterprises govern models, agents, and MCP servers behind their firewall.

Enterprise· Contact salesai-governanceself-hosted-llm
AX

Axtary

Agents

Content authorization and payload-binding for AI agents

Freemium· Local: $0 · Founding Team: $499 · Enterprise: CustomHuman-in-the-loop approval for AI code commitsGoverning MCP tool calls in Claude and Cursor
CW

Cloud World Model

Agents

Simulate AWS, GCP, Azure, OCI, and DigitalOcean infrastructure without provisioning real resources.

Freemium· Free tier: 1,000 credits/month auto-refreshed, no card required. Credit packs: Small $9, Medium $29, Large $79 (credits never expire). Usage: 1 credit/simulation step, 5 credits/chaos or multi-cloud call, 10 credits/AI explanation, read endpoints free.Agentic cloud architecture designMulti-cloud cost comparison
LA

Lakera

Evaluation · Proprietary in-house classifiers; model-agnostic (works in front of GPT-4o, Claude, Gemini, Llama, and custom LLMs)

Runtime security and guardrails for GenAI apps, agents, and RAG systems.

Freemium· Free community/developer tier at platform.lakera.ai; paid Enterprise plans (custom pricing, contact sales). No public price list.Prompt injection defense for chatbotsRAG guardrails against indirect injection
LA

LangWatch

Evaluation · Model-agnostic; supports OpenAI, Anthropic, AWS Bedrock, Azure OpenAI, Vertex AI, and any OpenTelemetry-instrumented LLM

Simulation-based testing, evaluation, and observability for LLM agents

Freemium· Developer: €0 · Growth: €29/ core-seat / month · Enterprise: CustomLLM agent regression testing in CIRAG answer-quality evaluation
MO

ModelBias

Evaluation · 100 models across Anthropic, OpenAI, Google, DeepSeek, Meta, xAI, Mistral, Qwen and others (via OpenRouter)

100 models, 100 prompts, 30,000 answers — an interactive look at AI defaults

Free· Free to browse and download the full dataset from GitHub.Comparing default model preferences across vendorsIllustrating RLHF homogenisation in talks and articles
MO

ModelFuzz

Evaluation · Model-agnostic; works with any OpenAI-compatible endpoint (Qwen 2.5 used in official examples).

Open-source red-teaming and execution-layer defense for AI agents against prompt injection.

Freemium· Free / open-source (MIT) via pip. Hosted dashboard with centralized policies, audit logs and continuous scanning coming soon via waitlist (pricing not yet public).Red-teaming OpenAI-compatible agent endpointsBlocking indirect prompt injection via retrieved documents