Skip to main content
📖 The AI Tool Bible

Best AI tools for safety testing

17 tools in the Evaluation category, filtered to safety testing.

All Evaluation
Arize AI preview image
Arize AI logo

Arize AI

Evaluation · Multi-model
8.2

Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-observabilityagent-evaluation
Giskard preview image
Giskard logo

Giskard

Evaluation · Multi-model
8.2

Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.

Freemium· Open-source free tier; Giskard Hub enterprise pricing on requestllm-red-teamingagent-security-testing
Great Expectations preview image
Great Expectations logo

Great Expectations

Evaluation
8.2

Open-source data quality framework for validating the datasets that feed your ML and analytics pipelines.

Freemium· Developer: Free · Team: Custom · Enterprise: Contact Salesdata-validationpipeline-testing
HoneyHive preview image
HoneyHive logo

HoneyHive

Evaluation · Multi-model
8.1

OpenTelemetry-native observability and evaluation platform for LLM agents in production.

Freemium· Free tier available; paid/enterprise tiers via salesagent-observabilityllm-evaluation
OpenAI Evals preview image
OpenAI Evals logo

OpenAI Evals

Evaluation · OpenAI GPT models (extensible)
8.1

OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.

Free· Free (MIT); you pay OpenAI API costs for eval runsllm-benchmarkingregression-testing
TruLens preview image
TruLens logo

TruLens

Evaluation · Multi-model (LLM-as-judge)
8.1

Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.

Free· Free, open source (Apache-licensed Python package)llm-evaluationrag-evaluation
W&B Weave preview image
W&B Weave logo

W&B Weave

Evaluation · Multi-model
8.1

Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

Freemium· Free tier available; paid and enterprise plans via W&Bllm-tracingagent-observability
Patronus preview image
Patronus logo

Patronus

Evaluation · Platform (any LLM)
7.8

Automated LLM evaluation for hallucinations, safety, and quality.

Paid· Individual: Free · Base: $25 · Enterprise: Contact us for Pricinghallucination detectionsafety
Opik preview image
Opik logo

Opik

Evaluation · Multi-model
7.3

Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Freemium· Free open-source self-host; free Cloud tier (no card); Enterprise contact salesllm-tracingagent-evaluation
Inspect AI preview image
Inspect AI logo

Inspect AI

Evaluation · Multi-model
7.2

Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.

Free· Free and open source (MIT-style license); you pay only for underlying model API usage.llm-benchmarkingagent-evaluation
LLMEval preview image
LLMEval logo

LLMEval

Evaluation · Multi-model
7.2

Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.

Free· Free; open-source academic benchmarksllm-benchmarkingacademic-evaluation
Promptfoo preview image
Promptfoo logo

Promptfoo

Evaluation · Multi-model
7.2

Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.

Freemium· Community: Free · Enterprise: Custom · On-Premise: Customllm-evalsred-teaming
Prompt Foundry preview image
Prompt Foundry logo

Prompt Foundry

Evaluation · OpenAI + Anthropic (multi-model)
7.1

Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.

Freemium· Free tier (10 prompts, 500 evals/mo); Pro $15/user/mo; Enterprise customprompt-managementmodel-comparison
LangFast preview image
LangFast logo

LangFast

Evaluation · Multi-model
7.0

No-signup LLM playground for testing, comparing, and versioning prompts against your own API keys.

Paid· One-time lifetime ~$60-$120; 14-day money-backprompt-testingprompt-versioning
Habibi preview image
Habibi logo

Habibi

Evaluation · GPT-4o, Perplexity Sonar, Gemini (bring-your-own API keys)

Self-hosted generative engine optimization for tracking brand mentions across ChatGPT, Perplexity, and Gemini

Freemium· Opsily Server: €40AI answer engine visibility trackingChatGPT brand mention monitoring
LangWatch preview image
LangWatch logo

LangWatch

Evaluation · Model-agnostic; supports OpenAI, Anthropic, AWS Bedrock, Azure OpenAI, Vertex AI, and any OpenTelemetry-instrumented LLM

Simulation-based testing, evaluation, and observability for LLM agents

Freemium· Developer: €0 · Growth: €29/ core-seat / month · Enterprise: CustomLLM agent regression testing in CIRAG answer-quality evaluation
ModelFuzz preview image
ModelFuzz logo

ModelFuzz

Evaluation · Model-agnostic; works with any OpenAI-compatible endpoint (Qwen 2.5 used in official examples).

Open-source red-teaming and execution-layer defense for AI agents against prompt injection.

Freemium· Free / open-source (MIT) via pip. Hosted dashboard with centralized policies, audit logs and continuous scanning coming soon via waitlist (pricing not yet public).Red-teaming OpenAI-compatible agent endpointsBlocking indirect prompt injection via retrieved documents