Best AI tools for safety testing
17 tools in the Evaluation category, filtered to safety testing.

Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.

Great Expectations
Open-source data quality framework for validating the datasets that feed your ML and analytics pipelines.

HoneyHive
OpenTelemetry-native observability and evaluation platform for LLM agents in production.

OpenAI Evals
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.

TruLens
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.

W&B Weave
Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

Patronus
Automated LLM evaluation for hallucinations, safety, and quality.

Opik
Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Inspect AI
Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.

LLMEval
Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.

Promptfoo
Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.

Prompt Foundry
Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.

LangFast
No-signup LLM playground for testing, comparing, and versioning prompts against your own API keys.

Habibi
Self-hosted generative engine optimization for tracking brand mentions across ChatGPT, Perplexity, and Gemini

LangWatch
Simulation-based testing, evaluation, and observability for LLM agents

ModelFuzz
Open-source red-teaming and execution-layer defense for AI agents against prompt injection.