Best AI tools for safety testing
19 tools in the Evaluation category, filtered to safety testing.
Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.
Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.
OpenAI Evals
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.
TruLens
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.
Patronus
Automated LLM evaluation for hallucinations, safety, and quality.
Inspect AI
Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.
LLMEval
Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.
Promptfoo
Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.
Arthur
Open-source toolkit for testing, tracing, and monitoring production AI agents.
Fiddler AI
Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.
Maxim AI
End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.
SEAL Leaderboard
Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.
VisualWebArena
Open benchmark for evaluating multimodal web agents on realistic visual browsing tasks.
OlympicArena
Olympiad-level multi-discipline benchmark for stress-testing reasoning in LLMs and multimodal models.
Superwise
Agentic management platform for runtime guardrails, policy enforcement, and observability across LLM agents.
Cleanlab TLM
Trustworthiness scoring layer that flags LLM hallucinations in real time.
Lakera
Runtime security and guardrails for GenAI apps, agents, and RAG systems.
ModelBias
100 models, 100 prompts, 30,000 answers — an interactive look at AI defaults
ModelFuzz
Open-source red-teaming and execution-layer defense for AI agents against prompt injection.