AI safety testing
Editorial picks for "ai safety eval".
22 tools
Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.
Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.
OpenAI Evals
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.
TruLens
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.
Patronus
Automated LLM evaluation for hallucinations, safety, and quality.
Inspect AI
Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.
Promptfoo
Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.
Arthur
Open-source toolkit for testing, tracing, and monitoring production AI agents.
Fiddler AI
Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.
Maxim AI
End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.
SEAL Leaderboard
Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.
VisualWebArena
Open benchmark for evaluating multimodal web agents on realistic visual browsing tasks.
Superwise
Agentic management platform for runtime guardrails, policy enforcement, and observability across LLM agents.
Guild AI
Control plane for deploying, governing, and auditing AI agents in production.
Cleanlab TLM
Trustworthiness scoring layer that flags LLM hallucinations in real time.
Prediction Guard
Self-hosted AI control plane that lets regulated enterprises govern models, agents, and MCP servers behind their firewall.
Axtary
Content authorization and payload-binding for AI agents
Cloud World Model
Simulate AWS, GCP, Azure, OCI, and DigitalOcean infrastructure without provisioning real resources.
Lakera
Runtime security and guardrails for GenAI apps, agents, and RAG systems.
LangWatch
Simulation-based testing, evaluation, and observability for LLM agents
ModelBias
100 models, 100 prompts, 30,000 answers — an interactive look at AI defaults
ModelFuzz
Open-source red-teaming and execution-layer defense for AI agents against prompt injection.