AI safety testing
Editorial picks for "ai safety eval".
39 tools

OpenAI Evals
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.

Inspect AI
Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.

Google Agent Development Kit (ADK)
Google's open-source framework for building, evaluating, and deploying production AI agents

LangWatch
Simulation-based testing, evaluation, and observability for LLM agents

Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.

Google AI Studio
Browser-based playground and API console for prototyping with Google's Gemini models.

Great Expectations
Open-source data quality framework for validating the datasets that feed your ML and analytics pipelines.

OpenAI Playground
OpenAI's official browser sandbox for prompting, tuning, and testing every model on the platform before you ship API code.

HoneyHive
OpenTelemetry-native observability and evaluation platform for LLM agents in production.

MLflow
Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.

TruLens
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.

W&B Weave
Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

PromptHub
Git-style prompt management, testing, and deployment platform for teams running multiple LLMs in production.

Patronus
Automated LLM evaluation for hallucinations, safety, and quality.

Gorilla
Open-source LLM purpose-built for function calling and API invocation across thousands of tools.

Opik
Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

LLMEval
Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.

Promptfoo
Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.

SAS Viya
Enterprise-grade data and AI analytics platform with built-in governance, MCP server, and a Copilot for regulated industries.

Wallaroo.AI
Production AI inference platform for deploying and monitoring models across cloud, on-prem, and edge.

AlpacaEval
Automatic LLM evaluator and leaderboard that benchmarks instruction-following with length-controlled win rates.

Arthur
Open-source toolkit for testing, tracing, and monitoring production AI agents.

Prompt Foundry
Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.

Respan (formerly Keywords AI)
LLM engineering platform combining a multi-model gateway with tracing, evals, and prompt management.

SEAL Leaderboard
Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.

OlympicArena
Olympiad-level multi-discipline benchmark for stress-testing reasoning in LLMs and multimodal models.

PySpur
Open-source agent builder with a drag-and-drop canvas, Python escape hatch, and a built-in test harness.

InfiBench
Stack Overflow-derived benchmark for evaluating code LLMs on real-world programming questions.

Parea AI
LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.

Ailin
One prompt, every AI model — plus a no-code builder for agents and workflows

Dify
Open-source LLMOps platform for building agentic workflows, RAG pipelines, and AI applications

Greptile
AI code review that understands the whole codebase, not just the diff

Nomic Atlas
Interactive maps and embeddings for unstructured text, image, and multimodal data.

Relevance AI
Build and deploy an AI workforce of specialized agents across your business tools

Retell AI
Build, test and deploy production-grade AI voice agents for inbound and outbound calls.

Same.new
Prompt-first AI web app builder that turns a URL or description into a live, hosted full-stack app.

The Email Game
An arena for autonomous email agents.

Vectara
Enterprise agent platform with built-in retrieval, grounding, and hallucination controls