AI hallucination detection
Editorial picks for "detect llm hallucination".
23 tools
Braintrust
FeaturedEval, monitor, and improve AI products end-to-end.
Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.
Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.
Athina AI
Collaborative LLM evaluation and observability platform for teams shipping AI features to production.
Cube
Semantic layer that grounds LLM agents in your real business metrics instead of letting them hallucinate SQL.
HoneyHive
OpenTelemetry-native observability and evaluation platform for LLM agents in production.
TruLens
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.
W&B Weave
Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.
Patronus
Automated LLM evaluation for hallucinations, safety, and quality.
Lamini
Memory-tuning platform for grounding LLMs in your facts.
Arthur
Open-source toolkit for testing, tracing, and monitoring production AI agents.
Fiddler AI
Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.
Maxim AI
End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.
SEAL Leaderboard
Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.
Superwise
Agentic management platform for runtime guardrails, policy enforcement, and observability across LLM agents.
Cleanlab TLM
Trustworthiness scoring layer that flags LLM hallucinations in real time.
Parea AI
LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.
Prediction Guard
Self-hosted AI control plane that lets regulated enterprises govern models, agents, and MCP servers behind their firewall.
Lagotto Meter
Measure the gap between what your site claims and what an AI agent actually finds
Lakera
Runtime security and guardrails for GenAI apps, agents, and RAG systems.
LangWatch
Simulation-based testing, evaluation, and observability for LLM agents
TokenPath
Token-level citation and attribution API for AI-generated answers
Vectara
Enterprise agent platform with built-in retrieval, grounding, and hallucination controls