AI hallucination detection
Editorial picks for "detect llm hallucination".
48 tools

Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.

Patronus
Automated LLM evaluation for hallucinations, safety, and quality.

Arthur
Open-source toolkit for testing, tracing, and monitoring production AI agents.

Fiddler AI
Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.

Cleanlab TLM
Trustworthiness scoring layer that flags LLM hallucinations in real time.

DataStax Astra DB
Serverless vector and document database for production RAG and AI agents

MongoDB Atlas Vector Search
Vector search built into the operational database you're already using.

PyCaret
Low-code Python AutoML library that wraps scikit-learn, XGBoost, LightGBM and friends behind a few-line API.

ScrapeGraphAI
LLM-driven web scraping API that turns natural-language prompts into structured JSON.

H2O.ai
Enterprise AI platform combining AutoML, generative AI, and vertical agents for regulated industries.

Great Expectations
Open-source data quality framework for validating the datasets that feed your ML and analytics pipelines.

Grammarly
Cross-app AI writing assistant that catches errors, rewrites paragraphs, and adjusts tone in the tools you already use.

Pathway
Live data framework for production RAG and streaming ETL pipelines in Python.

Seldon
Kubernetes-native MLOps platform for deploying and orchestrating ML and generative AI models in production.

Plano
Envoy-based data plane for AI agents that handles routing, guardrails, and observability outside your app code.

SAS Viya
Enterprise-grade data and AI analytics platform with built-in governance, MCP server, and a Copilot for regulated industries.

Wallaroo.AI
Production AI inference platform for deploying and monitoring models across cloud, on-prem, and edge.

OpenDataLoader PDF
Open-source PDF parser built for RAG pipelines, with reading-order detection, table extraction, and bounding-box citations.

CodeRabbit
AI pull-request reviewer that lives inside GitHub, GitLab, and your IDE.

OlympicArena
Olympiad-level multi-discipline benchmark for stress-testing reasoning in LLMs and multimodal models.

Superduper
Enterprise AI agent orchestration that brings RAG and agents to your existing data stack without migration.

Superwise
Agentic management platform for runtime guardrails, policy enforcement, and observability across LLM agents.

Callstack.ai PR Reviewer
AI pull request reviewer that flags bugs, security issues, and style drift against your team's own guidelines.

ChatGPT for Google
Browser extension that pipes GPT-4, Claude and other LLM answers into your existing search results page.
Firecrawl MCP Server
MCP server that gives Claude, Cursor, and other AI agents native web scrape, crawl, search, and extract tools via Firecrawl.

getdebug
AI-powered codebase analyzer and auto-fixer that ships validated PRs

Goose
Open-source, local-first AI agent for code, workflows, and everything in between.
Grafana MCP
Official Grafana Labs MCP server — dashboards, Prometheus, Loki, alerts and incidents in your LLM client

Greptile
AI code review that understands the whole codebase, not just the diff

TiDB
AI-native distributed SQL database with built-in vector search, agent memory, and RAG pipelines

AssemblyAI
Speech-to-text API with diarisation, summarisation, and topic detection.

W&B Weave
Production observability, tracing, and evaluation for LLM and agent systems from the Weights & Biases stack.

ChatGPT
OpenAI's flagship conversational assistant, and the default benchmark every other chatbot is measured against.

Resemble.ai
Enterprise voice cloning with deepfake-detection layer.

Wispr Flow
System-wide voice-to-text dictation that auto-edits filler words and learns your jargon.

Lamini
Memory-tuning platform for grounding LLMs in your facts.

STORM
Stanford's open-source research agent that turns a topic into a Wikipedia-style article with citations.

Headroom
Open-source context compression layer that strips 70-95% of boilerplate before it hits your LLM.

MMagic
OpenMMLab's research-grade toolbox for image and video generation, restoration, and editing.

Opik
Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.

Cald.AI
Voice AI agents that run inbound and outbound phone calls with sub-second latency.

Qodo (formerly CodiumAI)
AI code review and test generation platform that plugs into IDEs, pull requests, and CLI with deep multi-repo context.

AnkiDecks AI
AI flashcard generator that turns PDFs, slides, YouTube videos and handwritten notes into Anki-ready decks.

Hyperbrowser
Cloud browser infrastructure built for AI agents that need to scrape, click, and navigate the live web.

PhotoPrism
Self-hosted, AI-powered photo library with face recognition and content-based search.

QuillBot
AI paraphrasing and grammar suite that grew into an all-in-one writing toolkit.

SiteSpeakAI
Custom-trained chatbot that turns your website, docs, and PDFs into a multilingual support and lead-gen agent.

ChatLab
Local-first analytics and AI agent for your personal chat history, with bring-your-own model support.