AI eval dataset builder
Editorial picks for "llm eval datasets".
48 tools

MathEval
Holistic benchmark suite for evaluating mathematical reasoning in large language models.

LLMEval
Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.

InfiBench
Stack Overflow-derived benchmark for evaluating code LLMs on real-world programming questions.

LangSmith
LangChain's eval + observability platform.

Great Expectations
Open-source data quality framework for validating the datasets that feed your ML and analytics pipelines.

LiveBench
Contamination-free LLM benchmark that refreshes its questions monthly to keep frontier models honest.

Athina AI
Collaborative LLM evaluation and observability platform for teams shipping AI features to production.

HoneyHive
OpenTelemetry-native observability and evaluation platform for LLM agents in production.

Hugging Face AutoTrain
No-code fine-tuning and training pipeline that spins up state-of-the-art models on the Hugging Face Hub.

CAMEL-AI
Open-source Python framework for building multi-agent systems and synthetic data pipelines.

Valohai
MLOps platform for versioned pipelines, distributed training, and LLM evaluation across any cloud.

Inspect AI
Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.

Kiln AI
Open-source workbench for building, evaluating, and fine-tuning AI agents across 190+ models.

Maxim AI
End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.

Forefront
Fine-tune and serve open-source LLMs on your own data without managing GPUs.

OlympicArena
Olympiad-level multi-discipline benchmark for stress-testing reasoning in LLMs and multimodal models.

Phoenix
Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.

CompassRank
Public leaderboard from the OpenCompass project ranking open and closed LLMs across 100+ benchmarks.

MixEval
Dynamic LLM benchmark that mixes web queries with existing datasets to mirror Chatbot Arena rankings at a fraction of the cost.

ClickHouse
The open-source columnar database powering real-time analytics — and, increasingly, LLM observability and RAG backends.

Language Model Builder
Learn how LLMs work by building one on your Mac

LangWatch
Simulation-based testing, evaluation, and observability for LLM agents

Braintrust
FeaturedEval, monitor, and improve AI products end-to-end.

IBM watsonx
Enterprise AI platform for building, deploying, and governing models and agents

Glean
Work AI platform that unifies enterprise knowledge, search, and agents

Yi (01.AI)
Foundation models from 01.AI — open-weight Yi family plus frontier Yi-Lightning and Yi-Large

ClearML
End-to-end MLOps and GenAI platform with open-source experiment tracking and enterprise GPU orchestration.

Dataiku
Enterprise AI platform unifying data, ML, LLMs, and agents under one governed workflow.
Helicone
Open-source LLM observability — one-line proxy install.

Humanloop
Prompt management + evals for collaborative AI teams.

LanceDB
Open-source multimodal lakehouse and vector database built for AI training and retrieval at petabyte scale.

OpenPipe
Fine-tuning and reinforcement learning platform for turning expensive prompts into cheap, fast, task-specific models.

Writer
Enterprise generative AI platform built around in-house Palmyra LLMs for regulated, brand-consistent content.

Berkeley Function-Calling Leaderboard
Open benchmark from UC Berkeley that ranks LLMs on real-world tool-use and function-calling accuracy.

Langflow
Open-source visual builder for LangChain-style AI agents and RAG pipelines.

RAGFlow
Open-source RAG engine with deep document parsing, hybrid search, and visual agent orchestration.

PromptHub
Git-style prompt management, testing, and deployment platform for teams running multiple LLMs in production.

PromptLayer
Lightweight prompt logging + management for OpenAI/Claude apps.

Patronus
Automated LLM evaluation for hallucinations, safety, and quality.

AgentMemory
Open-source persistent memory runtime for AI coding agents, with hybrid retrieval and zero external dependencies.

Epsilla
Agent-as-a-Service platform with managed RAG and a no-code builder for vertical enterprise AI.

MMagic
OpenMMLab's research-grade toolbox for image and video generation, restoration, and editing.

Cognee
Open-source graph-memory layer that gives AI agents persistent, queryable context across sessions.

LLaMA Factory
Open-source, no-code WebUI for fine-tuning 100+ open LLMs with LoRA, QLoRA, DPO, and PPO.

LLM by Datasette
A CLI and Python library for running prompts against any LLM provider and logging everything to SQLite.

Prompt Foundry
Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.

Puzzlet AI
Git-native prompt management and observability platform for teams shipping LLM applications.

Tableau
Salesforce-owned BI platform that bolted generative AI onto enterprise dashboards via Tableau Pulse and Tableau Agent.