Skip to main content
📖 The AI Tool Bible

AI eval dataset builder

Editorial picks for "llm eval datasets".

48 tools

All Evaluation
MathEval preview image
MathEval logo

MathEval

Evaluation · GPT-4 grader / DeepSeek-LLM-7B verifier
7.3

Holistic benchmark suite for evaluating mathematical reasoning in large language models.

Free· Free; open-source benchmark with leaderboard submissions via matheval.aillm-math-benchmarkingmodel-leaderboards
LLMEval preview image
LLMEval logo

LLMEval

Evaluation · Multi-model
7.2

Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.

Free· Free; open-source academic benchmarksllm-benchmarkingacademic-evaluation
InfiBench preview image
InfiBench logo

InfiBench

Evaluation
6.9

Stack Overflow-derived benchmark for evaluating code LLMs on real-world programming questions.

Free· Free and open source (CC BY-SA 4.0)code-llm-evalmodel-benchmarking
LangSmith preview image
LangSmith logo

LangSmith

Evaluation · Platform (any LLM)
8.7

LangChain's eval + observability platform.

Freemium· Developer: $0 · Plus: $39 · Enterprise: Custom pricingLLM tracingevals
Great Expectations preview image
Great Expectations logo

Great Expectations

Evaluation
8.2

Open-source data quality framework for validating the datasets that feed your ML and analytics pipelines.

Freemium· Developer: Free · Team: Custom · Enterprise: Contact Salesdata-validationpipeline-testing
LiveBench preview image
LiveBench logo

LiveBench

Evaluation · Multi-model
8.2

Contamination-free LLM benchmark that refreshes its questions monthly to keep frontier models honest.

Free· Free and open source; self-hosted evaluation runnerllm-benchmarkingmodel-selection
Athina AI preview image
Athina AI logo

Athina AI

Evaluation · Multi-model
8.1

Collaborative LLM evaluation and observability platform for teams shipping AI features to production.

Freemium· Starter free (10k logs/mo); Pro & Enterprise customllm-evaluationprompt-management
HoneyHive preview image
HoneyHive logo

HoneyHive

Evaluation · Multi-model
8.1

OpenTelemetry-native observability and evaluation platform for LLM agents in production.

Freemium· Free tier available; paid/enterprise tiers via salesagent-observabilityllm-evaluation
Hugging Face AutoTrain preview image
Hugging Face AutoTrain logo

Hugging Face AutoTrain

Fine-tuning · Multi-model (Hugging Face Hub)
8.1

No-code fine-tuning and training pipeline that spins up state-of-the-art models on the Hugging Face Hub.

Paid· Per-minute billing based on hardware tier; self-hosted OSS version is freellm-fine-tuningtext-classification
CAMEL-AI preview image
CAMEL-AI logo

CAMEL-AI

Agents · Multi-model
8.0

Open-source Python framework for building multi-agent systems and synthetic data pipelines.

Free· Free, open-source; pay for the underlying LLM API callsmulti-agent-systemssynthetic-data-generation
Valohai preview image
Valohai logo

Valohai

Agents · Multi-model
7.3

MLOps platform for versioned pipelines, distributed training, and LLM evaluation across any cloud.

Enterprise· Free trial; contact sales for pricingmlopsllm-evaluation
Inspect AI preview image
Inspect AI logo

Inspect AI

Evaluation · Multi-model
7.2

Open-source LLM evaluation framework from the UK AI Security Institute with 200+ built-in benchmarks.

Free· Free and open source (MIT-style license); you pay only for underlying model API usage.llm-benchmarkingagent-evaluation
Kiln AI preview image
Kiln AI logo

Kiln AI

Evaluation · Multi-model
7.2

Open-source workbench for building, evaluating, and fine-tuning AI agents across 190+ models.

Freemium· Free Individual tier; Team (request access); Enterprise (custom)llm-evaluationfine-tuning
Maxim AI preview image
Maxim AI logo

Maxim AI

Evaluation · Multi-model
7.1

End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.

Freemium· Developer: Free · Professional: $29 /seat /month · Business: $49 /seat /month · Enterprise: Customagent-evaluationllm-observability
Forefront preview image
Forefront logo

Forefront

Fine-tuning · Multi-model (Mistral-7B, Mixtral, Phi-2)
7.0

Fine-tune and serve open-source LLMs on your own data without managing GPUs.

Paid· Basic: $20 · Pro: $50 · Enterprise: Contact salesfine-tuningopen-source-llms
OlympicArena preview image
OlympicArena logo

OlympicArena

Evaluation · GPT-4o, Claude-3.5-Sonnet, Doubao-Pro-32k, DeepSeek-Coder-V2, Qwen2-72B-Instruct
7.0

Olympiad-level multi-discipline benchmark for stress-testing reasoning in LLMs and multimodal models.

Free· Free, open-source research benchmarkllm-evaluationmultimodal-eval
Phoenix preview image
Phoenix logo

Phoenix

Evaluation · Multi-model
7.0

Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-tracingagent-debugging
CompassRank preview image
CompassRank logo

CompassRank

Evaluation · Multi-model
6.9

Public leaderboard from the OpenCompass project ranking open and closed LLMs across 100+ benchmarks.

Free· Free leaderboard; OpenCompass toolkit is Apache 2.0 open sourcellm-benchmarkingmodel-selection
MixEval preview image
MixEval logo

MixEval

Evaluation · GPT-3.5-Turbo-0125, GPT-4o-2024-05-13, Claude 3.5 Sonnet, MixEval, MixEval-Hard
6.9

Dynamic LLM benchmark that mixes web queries with existing datasets to mirror Chatbot Arena rankings at a fraction of the cost.

Free· Free and open sourcellm-benchmarkingmodel-ranking
ClickHouse preview image
ClickHouse logo

ClickHouse

RAG

The open-source columnar database powering real-time analytics — and, increasingly, LLM observability and RAG backends.

Freemium· Open-source self-managed: free. ClickHouse Cloud: from $50/month (usage-based on compute + storage, AWS/GCP/Azure). Enterprise tier available with dedicated support and BYOC options.LLM trace and cost analyticsRAG retrieval with hybrid vector + metadata filters
Language Model Builder preview image
Language Model Builder logo

Language Model Builder

Fine-tuning · In-house small transformer models trained by the user; exports to safetensors

Learn how LLMs work by building one on your Mac

Free· Free macOS download. No account, subscription, or fees. Mac App Store version listed as coming soon.Learn transformer internals hands-onPre-train a small language model locally
LangWatch preview image
LangWatch logo

LangWatch

Evaluation · Model-agnostic; supports OpenAI, Anthropic, AWS Bedrock, Azure OpenAI, Vertex AI, and any OpenTelemetry-instrumented LLM

Simulation-based testing, evaluation, and observability for LLM agents

Freemium· Developer: €0 · Growth: €29/ core-seat / month · Enterprise: CustomLLM agent regression testing in CIRAG answer-quality evaluation
Braintrust preview image
Braintrust logo

Braintrust

Featured
Evaluation · Platform (any LLM)
8.9

Eval, monitor, and improve AI products end-to-end.

Freemium· Starter: $0 · Pro: $249 · Enterprise: Custom pricingevalsmonitoring
IBM watsonx preview image
IBM watsonx logo

IBM watsonx

Agents · IBM Granite (3.x, Code, Time Series), Meta Llama 3.x, Mistral, plus other curated open models
8.6

Enterprise AI platform for building, deploying, and governing models and agents

Enterprise· watsonx.ai has a free tier on IBM Cloud with limited tokens; paid usage is metered per 1M tokens by model family (Granite, Llama, Mistral, etc.). watsonx.governance and watsonx.data are quoted per environment. Enterprise deals via IBM sales; on-prem/Cloud Pak for Data is separately licensed.Enterprise RAG chatbot over private documentsCustomer service agents with guardrails
Glean preview image
Glean logo

Glean

Agents · Model-agnostic: routes across 35+ LLMs including GPT-4o, Claude 3.5/4 Sonnet, Gemini 1.5/2, Llama 3, Mistral, plus Glean in-house models
8.5

Work AI platform that unifies enterprise knowledge, search, and agents

Enterprise· Enterprise pricing only; commonly reported in the $40-50/user/month range with a floor typically starting around 100 seats. No public self-serve tier. Contact sales for a quote.Enterprise search across Slack, Drive, Confluence and JiraCompany-wide AI assistant with citations
Yi (01.AI) preview image
Yi (01.AI) logo

Yi (01.AI)

Agents · Yi-Lightning (MoE), Yi-Large, Yi-1.5 (6B/9B/34B), Yi-VL, Yi-Coder — in-house 01.AI foundation models
8.4

Foundation models from 01.AI — open-weight Yi family plus frontier Yi-Lightning and Yi-Large

Freemium· Open-source Yi models free under permissive license; hosted API via platform.lingyiwanwu.com with pay-per-token pricing (Yi-Lightning positioned as a low-cost frontier tier; Yi-Large priced higher; exact per-token rates on the platform dashboard). Enterprise custom-training and consulting on quote.Self-hosted coding assistantBilingual English-Chinese chatbot
ClearML preview image
ClearML logo

ClearML

Agents · Model-agnostic
8.3

End-to-end MLOps and GenAI platform with open-source experiment tracking and enterprise GPU orchestration.

Freemium· Community: $0 · Pro: $15 Per User/Month + Usage · Scale: Custom Quote · Enterprise: Request a Quoteexperiment-trackinggpu-orchestration
Dataiku preview image
Dataiku logo

Dataiku

Agents · Multi-model (LLM Mesh: OpenAI, Anthropic, Bedrock, Vertex, OSS)
8.3

Enterprise AI platform unifying data, ML, LLMs, and agents under one governed workflow.

Enterprise· Basic: $10 · Pro: $20 · Enterprise: Contact salesenterprise-aiagent-orchestration
Helicone preview image
Helicone logo

Helicone

Evaluation · Platform (any LLM)
8.3

Open-source LLM observability — one-line proxy install.

Freemium· Free 100k req/mo; Pro from $25/moobservabilitycost tracking
Humanloop preview image
Humanloop logo

Humanloop

Evaluation · Platform (any LLM)
8.2

Prompt management + evals for collaborative AI teams.

Paid· From $200/mo teamprompt managementteam collab
LanceDB preview image
LanceDB logo

LanceDB

RAG
8.2

Open-source multimodal lakehouse and vector database built for AI training and retrieval at petabyte scale.

Freemium· Open-source free; LanceDB Cloud and Enterprise via contact salesvector-searchrag
OpenPipe preview image
OpenPipe logo

OpenPipe

Fine-tuning · Llama, Mistral, Qwen and other open-weight base models
8.2

Fine-tuning and reinforcement learning platform for turning expensive prompts into cheap, fast, task-specific models.

Freemium· Free tier available; usage-based pricing for training and hosted inference; enterprise plans on requestllm-cost-reductionfine-tuning
Writer preview image
Writer logo

Writer

Writing · Palmyra (in-house)
8.2

Enterprise generative AI platform built around in-house Palmyra LLMs for regulated, brand-consistent content.

Enterprise· Starter: Free · ENTERPRISE: Contact salesenterprise-writingbrand-voice
Berkeley Function-Calling Leaderboard preview image
Berkeley Function-Calling Leaderboard logo

Berkeley Function-Calling Leaderboard

Evaluation · Multi-model
8.1

Open benchmark from UC Berkeley that ranks LLMs on real-world tool-use and function-calling accuracy.

Free· Free and open source; you pay only for inference when reproducing runs.function-calling evaltool-use benchmarking
Langflow preview image
Langflow logo

Langflow

Agents · Multi-model
8.1

Open-source visual builder for LangChain-style AI agents and RAG pipelines.

Freemium· Open-source free; hosted free tier + paid enterprise via DataStaxagent-prototypingrag-pipelines
RAGFlow preview image
RAGFlow logo

RAGFlow

RAG · Multi-model
8.1

Open-source RAG engine with deep document parsing, hybrid search, and visual agent orchestration.

Freemium· Free tier; Starter $29/mo; Pro $129/mo; Enterprise customdocument-qaenterprise-search
PromptHub preview image
PromptHub logo

PromptHub

Writing · Multi-model (OpenAI, Anthropic, Google, Meta, Mistral, Bedrock, Azure)
8.0

Git-style prompt management, testing, and deployment platform for teams running multiple LLMs in production.

Freemium· Free signup; paid team plans (contact sales / in-app)prompt-managementprompt-versioning
PromptLayer preview image
PromptLayer logo

PromptLayer

Evaluation · Platform (any LLM)
7.9

Lightweight prompt logging + management for OpenAI/Claude apps.

Freemium· Free: $0 · Pro: $49 · Team: $500 · Enterprise: Customprompt loggingversioning
Patronus preview image
Patronus logo

Patronus

Evaluation · Platform (any LLM)
7.8

Automated LLM evaluation for hallucinations, safety, and quality.

Paid· Individual: Free · Base: $25 · Enterprise: Contact us for Pricinghallucination detectionsafety
AgentMemory preview image
AgentMemory logo

AgentMemory

Agents · Multi-model (Claude, Gemini, MiniMax, OpenRouter)
7.3

Open-source persistent memory runtime for AI coding agents, with hybrid retrieval and zero external dependencies.

Free· Free, open-source (Apache 2.0); bring your own LLM API keyagent-memorycoding-agents
Epsilla preview image
Epsilla logo

Epsilla

RAG · Multi-model
7.3

Agent-as-a-Service platform with managed RAG and a no-code builder for vertical enterprise AI.

Freemium· Free Tier: $0/month · Starter Tier: $29/month · Professional Tier: $249/month · AI Concierge: $2,499/month · Enterprise Tier: Custom/monthenterprise-ragai-agents
MMagic preview image
MMagic logo

MMagic

Image Generation · Multi-model (Stable Diffusion, ControlNet, StyleGAN, GANs, diffusion)
7.3

OpenMMLab's research-grade toolbox for image and video generation, restoration, and editing.

Free· Free and open source (Apache 2.0)text-to-imagesuper-resolution
Cognee preview image
Cognee logo

Cognee

RAG · Multi-model (Claude, OpenAI, others)
7.2

Open-source graph-memory layer that gives AI agents persistent, queryable context across sessions.

Freemium· Hobby free (1M tokens/mo); Growth $5/workspace/mo + token usage; Enterprise customagent-memoryknowledge-graphs
LLaMA Factory preview image
LLaMA Factory logo

LLaMA Factory

Fine-tuning · Multi-model (LLaMA, Mistral, Qwen, Gemma, Phi, LLaVA, ChatGLM, Yi)
7.2

Open-source, no-code WebUI for fine-tuning 100+ open LLMs with LoRA, QLoRA, DPO, and PPO.

Free· Free, open-source (Apache-2.0); self-hostedlora-fine-tuningqlora
LLM by Datasette preview image
LLM by Datasette logo

LLM by Datasette

Coding · Multi-model
7.2

A CLI and Python library for running prompts against any LLM provider and logging everything to SQLite.

Free· Free and open source (Apache 2.0); pay underlying model providers separatelycli-promptingprompt-logging
Prompt Foundry preview image
Prompt Foundry logo

Prompt Foundry

Evaluation · OpenAI + Anthropic (multi-model)
7.1

Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.

Freemium· Free tier (10 prompts, 500 evals/mo); Pro $15/user/mo; Enterprise customprompt-managementmodel-comparison
Puzzlet AI preview image
Puzzlet AI logo

Puzzlet AI

Agents · Multi-model
7.1

Git-native prompt management and observability platform for teams shipping LLM applications.

Freemium· Basic: $20 · Pro: $50 · Enterprise: Contact salesprompt-managementllm-observability
Tableau preview image
Tableau logo

Tableau

Agents · Salesforce Einstein (multi-model, incl. OpenAI via Einstein Trust Layer)
7.1

Salesforce-owned BI platform that bolted generative AI onto enterprise dashboards via Tableau Pulse and Tableau Agent.

Paid· Tableau Standard: $ 15 USD/User/Month (Billed annually) · Tableau Enterprise: $ 35 USD/User/Month (Billed annually) · Tableau Next: $ 40 USD/User/Month (Billed annually)business-intelligencenatural-language-queries