A framework for running standardised benchmark suites against a model to score capability, safety, and regression.
eval
Eval Harness
Related terms
Tools that implement Eval Harness
LangSmith
Evaluation · Platform (any LLM)
8.7
LangChain's eval + observability platform.
Freemium· Free starter; Plus $39/mo per seatLLM tracingevals
Braintrust
FeaturedEvaluation · Platform (any LLM)
8.9
Eval, monitor, and improve AI products end-to-end.
Freemium· Free up to 1k events/day; team from $249/moevalsmonitoring
OpenAI Evals
Evaluation · OpenAI GPT models (extensible)
8.1
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.
Free· Free (MIT); you pay OpenAI API costs for eval runsllm-benchmarkingregression-testing