A framework for running standardised benchmark suites against a model to score capability, safety, and regression.
eval
Eval Harness
Related terms
Tools that implement Eval Harness
LA
LangSmith
Evaluation · Platform (any LLM)
8.7
LangChain's eval + observability platform.
Freemium· Developer: $0 · Plus: $39 · Enterprise: Custom pricingLLM tracingevals
OE
OpenAI Evals
Evaluation · OpenAI GPT models (extensible)
8.1
OpenAI's open-source framework for benchmarking LLMs against a shared registry of evaluations.
Free· Free (MIT); you pay OpenAI API costs for eval runsllm-benchmarkingregression-testing
BR
Braintrust
FeaturedEvaluation · Platform (any LLM)
8.9
Eval, monitor, and improve AI products end-to-end.
Freemium· Starter: $0 · Pro: $249 · Enterprise: Custom pricingevalsmonitoring