Arthur alternatives
6 evaluation tools in the same lane as Arthur, ranked by editorial score.
OP
Opik
Evaluation · Multi-model
7.3
Open-source LLM observability and evaluation platform for debugging and monitoring AI agents in production.
Freemium· Free open-source self-host; free Cloud tier (no card); Enterprise contact salesllm-tracingagent-evaluation
GI
Giskard
Evaluation · Multi-model
8.2
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.
Freemium· Open-source free tier; Giskard Hub enterprise pricing on requestllm-red-teamingagent-security-testing
TR
TruLens
Evaluation · Multi-model (LLM-as-judge)
8.1
Open-source evaluation and tracing framework for LLM apps and agents, built on OpenTelemetry.
Free· Free, open source (Apache-licensed Python package)llm-evaluationrag-evaluation
PH
Phoenix
Evaluation · Multi-model
7.0
Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.
Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-tracingagent-debugging
AG
Agenta
Evaluation · Multi-model
6.9
Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.
Freemium· Hobby: $0 forever · Pro: $29 /month · Business: $299 /month · Enterprise: Custom · Open source: Free foreverprompt-engineeringllm-evaluation
AA
Arize AI
Evaluation · Multi-model
8.2
Enterprise observability and evaluation platform for LLM agents and generative AI applications.
Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-observabilityagent-evaluation