Tools

LLM observability & evaluation

15 tools compared: how each is priced, where it runs, and what to consider instead.

Tool Pricing model Free tier Open source
Arize Phoenix Open-source, local-first tracing and evaluation library for LLM and agent applications, with an optional path to Arize's commercial platform. Open source + paid Yes Yes
Braintrust Evaluation-first platform for testing, scoring, and comparing LLM application changes, with production tracing built in. Usage-based Yes No
DeepEval Open-source Python evaluation framework with 30+ research-backed metrics, with an optional paid platform (Confident AI) for CI/CD and hosted reporting. Free tier + paid Yes Yes
Galileo LLM observability and evaluation platform with real-time guardrails, aimed at monitoring agents and generative AI in production. Usage-based Yes No
Helicone Open-source LLM observability platform that logs and analyzes requests via a proxy or async SDK, with a free-tier hosted cloud. Free tier + paid Yes Yes
Humanloop Prompt management, evaluation, and observability platform aimed at product teams collaborating on LLM features with domain experts. Quote only Yes No
Langfuse Open-source LLM tracing and evaluation platform, available self-hosted for free or as a metered hosted cloud service. Free tier + paid Yes Yes
LangSmith Tracing, evaluation, and deployment platform for LLM applications, built by the team behind LangChain and LangGraph. Usage-based Yes No
Lunary GDPR-focused LLM observability platform for tracing, prompt management, and human review, hosted in Europe or self-hosted. Free tier + paid Yes Yes
OpenLLMetry Apache-2.0 open-source instrumentation library extending OpenTelemetry to LLM calls, usable with any OpenTelemetry-compatible backend. Free tier + paid Yes Yes
Patronus AI Evaluation-focused platform offering pre-built and custom evaluator models to score LLM outputs for accuracy, safety, and quality. Usage-based Yes No
Portkey AI gateway that routes and load-balances LLM calls across providers while logging traces, prompts, and feedback for observability. Free tier + paid Yes Yes
PromptLayer Prompt management and request-logging platform with versioning, playgrounds, and evaluation for LLM applications. Free tier + paid Yes No
Ragas Open-source Python framework of reference-free metrics purpose-built for evaluating retrieval-augmented generation (RAG) pipelines. Open source + paid Yes Yes
W&B Weave LLM tracing and evaluation toolkit from Weights & Biases, sharing billing and infrastructure with its ML experiment-tracking platform. Free tier + paid Yes No

In the index now