cd /news/ai-agents/iris-vs-langfuse-vs-phoenix-vs-promp… · home › topics › ai-agents › article
[ARTICLE · art-143505] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Iris vs Langfuse vs Phoenix vs Promptfoo: where each wins, where each loses

A comparison of four AI agent observability and evaluation tools — Langfuse, Arize Phoenix, Promptfoo, and Iris — finds they solve the same problem in fundamentally different ways: Langfuse and Phoenix are observability platforms that added evaluation, Promptfoo is a CLI test runner, and Iris is an MCP server that evaluates agent traces with deterministic rules and publishes each rule's precision and recall. The writeup notes that the deciding factors for teams are integration method, where evaluation runs, cost, and behavior when agents use tools. Langfuse is now part of ClickHouse, Phoenix's parent Arize was acquired by Dynatrace, and Promptfoo is now part of OpenAI.

by read6 min views2 publishedOct 1, 2026

Four tools, four different answers to the same question: how do you know what an AI agent did, and whether it was any good? Langfuse is an open-source AI engineering platform, part of ClickHouse since January 2026 (announcement). Arize Phoenix is the open-source half of Arize, whose acquisition by Dynatrace was announced in August 2026 (announcement). Promptfoo is an open-source CLI for evaluating and red-teaming LLM apps, now part of OpenAI (repository). Iris is an MCP server that evaluates agent traces with deterministic rules and publishes each rule's precision and recall.

They are not four flavours of one thing. Two are observability platforms that added evaluation, one is a test runner, and one is an evaluation server that speaks the agent's own protocol. The differences that matter to a team choosing between them are the boring ones: how it gets into the stack, where the evaluation runs, what it costs, and what happens when the agent uses tools.

Every vendor statement below links the vendor's own page, read on 2026-10-01; the compare pages carry the same cells with the sentence each was read from and whether that sentence was found on the page.

Langfuse is an SDK: its Python and JS SDKs wrap your functions with an observe decorator or wrapper, which "is an easy way to automatically capture inputs, outputs, timings, and errors of a wrapped function" (Langfuse SDK docs). Beyond the SDKs it lists 100+ library and framework integrations and OpenTelemetry (Langfuse docs).

Phoenix is OpenTelemetry: you instrument the app with the OpenTelemetry SDK plus OpenInference auto-instrumentation, and export spans to the Phoenix collector (Phoenix tracing docs). If your framework is already among its tracing integrations, that is a few lines (Phoenix integrations).

Promptfoo is a test runner: declarative test cases run from the CLI or as a library, locally or as a CI step, calling the model providers directly (Promptfoo intro). Output assertions need nothing inside the app; its trajectory assertions read OpenTelemetry spans the app sends to Promptfoo (tracing).

Iris is one block in the MCP client's config. The agent connects, discovers Iris's tools, and logs traces or asks for verdicts through them; frameworks that are not MCP clients send OpenTelemetry traces to Iris's OTLP door instead (clients).

Langfuse evaluates with LLM-as-a-judge, Jev as a judge (a decision model that "does not sample text, so the same state and question return the same verdict"), human annotation, and custom scores through the API and SDK (evaluation overview, Jev as a judge); when "you want Langfuse to run deterministic Python or TypeScript logic for you, use code evaluators" (scores via API/SDK). LLM-as-a-judge calls are model calls you pay for.

Phoenix runs evaluators on the server: "LLM-as-a-judge evaluators backed by Phoenix-managed prompts", and code evaluators whose local backends ship with Phoenix, so they run on a self-hosted deployment (server evals, code evaluators). The managed Arize AX adds agent-as-a-judge (Arize AX docs).

Promptfoo goes furthest into the agent's trajectory: its assertion library includes trajectory:* and tool-call assertions alongside model-graded rubrics and custom JavaScript or Python (assertions reference). The assertions run on your machine against what the run produced.

Iris runs its built-in rules in-process, on the trace, with no model call, and publishes every built-in rule's precision and recall on a labelled corpus at iris-eval.com/proof, regenerated from the code at each release. Judge templates exist for the cases a rule cannot decide, on a key you supply.

Prices are the vendors' own, read on 2026-10-01; the compare pages carry the sentence each was read from.

Langfuse self-hosts as web and worker containers backed by PostgreSQL, ClickHouse, Redis or Valkey, and S3 or blob storage (self-hosting). It is a real deployment.

Phoenix is pip install arize-phoenix and phoenix serve (terminal); "by default Phoenix starts with a file-based SQLite database in a temporary folder", with PostgreSQL as the other database (configuration), and Docker and Kubernetes among its deployment options.

Promptfoo runs locally; a Docker image hosts a results server, and the vendor's own page says self-hosting "is not recommended for production use cases" (self-hosting).

Iris is one process and one SQLite file, or the Docker image with a health check (README). Langfuse offers a hosted MCP server that can query observations, metrics and datasets and create scores (changelog), and create evaluators and evaluation rules (changelog) — a way for a coding agent to drive Langfuse.

Phoenix builds a remote MCP server into Phoenix 19 and later. "The operation catalog is generated from the Phoenix REST API", so an agent can work with projects, traces, datasets, experiments, prompts and annotations (remote MCP), including SQL over traces (Arize blog) — again, a way for a coding agent to drive Phoenix.

Promptfoo goes the other way: its mcp provider "calls Model Context Protocol (MCP) tools directly, so you can test or red team the server itself" (MCP provider) — it tests MCP servers.

Iris is an MCP server. The agent discovers it on connect and logs and evaluates through its tools; Iris grades what the agent did with its tools, not whether a server honours its contract — a server test harness like Promptfoo's answers that question, and Iris runs beside it (capabilities).

ee folders" ( Langfuse wins when you need prompt management with versions and labels (prompt management), broad framework coverage, and enterprise compliance on paper today (SOC 2, ISO 27001, HIPAA per its security page). It loses on weight: self-hosting means two application containers and four datastores (self-hosting).

Phoenix wins when your stack is already OpenTelemetry and you want an open-source tracing and evaluation UI that starts with one pip install (terminal). It loses for a team that needs an OSI-approved license: Phoenix is ELv2 rather than MIT (license), and agent-as-a-judge lives in the managed AX tier (Arize AX docs).

Promptfoo wins for pre-deployment testing: declarative cases, deterministic and trajectory assertions, red-teaming plugins, all from the CLI in CI (intro). It loses as an open-source production observer: the open-source tool tests before deployment, runtime protection is a separate commercial product (Guardrails), and the vendor does not recommend self-hosting its results server for production (self-hosting).

Iris wins when the agent speaks MCP and you want deterministic, local evaluation whose accuracy is published before you rely on it (proof) — one config block, one process, one file. It loses when you need prompt management, a compliance certificate today, or a hundred framework integrations; those are not what it is, and the compare pages say so in muted cells rather than pretending otherwise (compare).

Every vendor statement above was read from the linked page on 2026-10-01. The compare pages carry the same cells with the sentence each was read from, the date, and whether a plain download of the page still carries that sentence; the file behind each page is in the repository under website/src/lib/compare/. If a vendor's page has changed, the cell is wrong, and the fix is a pull request.

── more in #ai-agents 4 stories · sorted by recency
── more on @langfuse 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/iris-vs-langfuse-vs-…] indexed:0 read:6min 2026-10-01 · —