{"slug": "iris-vs-langfuse-vs-phoenix-vs-promptfoo-where-each-wins-where-each-loses", "title": "Iris vs Langfuse vs Phoenix vs Promptfoo: where each wins, where each loses", "summary": "A comparison of four AI agent observability and evaluation tools — Langfuse, Arize Phoenix, Promptfoo, and Iris — finds they solve the same problem in fundamentally different ways: Langfuse and Phoenix are observability platforms that added evaluation, Promptfoo is a CLI test runner, and Iris is an MCP server that evaluates agent traces with deterministic rules and publishes each rule's precision and recall. The writeup notes that the deciding factors for teams are integration method, where evaluation runs, cost, and behavior when agents use tools. Langfuse is now part of ClickHouse, Phoenix's parent Arize was acquired by Dynatrace, and Promptfoo is now part of OpenAI.", "body_md": "Four tools, four different answers to the same question: how do you know what an AI agent did, and whether it was any good? [Langfuse](https://langfuse.com/docs) is an open-source AI engineering platform, part of ClickHouse since January 2026 ([announcement](https://langfuse.com/blog/joining-clickhouse)). [Arize Phoenix](https://github.com/Arize-ai/phoenix) is the open-source half of Arize, whose acquisition by Dynatrace was announced in August 2026 ([announcement](https://arize.com/blog/a-new-chapter-with-dynatrace/)). [Promptfoo](https://www.promptfoo.dev/docs/intro/) is an open-source CLI for evaluating and red-teaming LLM apps, now part of OpenAI ([repository](https://github.com/promptfoo/promptfoo)). [Iris](https://iris-eval.com) is an MCP server that evaluates agent traces with deterministic rules and publishes each rule's precision and recall.\n\nThey are not four flavours of one thing. Two are observability platforms that added evaluation, one is a test runner, and one is an evaluation server that speaks the agent's own protocol. The differences that matter to a team choosing between them are the boring ones: how it gets into the stack, where the evaluation runs, what it costs, and what happens when the agent uses tools.\n\nEvery vendor statement below links the vendor's own page, read on 2026-10-01; the [compare pages](https://iris-eval.com/compare) carry the same cells with the sentence each was read from and whether that sentence was found on the page.\n\n**Langfuse** is an SDK: its Python and JS SDKs wrap your functions with an `observe` decorator or wrapper, which \"is an easy way to automatically capture inputs, outputs, timings, and errors of a wrapped function\" ([Langfuse SDK docs](https://langfuse.com/docs/observability/sdk/instrumentation)). Beyond the SDKs it lists 100+ library and framework integrations and OpenTelemetry ([Langfuse docs](https://langfuse.com/docs)).\n\n**Phoenix** is OpenTelemetry: you instrument the app with the OpenTelemetry SDK plus OpenInference auto-instrumentation, and export spans to the Phoenix collector ([Phoenix tracing docs](https://arize.com/docs/phoenix/get-started/get-started-tracing)). If your framework is already among its tracing integrations, that is a few lines ([Phoenix integrations](https://arize.com/docs/phoenix/integrations)).\n\n**Promptfoo** is a test runner: declarative test cases run from the CLI or as a library, locally or as a CI step, calling the model providers directly ([Promptfoo intro](https://www.promptfoo.dev/docs/intro/)). Output assertions need nothing inside the app; its trajectory assertions read OpenTelemetry spans the app sends to Promptfoo ([tracing](https://www.promptfoo.dev/docs/tracing/)).\n\n**Iris** is one block in the MCP client's config. The agent connects, discovers Iris's tools, and logs traces or asks for verdicts through them; frameworks that are not MCP clients send OpenTelemetry traces to Iris's OTLP door instead ([clients](https://iris-eval.com/clients)).\n\n**Langfuse** evaluates with LLM-as-a-judge, Jev as a judge (a decision model that \"does not sample text, so the same state and question return the same verdict\"), human annotation, and custom scores through the API and SDK ([evaluation overview](https://langfuse.com/docs/evaluation/overview), [Jev as a judge](https://langfuse.com/docs/evaluation/evaluation-methods/jev-as-a-judge)); when \"you want Langfuse to run deterministic Python or TypeScript logic for you, use code evaluators\" ([scores via API/SDK](https://langfuse.com/docs/evaluation/evaluation-methods/scores-via-sdk)). LLM-as-a-judge calls are model calls you pay for.\n\n**Phoenix** runs evaluators on the server: \"LLM-as-a-judge evaluators backed by Phoenix-managed prompts\", and code evaluators whose local backends ship with Phoenix, so they run on a self-hosted deployment ([server evals](https://arize.com/docs/phoenix/evaluation/server-evals), [code evaluators](https://arize.com/docs/phoenix/evaluation/server-evals/code-evaluators)). The managed Arize AX adds agent-as-a-judge ([Arize AX docs](https://arize.com/docs/ax)).\n\n**Promptfoo** goes furthest into the agent's trajectory: its assertion library includes `trajectory:*` and tool-call assertions alongside model-graded rubrics and custom JavaScript or Python ([assertions reference](https://www.promptfoo.dev/docs/configuration/expected-outputs/)). The assertions run on your machine against what the run produced.\n\n**Iris** runs its built-in rules in-process, on the trace, with no model call, and publishes every built-in rule's precision and recall on a labelled corpus at [iris-eval.com/proof](https://iris-eval.com/proof), regenerated from the code at each release. Judge templates exist for the cases a rule cannot decide, on a key you supply.\n\nPrices are the vendors' own, read on 2026-10-01; the compare pages carry the sentence each was read from.\n\n**Langfuse** self-hosts as web and worker containers backed by PostgreSQL, ClickHouse, Redis or Valkey, and S3 or blob storage ([self-hosting](https://langfuse.com/self-hosting)). It is a real deployment.\n\n**Phoenix** is `pip install arize-phoenix` and `phoenix serve` ([terminal](https://arize.com/docs/phoenix/self-hosting/deployment-options/terminal)); \"by default Phoenix starts with a file-based SQLite database in a temporary folder\", with PostgreSQL as the other database ([configuration](https://arize.com/docs/phoenix/self-hosting/configuration)), and Docker and Kubernetes among its deployment options.\n\n**Promptfoo** runs locally; a Docker image hosts a results server, and the vendor's own page says self-hosting \"is not recommended for production use cases\" ([self-hosting](https://www.promptfoo.dev/docs/usage/self-hosting/)).\n\n**Iris** is one process and one SQLite file, or the Docker image with a health check ([README](https://github.com/iris-eval/mcp-server)).\n\n**Langfuse** offers a hosted MCP server that can query observations, metrics and datasets and create scores ([changelog](https://langfuse.com/changelog/2026-05-29-mcp-update)), and create evaluators and evaluation rules ([changelog](https://langfuse.com/changelog/2026-06-10-evaluators-via-mcp)) — a way for a coding agent to drive Langfuse.\n\n**Phoenix** builds a remote MCP server into Phoenix 19 and later. \"The operation catalog is generated from the Phoenix REST API\", so an agent can work with projects, traces, datasets, experiments, prompts and annotations ([remote MCP](https://arize.com/docs/phoenix/integrations/remote-mcp)), including SQL over traces ([Arize blog](https://arize.com/blog/phoenix-mcp-sql-code-mode/)) — again, a way for a coding agent to drive Phoenix.\n\n**Promptfoo** goes the other way: its `mcp` provider \"calls Model Context Protocol (MCP) tools directly, so you can test or red team the server itself\" ([MCP provider](https://www.promptfoo.dev/docs/providers/mcp/)) — it tests MCP servers.\n\n**Iris** is an MCP server. The agent discovers it on connect and logs and evaluates through its tools; Iris grades what the agent did with its tools, not whether a server honours its contract — a server test harness like Promptfoo's answers that question, and Iris runs beside it ([capabilities](https://iris-eval.com/capabilities)).\n\n`ee` folders\" (\n**Langfuse wins** when you need prompt management with versions and labels ([prompt management](https://langfuse.com/docs/prompt-management/overview)), broad framework coverage, and enterprise compliance on paper today (SOC 2, ISO 27001, HIPAA per its [security page](https://langfuse.com/security)). **It loses** on weight: self-hosting means two application containers and four datastores ([self-hosting](https://langfuse.com/self-hosting)).\n\n**Phoenix wins** when your stack is already OpenTelemetry and you want an open-source tracing and evaluation UI that starts with one `pip install` ([terminal](https://arize.com/docs/phoenix/self-hosting/deployment-options/terminal)). **It loses** for a team that needs an OSI-approved license: Phoenix is ELv2 rather than MIT ([license](https://github.com/Arize-ai/phoenix/blob/main/LICENSE)), and agent-as-a-judge lives in the managed AX tier ([Arize AX docs](https://arize.com/docs/ax)).\n\n**Promptfoo wins** for pre-deployment testing: declarative cases, deterministic and trajectory assertions, red-teaming plugins, all from the CLI in CI ([intro](https://www.promptfoo.dev/docs/intro/)). **It loses** as an open-source production observer: the open-source tool tests before deployment, runtime protection is a separate commercial product ([Guardrails](https://www.promptfoo.dev/guardrails/)), and the vendor does not recommend self-hosting its results server for production ([self-hosting](https://www.promptfoo.dev/docs/usage/self-hosting/)).\n\n**Iris wins** when the agent speaks MCP and you want deterministic, local evaluation whose accuracy is published before you rely on it ([proof](https://iris-eval.com/proof)) — one config block, one process, one file. **It loses** when you need prompt management, a compliance certificate today, or a hundred framework integrations; those are not what it is, and the compare pages say so in muted cells rather than pretending otherwise ([compare](https://iris-eval.com/compare)).\n\nEvery vendor statement above was read from the linked page on 2026-10-01. The [compare pages](https://iris-eval.com/compare) carry the same cells with the sentence each was read from, the date, and whether a plain download of the page still carries that sentence; the file behind each page is in the [repository](https://github.com/iris-eval/mcp-server) under `website/src/lib/compare/`. If a vendor's page has changed, the cell is wrong, and the fix is a pull request.", "url": "https://wpnews.pro/news/iris-vs-langfuse-vs-phoenix-vs-promptfoo-where-each-wins-where-each-loses", "canonical_source": "https://dev.to/irparent/iris-vs-langfuse-vs-phoenix-vs-promptfoo-where-each-wins-where-each-loses-3kca", "published_at": "2026-10-01 21:55:28+00:00", "updated_at": "2026-10-01 22:14:37.947439+00:00", "lang": "en", "topics": ["ai-agents", "mlops", "developer-tools", "agent-protocols", "ai-tools"], "entities": ["Langfuse", "Arize Phoenix", "Promptfoo", "Iris", "ClickHouse", "Dynatrace", "OpenAI", "Arize"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/iris-vs-langfuse-vs-phoenix-vs-promptfoo-where-each-wins-where-each-loses", "markdown": "https://wpnews.pro/news/iris-vs-langfuse-vs-phoenix-vs-promptfoo-where-each-wins-where-each-loses.md", "text": "https://wpnews.pro/news/iris-vs-langfuse-vs-phoenix-vs-promptfoo-where-each-wins-where-each-loses.txt", "jsonld": "https://wpnews.pro/news/iris-vs-langfuse-vs-phoenix-vs-promptfoo-where-each-wins-where-each-loses.jsonld"}}