{"slug": "what-agent-traces-can-tell-you-without-an-llm-judge", "title": "What agent traces can tell you without an LLM judge", "summary": "Arize Phoenix contributor Ashwin Govind Ugale released tracelint, an open-source linter that runs deterministic structural checks on OpenInference traces and returns a CI-usable exit code, avoiding LLM judges for failures provable from the trace alone. In a LangGraph release agent test using gpt-4o-mini, the agent deployed build 4.12.0 to production despite a Jenkins result of UNSTABLE with 3 of 215 tests failed; with \"deploy only if its pipeline passes\" in the system prompt it never deployed in ten runs, and with that sentence removed it deployed in all ten. tracelint flags deterministic defects such as tool arguments violating their JSON Schema, calls to nonexistent tools, and failed results fed into side-effecting actions, while surfacing repeated no-progress calls as candidates for review.", "body_md": "This is a guest post by Ashwin Govind Ugale, a contributor to the [Arize Phoenix](https://arize.com/phoenix/) open source AI observability platform.\n\n### TL;DR\n\nSome agent failures can be proven from the trace alone, such as a tool call that violates its schema or a failed result reused in a side effect. Other patterns, like repeated calls with no progress, can be surfaced as candidates without claiming they’re definitely bugs.\n\nA useful rule for deciding which is which: if the verdict depends on what the agent should have meant, it’s an eval. If it depends only on what the trace records, it can be a deterministic check.\n\ntracelint is a small open-source linter that runs those checks on [OpenInference](https://arize.com/glossary/openinference/) traces, including the ones [Phoenix](https://arize.com/phoenix/) collects, and returns an exit code you can use in CI. [tracelint](https://github.com/AshwinUgale/tracelint)\n\n## A run can look successful and still have a broken trajectory\n\nIn a small experiment, I asked a [LangGraph](https://arize.com/docs/phoenix/integrations/python/langgraph) release agent (gpt-4o-mini) to ship release 4.12.0. Its stubbed pipeline tool returns HTTP 200 with a Jenkins result of `UNSTABLE` (3 of 215 tests failed). The agent deploys that build to production anyway and replies:\n\nThe release version 4.12.0 has been shipped to production. However, please note that the Jenkins build result was UNSTABLE, with 3 test failures out of 215 tests.\n\nEvery span reports success. The agent noticed the failure, and shipped anyway.\n\nWith “deploy only if its pipeline passes” in the system prompt, it never deployed in ten runs. With that sentence removed, it deployed in all ten. That’s the kind of regression CI exists to catch.\n\nDo you need an [LLM judge](https://arize.com/guides/llm-as-a-judge/) here? No. The pipeline failed by your team’s own definition, the same build ID shows up in the deploy call, and deploying has side effects. The verdict is already in the trace, plus two facts about your tools. I built tracelint to turn failures like this into repeatable checks you can run locally or in CI.\n\nThink of these as structural regression tests for agent behavior: they check claims the trace can prove and leave semantic judgement to [agent evals](https://arize.com/ai-agents/agent-evaluation/).\n\n## What agent failures can you prove from a trace?\n\nA useful rule of thumb is simple: if the verdict requires knowing what the agent was supposed to do, or what it meant, it probably isn’t a lint rule. If the verdict depends only on what the [trace](https://arize.com/docs/phoenix/tracing/concepts-tracing/what-are-traces) records, plus facts you declare about your tools, it can be checked deterministically.\n\nA few examples along that line:\n\n- Arguments that violate the [tool’s JSON Schema](https://arize.com/blog/how-to-evaluate-tool-calling-agents/) : deterministic.\n- A call to a tool that doesn’t exist: deterministic.\n- A result declared as failed that’s fed into a side-effecting action: deterministic.\n- The same call repeatedly with no apparent progress: not necessarily a bug, since it might be polling or a legitimate retry. This gets surfaced for review.\n- The agent picked the wrong strategy: a question for an eval.\n\ntracelint includes several other structural checks, but they all follow the same principle. It reports a defect only when the trace provides enough evidence to prove it.\n\n## How to run tracelint on an Arize Phoenix trace\n\nIf you’re already [tracing your agents with Arize Phoenix](https://arize.com/docs/phoenix/tracing/how-to-tracing/setup-tracing), tracelint can read the OpenInference spans you collect without any changes to the agent:\n\n``` python\nfrom phoenix.client import Client\nfrom tracelint import lint_otel_traces, render_report\n\nspans = Client().spans.get_spans_dataframe(project_name=\"release-agent\")\nfor report in lint_otel_traces(spans.to_dict(\"records\")):  # one report per run\n    print(render_report(report))\n```\n\nRun tracelint on this trace as-is and it won’t fail it, and that’s correct: nothing in the trace says `UNSTABLE` means failure. That’s your CI’s convention. Two pieces of domain knowledge still have to come from you:\n\n- `UNSTABLE` (like`FAILURE` and`ABORTED` ) means the pipeline failed\n- `deploy` takes a real action (it changes production), so acting on a failed result there is a defect, not just something to review\n\nDeclare those facts in a small `tools.json`:\n\n```\n{\n  \"tools\": {\n    \"run_release_pipeline\": {\n      \"metadata\": {\n        \"failure_when\": {\"pointer\": \"/result\", \"in\": [\"FAILURE\", \"UNSTABLE\", \"ABORTED\"]}\n      }\n    },\n    \"deploy\": {\n      \"metadata\": {\"side_effecting\": true}\n    }\n  }\n}\n```\n\nRun the same trace again with those tool semantics declared:\n\n``` bash\n$ tracelint check spans.json --format openinference --tools tools.json\n689451f502ecfcb8bc0063e2615ed8c5: 2 finding(s), exit 2\n[hard_event] R2a tool_error_event  (step 3)\n'run_release_pipeline' returned a declared failure (/result='UNSTABLE')\n[hard_defect] R2b error_mishandled  (step 3,4)\nvalue(s) from the errored 'run_release_pipeline' result (b-4120-7f3a) reused\nas arguments to 'deploy' (a side-effecting action, no fallback)\n…\n```\n\nIn plain terms:\n\n1. The pipeline failed. Its result was `UNSTABLE` , which your`tools.json` says counts as a failure.\n2. Its build was deployed anyway. The build ID `b-4120-7f3a` from that failed run was passed to`deploy` , which changes production. That’s a provable defect, so tracelint exits with code 2 and fails the CI job.\n\nYou don’t have to write `tools.json` from scratch. `tracelint init spans.json --format openinference -o tools.json` drafts one from the trace, including tool schemas where the instrumentation records them.\n\n## How tracelint classifies findings: hard defects, candidates, and not checked\n\nEvery finding lands in one of three buckets, and the distinction matters more than any individual rule.\n\n- **Hard defects** are proven by the trace. In CI,`tracelint check` returns exit code 2 and fails the job.\n- **Candidates** are possible problems, such as repeated calls that might be a loop or might be a legitimate retry. They’re shown for review but never fail the build. That’s deliberate: one noisy heuristic is enough to make engineers disable a linter.\n- **Not checked** means a rule lacked the evidence it needed, and tracelint says so instead of counting it as a pass.\n\nFor example, the `tools.json` above doesn’t say what a failed deploy looks like, so the same run’s output includes:\n\n```\nsuppressed (3) — not checked, not a clean pass:\nR2a tool_error_event: side-effecting tool 'deploy' returned an unclassifiable\nresult and declares no failure_when predicate — cannot verify it did not fail\n…\n```\n\ntracelint won’t claim the deploy itself succeeded, because nothing tells it what failure would look like.\n\nIf you don’t save traces in your tests yet, tracelint has a capture helper and a pytest fixture that record a run through your framework’s OpenInference instrumentor, plus a GitHub Action for CI. The [repo](https://github.com/AshwinUgale/tracelint) has setup details for each.\n\n## Where deterministic trace checks stop\n\ntracelint’s structural checks don’t tell you whether the final answer is [semantically correct](https://arize.com/docs/phoenix/evaluation/llm-evals). That’s a separate, task-specific evaluation problem.\n\n- Some checks depend on facts you declare, such as which tools have side effects. If those declarations are wrong, the verdicts will be too.\n- Repetition and unexplained argument values stay candidates, because retries and value transformations can be legitimate.\n\n## When to use deterministic trace checks, agent evals, and production investigation\n\nUse deterministic trace checks for structural facts you can prove from a run, evals for task-specific quality and behavior, and [production investigation](https://arize.com/resources/whats-an-agent-observability-platform/) for recurring or previously unknown failure patterns across many traces.\n\nThe three layers overlap, but each is best suited to a different part of the reliability problem:\n\n| Layer | Best suited for | \n|---|---|\n| Deterministic trace checks | Explicit structural rules that can be proven from a single run | \n| Evals | Application-specific quality, behavior, grounding, tool choice, and policy | \n| Production investigation | Discovering recurring or previously unknown failure patterns across many traces | \n\nThey also feed each other. [Signal](https://arize.com/blog/how-to-find-and-debug-agent-failures-your-evals-are-missing/) can surface a recurring trajectory problem in production. If that failure can then be expressed as a structural rule, a deterministic regression check can help keep it from coming back.\n\n## Try tracelint on a Phoenix trace\n\n```\npip install tracelint\ntracelint check spans.json --format openinference\n```\n\nThe [repo](https://github.com/AshwinUgale/tracelint) has setup details and a [Phoenix guide](https://github.com/AshwinUgale/tracelint/blob/main/docs/integrations/phoenix.md). Feedback and issues are welcome, especially traces that tracelint doesn’t handle well.", "url": "https://wpnews.pro/news/what-agent-traces-can-tell-you-without-an-llm-judge", "canonical_source": "https://arize.com/blog/agent-traces-without-llm-judge/", "published_at": "2026-10-06 14:00:07+00:00", "updated_at": "2026-10-06 14:16:57.388704+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "mlops", "ai-safety"], "entities": ["Arize Phoenix", "Ashwin Govind Ugale", "tracelint", "OpenInference", "LangGraph", "gpt-4o-mini", "Jenkins"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-agent-traces-can-tell-you-without-an-llm-judge", "markdown": "https://wpnews.pro/news/what-agent-traces-can-tell-you-without-an-llm-judge.md", "text": "https://wpnews.pro/news/what-agent-traces-can-tell-you-without-an-llm-judge.txt", "jsonld": "https://wpnews.pro/news/what-agent-traces-can-tell-you-without-an-llm-judge.jsonld"}}