What agent traces can tell you without an LLM judge Arize Phoenix contributor Ashwin Govind Ugale released tracelint, an open-source linter that runs deterministic structural checks on OpenInference traces and returns a CI-usable exit code, avoiding LLM judges for failures provable from the trace alone. In a LangGraph release agent test using gpt-4o-mini, the agent deployed build 4.12.0 to production despite a Jenkins result of UNSTABLE with 3 of 215 tests failed; with "deploy only if its pipeline passes" in the system prompt it never deployed in ten runs, and with that sentence removed it deployed in all ten. tracelint flags deterministic defects such as tool arguments violating their JSON Schema, calls to nonexistent tools, and failed results fed into side-effecting actions, while surfacing repeated no-progress calls as candidates for review. This is a guest post by Ashwin Govind Ugale, a contributor to the Arize Phoenix https://arize.com/phoenix/ open source AI observability platform. TL;DR Some agent failures can be proven from the trace alone, such as a tool call that violates its schema or a failed result reused in a side effect. Other patterns, like repeated calls with no progress, can be surfaced as candidates without claiming they’re definitely bugs. A useful rule for deciding which is which: if the verdict depends on what the agent should have meant, it’s an eval. If it depends only on what the trace records, it can be a deterministic check. tracelint is a small open-source linter that runs those checks on OpenInference https://arize.com/glossary/openinference/ traces, including the ones Phoenix https://arize.com/phoenix/ collects, and returns an exit code you can use in CI. tracelint https://github.com/AshwinUgale/tracelint A run can look successful and still have a broken trajectory In a small experiment, I asked a LangGraph https://arize.com/docs/phoenix/integrations/python/langgraph release agent gpt-4o-mini to ship release 4.12.0. Its stubbed pipeline tool returns HTTP 200 with a Jenkins result of UNSTABLE 3 of 215 tests failed . The agent deploys that build to production anyway and replies: The release version 4.12.0 has been shipped to production. However, please note that the Jenkins build result was UNSTABLE, with 3 test failures out of 215 tests. Every span reports success. The agent noticed the failure, and shipped anyway. With “deploy only if its pipeline passes” in the system prompt, it never deployed in ten runs. With that sentence removed, it deployed in all ten. That’s the kind of regression CI exists to catch. Do you need an LLM judge https://arize.com/guides/llm-as-a-judge/ here? No. The pipeline failed by your team’s own definition, the same build ID shows up in the deploy call, and deploying has side effects. The verdict is already in the trace, plus two facts about your tools. I built tracelint to turn failures like this into repeatable checks you can run locally or in CI. Think of these as structural regression tests for agent behavior: they check claims the trace can prove and leave semantic judgement to agent evals https://arize.com/ai-agents/agent-evaluation/ . What agent failures can you prove from a trace? A useful rule of thumb is simple: if the verdict requires knowing what the agent was supposed to do, or what it meant, it probably isn’t a lint rule. If the verdict depends only on what the trace https://arize.com/docs/phoenix/tracing/concepts-tracing/what-are-traces records, plus facts you declare about your tools, it can be checked deterministically. A few examples along that line: - Arguments that violate the tool’s JSON Schema https://arize.com/blog/how-to-evaluate-tool-calling-agents/ : deterministic. - A call to a tool that doesn’t exist: deterministic. - A result declared as failed that’s fed into a side-effecting action: deterministic. - The same call repeatedly with no apparent progress: not necessarily a bug, since it might be polling or a legitimate retry. This gets surfaced for review. - The agent picked the wrong strategy: a question for an eval. tracelint includes several other structural checks, but they all follow the same principle. It reports a defect only when the trace provides enough evidence to prove it. How to run tracelint on an Arize Phoenix trace If you’re already tracing your agents with Arize Phoenix https://arize.com/docs/phoenix/tracing/how-to-tracing/setup-tracing , tracelint can read the OpenInference spans you collect without any changes to the agent: python from phoenix.client import Client from tracelint import lint otel traces, render report spans = Client .spans.get spans dataframe project name="release-agent" for report in lint otel traces spans.to dict "records" : one report per run print render report report Run tracelint on this trace as-is and it won’t fail it, and that’s correct: nothing in the trace says UNSTABLE means failure. That’s your CI’s convention. Two pieces of domain knowledge still have to come from you: - UNSTABLE like FAILURE and ABORTED means the pipeline failed - deploy takes a real action it changes production , so acting on a failed result there is a defect, not just something to review Declare those facts in a small tools.json : { "tools": { "run release pipeline": { "metadata": { "failure when": {"pointer": "/result", "in": "FAILURE", "UNSTABLE", "ABORTED" } } }, "deploy": { "metadata": {"side effecting": true} } } } Run the same trace again with those tool semantics declared: bash $ tracelint check spans.json --format openinference --tools tools.json 689451f502ecfcb8bc0063e2615ed8c5: 2 finding s , exit 2 hard event R2a tool error event step 3 'run release pipeline' returned a declared failure /result='UNSTABLE' hard defect R2b error mishandled step 3,4 value s from the errored 'run release pipeline' result b-4120-7f3a reused as arguments to 'deploy' a side-effecting action, no fallback … In plain terms: 1. The pipeline failed. Its result was UNSTABLE , which your tools.json says counts as a failure. 2. Its build was deployed anyway. The build ID b-4120-7f3a from that failed run was passed to deploy , which changes production. That’s a provable defect, so tracelint exits with code 2 and fails the CI job. You don’t have to write tools.json from scratch. tracelint init spans.json --format openinference -o tools.json drafts one from the trace, including tool schemas where the instrumentation records them. How tracelint classifies findings: hard defects, candidates, and not checked Every finding lands in one of three buckets, and the distinction matters more than any individual rule. - Hard defects are proven by the trace. In CI, tracelint check returns exit code 2 and fails the job. - Candidates are possible problems, such as repeated calls that might be a loop or might be a legitimate retry. They’re shown for review but never fail the build. That’s deliberate: one noisy heuristic is enough to make engineers disable a linter. - Not checked means a rule lacked the evidence it needed, and tracelint says so instead of counting it as a pass. For example, the tools.json above doesn’t say what a failed deploy looks like, so the same run’s output includes: suppressed 3 — not checked, not a clean pass: R2a tool error event: side-effecting tool 'deploy' returned an unclassifiable result and declares no failure when predicate — cannot verify it did not fail … tracelint won’t claim the deploy itself succeeded, because nothing tells it what failure would look like. If you don’t save traces in your tests yet, tracelint has a capture helper and a pytest fixture that record a run through your framework’s OpenInference instrumentor, plus a GitHub Action for CI. The repo https://github.com/AshwinUgale/tracelint has setup details for each. Where deterministic trace checks stop tracelint’s structural checks don’t tell you whether the final answer is semantically correct https://arize.com/docs/phoenix/evaluation/llm-evals . That’s a separate, task-specific evaluation problem. - Some checks depend on facts you declare, such as which tools have side effects. If those declarations are wrong, the verdicts will be too. - Repetition and unexplained argument values stay candidates, because retries and value transformations can be legitimate. When to use deterministic trace checks, agent evals, and production investigation Use deterministic trace checks for structural facts you can prove from a run, evals for task-specific quality and behavior, and production investigation https://arize.com/resources/whats-an-agent-observability-platform/ for recurring or previously unknown failure patterns across many traces. The three layers overlap, but each is best suited to a different part of the reliability problem: | Layer | Best suited for | |---|---| | Deterministic trace checks | Explicit structural rules that can be proven from a single run | | Evals | Application-specific quality, behavior, grounding, tool choice, and policy | | Production investigation | Discovering recurring or previously unknown failure patterns across many traces | They also feed each other. Signal https://arize.com/blog/how-to-find-and-debug-agent-failures-your-evals-are-missing/ can surface a recurring trajectory problem in production. If that failure can then be expressed as a structural rule, a deterministic regression check can help keep it from coming back. Try tracelint on a Phoenix trace pip install tracelint tracelint check spans.json --format openinference The repo https://github.com/AshwinUgale/tracelint has setup details and a Phoenix guide https://github.com/AshwinUgale/tracelint/blob/main/docs/integrations/phoenix.md . Feedback and issues are welcome, especially traces that tracelint doesn’t handle well.