This is a guest post by Ashwin Govind Ugale, a contributor to the Arize Phoenix open source AI observability platform.
TL;DR
Some agent failures can be proven from the trace alone, such as a tool call that violates its schema or a failed result reused in a side effect. Other patterns, like repeated calls with no progress, can be surfaced as candidates without claiming they’re definitely bugs.
A useful rule for deciding which is which: if the verdict depends on what the agent should have meant, it’s an eval. If it depends only on what the trace records, it can be a deterministic check.
tracelint is a small open-source linter that runs those checks on OpenInference traces, including the ones Phoenix collects, and returns an exit code you can use in CI. tracelint
A run can look successful and still have a broken trajectory #
In a small experiment, I asked a LangGraph release agent (gpt-4o-mini) to ship release 4.12.0. Its stubbed pipeline tool returns HTTP 200 with a Jenkins result of UNSTABLE (3 of 215 tests failed). The agent deploys that build to production anyway and replies:
The release version 4.12.0 has been shipped to production. However, please note that the Jenkins build result was UNSTABLE, with 3 test failures out of 215 tests.
Every span reports success. The agent noticed the failure, and shipped anyway.
With “deploy only if its pipeline passes” in the system prompt, it never deployed in ten runs. With that sentence removed, it deployed in all ten. That’s the kind of regression CI exists to catch.
Do you need an LLM judge here? No. The pipeline failed by your team’s own definition, the same build ID shows up in the deploy call, and deploying has side effects. The verdict is already in the trace, plus two facts about your tools. I built tracelint to turn failures like this into repeatable checks you can run locally or in CI.
Think of these as structural regression tests for agent behavior: they check claims the trace can prove and leave semantic judgement to agent evals.
What agent failures can you prove from a trace? #
A useful rule of thumb is simple: if the verdict requires knowing what the agent was supposed to do, or what it meant, it probably isn’t a lint rule. If the verdict depends only on what the trace records, plus facts you declare about your tools, it can be checked deterministically.
A few examples along that line:
- Arguments that violate the tool’s JSON Schema : deterministic.
- A call to a tool that doesn’t exist: deterministic.
- A result declared as failed that’s fed into a side-effecting action: deterministic.
- The same call repeatedly with no apparent progress: not necessarily a bug, since it might be polling or a legitimate retry. This gets surfaced for review.
- The agent picked the wrong strategy: a question for an eval.
tracelint includes several other structural checks, but they all follow the same principle. It reports a defect only when the trace provides enough evidence to prove it.
How to run tracelint on an Arize Phoenix trace #
If you’re already tracing your agents with Arize Phoenix, tracelint can read the OpenInference spans you collect without any changes to the agent:
from phoenix.client import Client
from tracelint import lint_otel_traces, render_report
spans = Client().spans.get_spans_dataframe(project_name="release-agent")
for report in lint_otel_traces(spans.to_dict("records")): # one report per run
print(render_report(report))
Run tracelint on this trace as-is and it won’t fail it, and that’s correct: nothing in the trace says UNSTABLE means failure. That’s your CI’s convention. Two pieces of domain knowledge still have to come from you:
UNSTABLE(likeFAILUREandABORTED) means the pipeline faileddeploytakes a real action (it changes production), so acting on a failed result there is a defect, not just something to review
Declare those facts in a small tools.json:
{
"tools": {
"run_release_pipeline": {
"metadata": {
"failure_when": {"pointer": "/result", "in": ["FAILURE", "UNSTABLE", "ABORTED"]}
}
},
"deploy": {
"metadata": {"side_effecting": true}
}
}
}
Run the same trace again with those tool semantics declared:
$ tracelint check spans.json --format openinference --tools tools.json
689451f502ecfcb8bc0063e2615ed8c5: 2 finding(s), exit 2
[hard_event] R2a tool_error_event (step 3)
'run_release_pipeline' returned a declared failure (/result='UNSTABLE')
[hard_defect] R2b error_mishandled (step 3,4)
value(s) from the errored 'run_release_pipeline' result (b-4120-7f3a) reused
as arguments to 'deploy' (a side-effecting action, no fallback)
…
In plain terms:
- The pipeline failed. Its result was
UNSTABLE, which yourtools.jsonsays counts as a failure. - Its build was deployed anyway. The build ID
b-4120-7f3afrom that failed run was passed todeploy, which changes production. That’s a provable defect, so tracelint exits with code 2 and fails the CI job.
You don’t have to write tools.json from scratch. tracelint init spans.json --format openinference -o tools.json drafts one from the trace, including tool schemas where the instrumentation records them.
How tracelint classifies findings: hard defects, candidates, and not checked #
Every finding lands in one of three buckets, and the distinction matters more than any individual rule.
- Hard defects are proven by the trace. In CI,
tracelint checkreturns exit code 2 and fails the job. - Candidates are possible problems, such as repeated calls that might be a loop or might be a legitimate retry. They’re shown for review but never fail the build. That’s deliberate: one noisy heuristic is enough to make engineers disable a linter.
- Not checked means a rule lacked the evidence it needed, and tracelint says so instead of counting it as a pass.
For example, the tools.json above doesn’t say what a failed deploy looks like, so the same run’s output includes:
suppressed (3) — not checked, not a clean pass:
R2a tool_error_event: side-effecting tool 'deploy' returned an unclassifiable
result and declares no failure_when predicate — cannot verify it did not fail
…
tracelint won’t claim the deploy itself succeeded, because nothing tells it what failure would look like.
If you don’t save traces in your tests yet, tracelint has a capture helper and a pytest fixture that record a run through your framework’s OpenInference instrumentor, plus a GitHub Action for CI. The repo has setup details for each.
Where deterministic trace checks stop #
tracelint’s structural checks don’t tell you whether the final answer is semantically correct. That’s a separate, task-specific evaluation problem.
- Some checks depend on facts you declare, such as which tools have side effects. If those declarations are wrong, the verdicts will be too.
- Repetition and unexplained argument values stay candidates, because retries and value transformations can be legitimate.
When to use deterministic trace checks, agent evals, and production investigation #
Use deterministic trace checks for structural facts you can prove from a run, evals for task-specific quality and behavior, and production investigation for recurring or previously unknown failure patterns across many traces.
The three layers overlap, but each is best suited to a different part of the reliability problem:
| Layer | Best suited for |
|---|---|
| Deterministic trace checks | Explicit structural rules that can be proven from a single run |
| Evals | Application-specific quality, behavior, grounding, tool choice, and policy |
| Production investigation | Discovering recurring or previously unknown failure patterns across many traces |
They also feed each other. Signal can surface a recurring trajectory problem in production. If that failure can then be expressed as a structural rule, a deterministic regression check can help keep it from coming back.
Try tracelint on a Phoenix trace #
pip install tracelint
tracelint check spans.json --format openinference
The repo has setup details and a Phoenix guide. Feedback and issues are welcome, especially traces that tracelint doesn’t handle well.