cd /news/ai-products/evaluate-production-traces-with-jev-… · home › topics › ai-products › article
[ARTICLE · art-141287] src=arize.com ↗ pub= topic=ai-products verified=true sentiment=↑ positive

Evaluate production traces with Jev-as-a-Judge directly in Arize AX

Arize AX now integrates directly with TypeSafe AI, bringing native Jev-as-a-Judge evaluations into the Arize evaluation workflow, the company announced. In an Arize benchmark for hallucination detection, a threshold-tuned Jev matched Claude Opus 5 at 87% accuracy while running at roughly 1/300 of the cost and 23x the speed, though Jev's default cutoff performed materially worse without threshold tuning. The integration lets teams evaluate production traces with boolean, choice, and score-based questions and receive structured labels and confidence scores directly in Arize AX.

by read3 min views1 publishedSep 28, 2026
Evaluate production traces with Jev-as-a-Judge directly in Arize AX
Image: Arize (auto-discovered)

Arize AX now integrates directly with TypeSafe AI, bringing native Jev-as-a-Judge evaluations into the Arize evaluation workflow.

Jev is TypeSafe AI’s System One decision model, built for structured decisions rather than text generation. With the new integration, teams can evaluate traces with boolean, choice, and score-based questions and get structured labels and confidence scores back directly in Arize AX.

A faster option for bounded evaluation tasks #

Not every evaluation needs a general-purpose LLM to reason through a response and generate an explanation.

Many production evals ask relatively bounded questions:

  • Did the agent resolve the user’s request?
  • Did it choose the right tool?
  • Does this response satisfy a defined policy?
  • Which category does this interaction belong to?

Jev is designed for this kind of classification, scoring, and routing. You provide a shared state from the trace and one or more typed questions, and Jev returns structured results in a single call.

That gives AI teams another tool for matching the evaluator to the job. You can start by testing Jev on evaluations with explicit criteria and a fixed set of possible outcomes. Compare its judgment with human labels on representative examples before deciding whether its accuracy, cost, and latency meet your requirements. An LLM judge may be useful when you need a written explanation, while an agent judge can gather additional context or investigate before reaching a verdict.

Lower evaluation costs can make it practical to assess more production traffic, giving teams more examples of where agents fail and what to improve next.

What we learned testing Jev #

We tested Jev on hallucination detection to compare its accuracy, cost, and latency with LLM judges.

In one Arize benchmark for hallucination detection, a threshold-tuned Jev matched Claude Opus 5 at 87% accuracy, while running at roughly 1/300 of the cost and 23x the speed in that test. The important caveat was threshold tuning: Jev’s default cutoff performed materially worse, reinforcing the importance of calibrating evaluators against human-labeled data before deploying them broadly.

We’ve also explored Jev for real-time guardrails and previously showed how to connect it to AX using a remote evaluator. Native Jev-as-a-Judge support removes that additional service layer and brings the workflow directly into AX.

Go deeper on Jev

We’ve been exploring where decision models like Jev fit into the AI evaluation stack. Catch up on the latest Arize research and tutorials:

Get started #

To use Jev-as-a-Judge in Arize AX:

  1. Add your TypeSafe AI API key under Settings → AI Providers → TypeSafe AI .
  2. Create a new Jev-as-a-Judge evaluator.
  3. Define the trace data Jev should evaluate and the questions it should answer.
  4. Test the evaluator against human-labeled examples from your application.
  5. Attach it to an online evaluation task to evaluate incoming production traces.

Follow the Jev-as-a-Judge setup guide in our docs. You can also learn how to set up the TypeSafe AI integration in our docs.

── more in #ai-products 4 stories · sorted by recency
── more on @arize 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evaluate-production-…] indexed:0 read:3min 2026-09-28 · —