# Evaluate production traces with Jev-as-a-Judge directly in Arize AX

> Source: <https://arize.com/blog/arize-ax-jev-as-a-judge/>
> Published: 2026-09-28 21:00:35+00:00

[Arize AX](https://arize.com/products/ax/) now [integrates directly with TypeSafe AI](https://arize.com/docs/ax/security-and-settings/integrations-playground/typesafe), bringing native [**Jev-as-a-Judge** evaluations](https://arize.com/blog/jev-as-a-judge/) into the Arize evaluation workflow.

[Jev is TypeSafe AI’s System One decision model](https://arize.com/blog/typesafe-jev-llm-judge/), built for structured decisions rather than text generation. With the new integration, teams can evaluate traces with boolean, choice, and score-based questions and get structured labels and confidence scores back directly in Arize AX.

## **A faster option for bounded evaluation tasks**

Not every evaluation needs a general-purpose LLM to reason through a response and generate an explanation.

Many production evals ask relatively bounded questions:

- Did the agent resolve the user’s request?
- Did it choose the right tool?
- Does this response satisfy a defined policy?
- Which category does this interaction belong to?

Jev is designed for this kind of classification, scoring, and routing. You provide a shared state from the trace and one or more typed questions, and Jev returns structured results in a single call.

That gives AI teams another tool for matching the evaluator to the job. You can start by testing Jev on evaluations with explicit criteria and a fixed set of possible outcomes. Compare its judgment with [human labels](https://arize.com/glossary/human-evaluation/) on representative examples before deciding whether its accuracy, cost, and latency meet your requirements. An LLM judge may be useful when you need a written explanation, while an agent judge can gather additional context or investigate before reaching a verdict.

Lower evaluation costs can make it practical to assess more production traffic, giving teams more examples of where agents fail and what to improve next.

## **What we learned testing Jev**

We tested Jev on hallucination detection to compare its accuracy, cost, and latency with LLM judges.

In one Arize benchmark for hallucination detection, a threshold-tuned Jev matched Claude Opus 5 at **87% accuracy**, while running at roughly **1/300 of the cost and 23x the speed** in that test. The important caveat was threshold tuning: Jev’s default cutoff performed materially worse, reinforcing the importance of calibrating evaluators against human-labeled data before deploying them broadly.

We’ve also explored Jev for real-time guardrails and previously showed how to connect it to AX using a remote evaluator. Native Jev-as-a-Judge support removes that additional service layer and brings the workflow directly into AX.

### **Go deeper on Jev**

We’ve been exploring where decision models like Jev fit into the AI evaluation stack. Catch up on the latest Arize research and tutorials:

- [**TypeSafe’s Jev: Can decision models replace LLM judges?**](https://arize.com/blog/typesafe-jev-llm-judge/) An introduction to Jev, how decision models differ from generative LLMs, and where they can fit into AI evaluation workflows.
- [**Jev vs. LLM-as-a-Judge: Accuracy and cost benchmarks**](https://arize.com/blog/jev-as-a-judge/) Our benchmark comparing Jev with Claude Opus 5 and GPT-5.6 Terra across accuracy, latency, and cost, including why evaluator threshold tuning matters.
- [**Real-time LLM guardrails with Jev: comparing latency and cost**](https://arize.com/blog/llm-guardrails-jev/) A practical look at using Jev for latency-sensitive guardrails on agent inputs, responses, and tool calls.

## **Get started**

To use Jev-as-a-Judge in Arize AX:

1. Add your TypeSafe AI API key under **Settings → AI Providers → TypeSafe AI** .
2. Create a new **Jev-as-a-Judge** evaluator.
3. Define the trace data Jev should evaluate and the questions it should answer.
4. Test the evaluator against human-labeled examples from your application.
5. Attach it to an online evaluation task to evaluate incoming production traces.

[**Follow the Jev-as-a-Judge setup guide in our docs**](https://arize.com/docs/ax/evaluate/jev-as-a-judge#jev-as-a-judge).

You can also learn how to **set up the TypeSafe AI integration in our docs.**
