Evaluate production traces with Jev-as-a-Judge directly in Arize AX Arize AX now integrates directly with TypeSafe AI, bringing native Jev-as-a-Judge evaluations into the Arize evaluation workflow, the company announced. In an Arize benchmark for hallucination detection, a threshold-tuned Jev matched Claude Opus 5 at 87% accuracy while running at roughly 1/300 of the cost and 23x the speed, though Jev's default cutoff performed materially worse without threshold tuning. The integration lets teams evaluate production traces with boolean, choice, and score-based questions and receive structured labels and confidence scores directly in Arize AX. Arize AX https://arize.com/products/ax/ now integrates directly with TypeSafe AI https://arize.com/docs/ax/security-and-settings/integrations-playground/typesafe , bringing native Jev-as-a-Judge evaluations https://arize.com/blog/jev-as-a-judge/ into the Arize evaluation workflow. Jev is TypeSafe AI’s System One decision model https://arize.com/blog/typesafe-jev-llm-judge/ , built for structured decisions rather than text generation. With the new integration, teams can evaluate traces with boolean, choice, and score-based questions and get structured labels and confidence scores back directly in Arize AX. A faster option for bounded evaluation tasks Not every evaluation needs a general-purpose LLM to reason through a response and generate an explanation. Many production evals ask relatively bounded questions: - Did the agent resolve the user’s request? - Did it choose the right tool? - Does this response satisfy a defined policy? - Which category does this interaction belong to? Jev is designed for this kind of classification, scoring, and routing. You provide a shared state from the trace and one or more typed questions, and Jev returns structured results in a single call. That gives AI teams another tool for matching the evaluator to the job. You can start by testing Jev on evaluations with explicit criteria and a fixed set of possible outcomes. Compare its judgment with human labels https://arize.com/glossary/human-evaluation/ on representative examples before deciding whether its accuracy, cost, and latency meet your requirements. An LLM judge may be useful when you need a written explanation, while an agent judge can gather additional context or investigate before reaching a verdict. Lower evaluation costs can make it practical to assess more production traffic, giving teams more examples of where agents fail and what to improve next. What we learned testing Jev We tested Jev on hallucination detection to compare its accuracy, cost, and latency with LLM judges. In one Arize benchmark for hallucination detection, a threshold-tuned Jev matched Claude Opus 5 at 87% accuracy , while running at roughly 1/300 of the cost and 23x the speed in that test. The important caveat was threshold tuning: Jev’s default cutoff performed materially worse, reinforcing the importance of calibrating evaluators against human-labeled data before deploying them broadly. We’ve also explored Jev for real-time guardrails and previously showed how to connect it to AX using a remote evaluator. Native Jev-as-a-Judge support removes that additional service layer and brings the workflow directly into AX. Go deeper on Jev We’ve been exploring where decision models like Jev fit into the AI evaluation stack. Catch up on the latest Arize research and tutorials: - TypeSafe’s Jev: Can decision models replace LLM judges? https://arize.com/blog/typesafe-jev-llm-judge/ An introduction to Jev, how decision models differ from generative LLMs, and where they can fit into AI evaluation workflows. - Jev vs. LLM-as-a-Judge: Accuracy and cost benchmarks https://arize.com/blog/jev-as-a-judge/ Our benchmark comparing Jev with Claude Opus 5 and GPT-5.6 Terra across accuracy, latency, and cost, including why evaluator threshold tuning matters. - Real-time LLM guardrails with Jev: comparing latency and cost https://arize.com/blog/llm-guardrails-jev/ A practical look at using Jev for latency-sensitive guardrails on agent inputs, responses, and tool calls. Get started To use Jev-as-a-Judge in Arize AX: 1. Add your TypeSafe AI API key under Settings → AI Providers → TypeSafe AI . 2. Create a new Jev-as-a-Judge evaluator. 3. Define the trace data Jev should evaluate and the questions it should answer. 4. Test the evaluator against human-labeled examples from your application. 5. Attach it to an online evaluation task to evaluate incoming production traces. Follow the Jev-as-a-Judge setup guide in our docs https://arize.com/docs/ax/evaluate/jev-as-a-judge jev-as-a-judge . You can also learn how to set up the TypeSafe AI integration in our docs.