{"slug": "evaluate-production-traces-with-jev-as-a-judge-directly-in-arize-ax", "title": "Evaluate production traces with Jev-as-a-Judge directly in Arize AX", "summary": "Arize AX now integrates directly with TypeSafe AI, bringing native Jev-as-a-Judge evaluations into the Arize evaluation workflow, the company announced. In an Arize benchmark for hallucination detection, a threshold-tuned Jev matched Claude Opus 5 at 87% accuracy while running at roughly 1/300 of the cost and 23x the speed, though Jev's default cutoff performed materially worse without threshold tuning. The integration lets teams evaluate production traces with boolean, choice, and score-based questions and receive structured labels and confidence scores directly in Arize AX.", "body_md": "[Arize AX](https://arize.com/products/ax/) now [integrates directly with TypeSafe AI](https://arize.com/docs/ax/security-and-settings/integrations-playground/typesafe), bringing native [**Jev-as-a-Judge** evaluations](https://arize.com/blog/jev-as-a-judge/) into the Arize evaluation workflow.\n\n[Jev is TypeSafe AI’s System One decision model](https://arize.com/blog/typesafe-jev-llm-judge/), built for structured decisions rather than text generation. With the new integration, teams can evaluate traces with boolean, choice, and score-based questions and get structured labels and confidence scores back directly in Arize AX.\n\n## **A faster option for bounded evaluation tasks**\n\nNot every evaluation needs a general-purpose LLM to reason through a response and generate an explanation.\n\nMany production evals ask relatively bounded questions:\n\n- Did the agent resolve the user’s request?\n- Did it choose the right tool?\n- Does this response satisfy a defined policy?\n- Which category does this interaction belong to?\n\nJev is designed for this kind of classification, scoring, and routing. You provide a shared state from the trace and one or more typed questions, and Jev returns structured results in a single call.\n\nThat gives AI teams another tool for matching the evaluator to the job. You can start by testing Jev on evaluations with explicit criteria and a fixed set of possible outcomes. Compare its judgment with [human labels](https://arize.com/glossary/human-evaluation/) on representative examples before deciding whether its accuracy, cost, and latency meet your requirements. An LLM judge may be useful when you need a written explanation, while an agent judge can gather additional context or investigate before reaching a verdict.\n\nLower evaluation costs can make it practical to assess more production traffic, giving teams more examples of where agents fail and what to improve next.\n\n## **What we learned testing Jev**\n\nWe tested Jev on hallucination detection to compare its accuracy, cost, and latency with LLM judges.\n\nIn one Arize benchmark for hallucination detection, a threshold-tuned Jev matched Claude Opus 5 at **87% accuracy**, while running at roughly **1/300 of the cost and 23x the speed** in that test. The important caveat was threshold tuning: Jev’s default cutoff performed materially worse, reinforcing the importance of calibrating evaluators against human-labeled data before deploying them broadly.\n\nWe’ve also explored Jev for real-time guardrails and previously showed how to connect it to AX using a remote evaluator. Native Jev-as-a-Judge support removes that additional service layer and brings the workflow directly into AX.\n\n### **Go deeper on Jev**\n\nWe’ve been exploring where decision models like Jev fit into the AI evaluation stack. Catch up on the latest Arize research and tutorials:\n\n- [**TypeSafe’s Jev: Can decision models replace LLM judges?**](https://arize.com/blog/typesafe-jev-llm-judge/) An introduction to Jev, how decision models differ from generative LLMs, and where they can fit into AI evaluation workflows.\n- [**Jev vs. LLM-as-a-Judge: Accuracy and cost benchmarks**](https://arize.com/blog/jev-as-a-judge/) Our benchmark comparing Jev with Claude Opus 5 and GPT-5.6 Terra across accuracy, latency, and cost, including why evaluator threshold tuning matters.\n- [**Real-time LLM guardrails with Jev: comparing latency and cost**](https://arize.com/blog/llm-guardrails-jev/) A practical look at using Jev for latency-sensitive guardrails on agent inputs, responses, and tool calls.\n\n## **Get started**\n\nTo use Jev-as-a-Judge in Arize AX:\n\n1. Add your TypeSafe AI API key under **Settings → AI Providers → TypeSafe AI** .\n2. Create a new **Jev-as-a-Judge** evaluator.\n3. Define the trace data Jev should evaluate and the questions it should answer.\n4. Test the evaluator against human-labeled examples from your application.\n5. Attach it to an online evaluation task to evaluate incoming production traces.\n\n[**Follow the Jev-as-a-Judge setup guide in our docs**](https://arize.com/docs/ax/evaluate/jev-as-a-judge#jev-as-a-judge).\n\nYou can also learn how to **set up the TypeSafe AI integration in our docs.**", "url": "https://wpnews.pro/news/evaluate-production-traces-with-jev-as-a-judge-directly-in-arize-ax", "canonical_source": "https://arize.com/blog/arize-ax-jev-as-a-judge/", "published_at": "2026-09-28 21:00:35+00:00", "updated_at": "2026-09-28 21:18:19.314915+00:00", "lang": "en", "topics": ["ai-products", "ai-tools", "large-language-models", "mlops", "ai-agents"], "entities": ["Arize", "Arize AX", "TypeSafe AI", "Jev", "Claude Opus 5", "GPT-5.6 Terra"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/evaluate-production-traces-with-jev-as-a-judge-directly-in-arize-ax", "markdown": "https://wpnews.pro/news/evaluate-production-traces-with-jev-as-a-judge-directly-in-arize-ax.md", "text": "https://wpnews.pro/news/evaluate-production-traces-with-jev-as-a-judge-directly-in-arize-ax.txt", "jsonld": "https://wpnews.pro/news/evaluate-production-traces-with-jev-as-a-judge-directly-in-arize-ax.jsonld"}}