Jev-as-a-Judge for Agent Evaluations A tutorial published by Jev introduces Jev-as-a-Judge, a method that uses the Jev model as an LLM-as-a-judge to evaluate AI agent runs by reading the request, each tool call and its result, and the final reply, then returning a verdict with a probability attached. The approach targets cases where an agent sounds correct despite failing, such as a refund agent claiming "Your refund has been processed" when the refund tool timed out, and is presented through one refund example plus a live playground, with a full lab building the complete evaluation workflow. LLM-as-a-judge means using a large language model LLM to evaluate the work of another model. You give the judge an output, or a pair of outputs, along with the criteria that matter, and it scores, ranks, or compares them. Did the answer follow the policy? Is response A better than response B? Did the agent actually finish the task? This approach is useful because manual review does not scale. A person can carefully read a few dozen outputs, but a judge model can apply the same criteria to thousands of runs, every time you change a prompt, a tool, or a model. Jev is a new model built specifically for this kind of decision, which makes it a natural fit for the judge role. Using Jev as the judging LLM is what this tutorial means by Jev-as-a-Judge . Judging agents adds a twist, because an agent can sound correct even when its work failed. Imagine a refund agent telling a customer, "Your refund has been processed," when the refund tool actually timed out. Jev-as-a-Judge catches this by checking what the agent did, not just what it said. Jev reads the request, each tool call and its result, and the final reply, then returns a verdict with a probability attached. This guide walks through the idea with one refund example and a live playground. The full lab https://academy.dair.ai/labs/jev-as-a-judge-for-agent-evals builds the complete evaluation workflow.