# Jev-as-a-Judge for Agent Evaluations

> Source: <https://academy.dair.ai/resources/jev-as-a-judge>
> Published: 2026-10-07 01:07:17+00:00

**LLM-as-a-judge** means using a large language model (LLM) to evaluate the work of another model. You give the judge an output, or a pair of outputs, along with the criteria that matter, and it scores, ranks, or compares them. Did the answer follow the policy? Is response A better than response B? Did the agent actually finish the task?

This approach is useful because manual review does not scale. A person can carefully read a few dozen outputs, but a judge model can apply the same criteria to thousands of runs, every time you change a prompt, a tool, or a model.

Jev is a new model built specifically for this kind of decision, which makes it a natural fit for the judge role. Using Jev as the judging LLM is what this tutorial means by **Jev-as-a-Judge**.

Judging agents adds a twist, because an agent can sound correct even when its work failed. Imagine a refund agent telling a customer, "Your refund has been processed," when the refund tool actually timed out.

Jev-as-a-Judge catches this by checking what the agent did, not just what it said. Jev reads the request, each tool call and its result, and the final reply, then returns a verdict with a probability attached.

This guide walks through the idea with one refund example and a live playground. The [full lab](https://academy.dair.ai/labs/jev-as-a-judge-for-agent-evals) builds the complete evaluation workflow.
