How I Test an LLM Feature: Field Notes on Golden Sets, LLM-as-Judge, and Eval Gates in CI A developer describes a three-part approach to testing LLM features that cannot be verified with string equality: a version-controlled golden set of real inputs, deterministic property assertions that run before any LLM-as-judge call, and a CI gate on pass-rate regression rather than a perfect score. The method was prompted by an incident in which a colleague's unrelated system-prompt edit caused a ticket-extraction route to invent order IDs while every existing test still passed. The author recommends splitting code at the model boundary so only the single model call needs an eval, and keeping fixtures as JSON/JSONL alongside the prompt file for reviewable diffs. Headline: An LLM feature cannot be unit-tested with string equality, because the same prompt returns different text on every call, so I assert on properties of the output instead of the output itself. The three changes that made my AI features safe to refactor were a version-controlled golden set of real inputs, deterministic assertions running before any LLM-as-judge call, and a CI gate on pass-rate regression rather than on a perfect score. toEqual assertions for any code path whose output comes from a model. I did not plan to build an eval harness. I shipped a route that turns a free-text support message into a structured ticket, and it worked. Weeks later a colleague edited the system prompt for an unrelated reason — a wording tweak in the tone instructions — and the extraction started inventing order IDs for messages that contained none. Every test in the repository still passed, because the only tests were for the code around the model. Nothing in CI was looking at what the model said. That is the gap evals fill. A large language model is a non-deterministic function, so expect output .toBe '...' fails for reasons unrelated to correctness. Setting temperature: 0 makes output far more stable but does not make it byte-identical: providers batch requests and floating-point accumulation differs between batches, so no major provider guarantees reproducible tokens. Asserting on exact strings produces a suite that is red on Tuesday and green on Wednesday, which teams learn to ignore within a sprint. The useful move is to split the code at the model boundary. Prompt assembly, retry logic, output parsing, tool dispatch, and every branch that consumes the parsed result are ordinary deterministic functions and deserve ordinary unit tests. Only the single call that crosses into the model needs an eval. Once the seam exists, the eval asserts on properties that are true of any correct answer: does it parse against the schema, does it stay in the requested language, does it cite only identifiers that appear in the input, does it refuse when the context contains no answer, does it never echo the system prompt. A golden set is a version-controlled file of real inputs paired with the properties their outputs must satisfy. Mine started at roughly a dozen cases — one per input shape I could name — and grew every time production surprised me. I do not write synthetic cases when real ones exist, because the cases that catch regressions are the awkward real ones: the empty message, the message written in Arabic, the message that is only an order number, the message that is a prompt-injection attempt pasted out of an email. Two storage rules have paid for themselves. First, keep the fixtures as JSON or JSONL next to the prompt they test, so the case list is reviewable in a pull request. Second, keep the prompt in its own file rather than in a template literal inside a route handler, so a prompt change and its eval results show up in the same diff. js // evals/ticket-extraction.eval.test.ts import { describe, it, expect } from 'vitest'; import { generateObject } from 'ai'; import { z } from 'zod'; import cases from './fixtures/tickets.json'; import { TICKET PROMPT } from '../src/prompts/ticket'; const Ticket = z.object { category: z.enum 'billing', 'bug', 'feature', 'other' , urgency: z.number .int .min 1 .max 5 , summary: z.string .max 200 , orderIds: z.array z.string , } ; describe 'ticket extraction', = { for const c of cases { it.concurrent c.name, async = { const { object } = await generateObject { model: 'anthropic/claude-sonnet-5', // pinned ID, routed via AI Gateway schema: Ticket, temperature: 0, system: TICKET PROMPT, prompt: c.input, } ; expect object.category .toBe c.expected.category ; // grounding: never invent an order ID that is not in the message for const id of object.orderIds expect c.input .toContain id ; } ; } } ; generateObject from the Vercel AI SDK is doing double duty here: it validates the model output against the Zod schema before returning, so a shape regression throws instead of flowing into an assertion. That is the cheapest eval in the file and it is the one that would have caught my invented-order-ID bug. Use a deterministic check whenever the failure can be expressed as a predicate over the string, and reach for a judge only for properties that require reading comprehension. Deterministic checks are free, instant, and never disagree with themselves; a judge is a second model call with its own error rate and its own bill. | Property to check | Method | Why | |---|---|---| | Output shape and JSON validity | Deterministic Zod schema | Schema validation is exact; a judge adds cost and can only be worse. | | Grounding: every ID, price, or quote appears in the input | Deterministic substring or set check | Hallucinated identifiers are string-detectable, and this is the most damaging class of error. | | Forbidden content: prompt leakage, internal URLs, competitor names | Deterministic regex denylist | Safety properties must be exact, not probabilistic. | | Language, length, and format constraints | Deterministic | Word counts and script ranges are trivially computable. | | Did the answer actually address the question? | LLM-as-judge | Requires comprehension; no predicate expresses it. | | Is the answer semantically equivalent to a reference? | LLM-as-judge or embedding similarity | Many correct phrasings exist, so exact matching under-reports success. | Treat the judge as production code with a version, not as an oracle. Pin the exact model ID — claude-sonnet-5 , never a floating alias — because a silent provider-side model change reads as a quality regression in your product. Ask binary questions instead of 1–10 scores: "does the answer state a refund deadline, yes or no" is reproducible, while "rate helpfulness out of ten" drifts and turns your threshold into an arbitrary number. Keep the rubric in a committed file that is reviewed like any other source. And calibrate: hand-label twenty to thirty outputs yourself, run the judge over the same outputs, and look at where the two disagree. A judge you have never compared to a human is a number, not a measurement. Where budget allows, use a different model family for the judge than for the generator, because models tend to rate their own phrasing style favourably. js const Verdict = z.object { grounded: z.boolean , answersQuestion: z.boolean , reason: z.string .max 300 , } ; export async function judge question: string, answer: string, context: string { const { object } = await generateObject { model: 'anthropic/claude-sonnet-5', // pinned; bump in a reviewed commit schema: Verdict, temperature: 0, system: JUDGE RUBRIC, // committed file, binary criteria only prompt: