{"slug": "no-need-to-sample", "title": "No need to sample", "summary": "A blogger's comparison test found that TypeSafe's Jev model scored 64 generated game levels per schema variant in about a quarter of a second for a fraction of a cent, agreeing with an earlier Haiku judge on variant rankings (Spearman coefficient 0.58 on the weighted total) while running inline rather than after the fact. Jev's Likert-style scores ran more lenient and compressed than Haiku's — for example nested_narrative scored 4.64 versus 4.32 and append_order broke 7.8% versus 8.9% — but both judges ranked nested_narrative cleanest and append_order most broken. The author argues the cost and latency mean every generation can be judged inline instead of sampling a few percent of outputs.", "body_md": "# No need to sample\n\nYesterday I did some [initial exploration](https://jdhornsby.com/fifty-cents-of-jev/) of [Jev](https://docs.typesafe.ai/introduction). I tried a few use cases and was impressed by all of them except chess, which it played terribly, at least under my harness. But I wanted to try what I think is the killer use case: judging LLM output.\n\nAlongside programmatic checks, it’s common to use one LLM to judge the output of another. That comes with two constraints. A judge call adds seconds, so you run it after the fact, which is fine for a slow feedback loop and bad for catching a mistake as it happens. And a judge call is relatively expensive, so you score a few percent of outputs and assume the rest look like the sample. Basic economics.\n\nIn this post, I take a Haiku judge I’d already built and run, a [reachability check](https://jdhornsby.com/a-number-that-looks-fine/) on generated game levels, and point Jev at the same dataset. It’s fast, it’s dirt cheap, and it mostly agrees with Haiku.\n\nI think the economics just changed.\n\n## When you can judge everything\n\nJev obliterates both constraints. It can score a piece of content on multiple dimensions in a quarter of a second for a fraction of a cent. At that price there is no reason to sample. At that latency you can do it before sending the result to a user.\n\nThis is not an original idea. TypeSafe is pretty clear that this is what they built Jev for. They even named the model after the [economist](https://en.wikipedia.org/wiki/Jevons_paradox) who noticed that making something cheaper tends to mean you use far more of it.\n\nJev is well suited to use cases that require clear judgment calls about a piece of content. Does every claim in this answer have a citation? Does every citation point at a real source? Does this summary introduce a number that wasn’t in the source material?\n\nIt is also competent at use cases where semantic meaning matters and the judgment is fuzzier. Is the tone right for a customer who is already angry? Does this response answer the question or restate it? Is this within policy, or does it need a human?\n\nIt does take some design work. My first attempt was a straight port of an old six-criterion rubric, and the scores came back compressed: the same rankings as the original judge, but with much less spread. I’d bet that’s a me problem. Narrow questions with explicit criteria are what Jev is built for, and a rubric written for a generative judge isn’t that. It’s a different skill, and I’m still learning it.\n\nThe pattern is simple. Judge every important generation inline. Act when confidence is high, escalate when it isn’t. Catch mistakes early.\n\n## The numbers\n\nI had a dataset from another experiment I ran a while ago. I was exploring the impact of schema design on structured output (my favorite topic lately). I had an LLM generate level configs for a fake game using six schema variants ranging from intentionally confounding to hopefully optimal. It included two types of judgments: a Likert-style score on six dimensions that tried to measure how good the level was and a more mechanical check on whether the end of the level was reachable.\n\nI converted these judgments to a Jev `Score` and `Noul` respectively. Here is how they compare to the original Haiku results (n=64 levels per variant, three encounters each):\n\n| Variant | Original Likert | Jev Likert | Original break % | Jev break % | \n|---|---|---|---|---|\n| flat_alpha | 3.84 | 4.46 | 5.2 | 3.6 | \n| alpha_nested | 3.84 | 4.47 | 4.2 | 2.6 | \n| ui_contract | 3.89 | 4.56 | 1.0 | 1.0 | \n| append_order | 4.18 | 4.55 | 8.9 | 7.8 | \n| grouped_by_type | 4.18 | 4.62 | 1.6 | 1.0 | \n| nested_narrative | 4.32 | 4.64 | 0.0 | 0.0 | \n\nBoth judges rank the variants the same way, with nested_narrative clean and append_order breaking most often.\n\nOn the Likert-style rubric, Jev was more lenient and compressed. Zooming in, the two judges did have a positive Spearman coefficient (0.58 on the weighted total), so they generally ranked them similarly. In a real world scenario, I could try moving my threshold, but I think I’d redesign the rubric in this case.\n\nThe break percentage measures how many levels are not logically completable. These line up closely, with Jev a little more lenient on most variants. I was curious about where it differed, so I compared each and grouped them by what happened:\n\n| Category | Verdict | Jev confidence | n | \n|---|---|---|---|\n| semantic match | Jev more accurate - Haiku missed it reading literally | 0.71–0.96 | 6 | \n| plausible inference | debatable | 0.60–0.83 | 3 | \n| semantic overreach | Haiku more accurate - Jev too generous | 0.50–0.67 | 2 | \n\nThe two calls Jev got wrong are also the two it was least sure about. Everything defensible starts at 0.71. That’s the threshold behavior you’d want from a gate, and it’s why this felt natural to use for this kind of check.\n\nI then reran the full set 9 more times. Variant means held to within 0.006, though individual scores did move between runs (mean spread 0.11 on a 0–4 scale). It’s stable in aggregate, but not per item. The reruns also gave me latency numbers:\n\nThis is calling from my laptop on the East Coast. I’d love to get numbers with more favorable regional colocation and networking.\n\nI spent $0.56 on the 5,866 requests I made in this test and related probing. That’s about a hundredth of a cent per call.\n\n## Scope and limits\n\nOne dataset, one domain, and one I built myself to stress exactly this kind of interdependency. The results are specific to my example. YMMV.\n\nThe comparison is against a Haiku judge, not ground truth, and where they disagreed I adjudicated the cases myself. The original judgments also have the problems I wrote about in [a number that looks fine](https://jdhornsby.com/a-number-that-looks-fine/).\n\nAll timings are from my laptop on the East Coast, against `jev-1.13`. I expect things will change rapidly.\n\nRun it yourself: [repo](https://github.com/jdhornsby/typesafe-jev)", "url": "https://wpnews.pro/news/no-need-to-sample", "canonical_source": "https://jdhornsby.com/no-need-to-sample/", "published_at": "2026-09-20 20:00:00+00:00", "updated_at": "2026-09-28 12:49:39.548811+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "ai-products", "artificial-intelligence"], "entities": ["Jev", "TypeSafe", "Haiku", "Anthropic"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/no-need-to-sample", "markdown": "https://wpnews.pro/news/no-need-to-sample.md", "text": "https://wpnews.pro/news/no-need-to-sample.txt", "jsonld": "https://wpnews.pro/news/no-need-to-sample.jsonld"}}