For as long as software has been able to ask a model a question, almost every answer has come back as prose. Ask whether a support ticket is billing or account, and you get a paragraph. "This looks like a billing issue, though it could also be an account access problem — it depends on whether the charge was a renewal." A person reads that and understands it in a second. A program cannot. Somebody has to turn writing back into a decision, and the machinery for that is regex, keyword lists, a second smaller model, or hope. The judgment is real. It simply arrives in a container built for humans.
On September 15, 2026, TypeSafe AI shipped Jev, and the container came off. You send program state plus typed questions, and you get back choices, scores and probabilities, computed in parallel, each carrying a confidence value — and not one sentence of prose. That is a small thing to describe and a large thing to build. Before I say anything about where it stops, I want to state clearly what it got right, because it is a genuine engineering result and not a marketing claim:
It turned judgment into an interface.
The first consequence is about shape.
A paragraph is an open surface. Anything plausible can be written on it, including a confident summary of a decision that was never actually made. A typed value is a closed surface: the set of things it can be was defined by the caller. You cannot improvise the shape of a value. You can only be wrong about which value it is.
That moves the failure mode from ambiguous to incorrect, and those are different engineering problems. Ambiguity is resolved by reading harder. Incorrectness is resolved by measurement, which means you can write a test for it. Once the answer to "which of these four data sources should I read" arrives as a field with four legal values and a probability, the surrounding system stops being a pipeline with a language model in the middle and becomes a pipeline with a component in it — a component that has an input contract and an output contract, like every other component you own.
Plenty of people will say this is just JSON-mode classification, and that classifiers are old. Fine. Then notice that nobody put a classifier in the path of every tool call before, because the classifier was either too weak to trust or too expensive to run per call. Which is the second thing.
TypeSafe's own published figures put Jev at 70–500 milliseconds and $0.042 per million input tokens, with output free. Those are vendor numbers, not ours, and they are the whole point.
At a price measured in millionths of a dollar, judgment stops being a budget line and becomes a rounding error. At a latency measured in milliseconds, it stops being a request and becomes a step. The interesting effect is not that the same decisions got cheaper. It is that a class of decisions which was previously never made at all now gets made. Nobody was going to pay a person, or a full model call, to ask whether this particular merge should wait for review. So it did not get asked. Now it does.
That is a real expansion of what software is permitted to notice about itself, and it belongs to Jev.
Adoption figures deserve suspicion, so I will label them for what they are. The company raised a $40M seed led by DCVC, and its founder, Diogo Almeida, co-authored InstructGPT and the RLHF work at OpenAI — company and public record. The launch thread on Hacker News has run to 1,989 points and 520 comments (as of October 1), which you can read straight off Hacker News' own public API. Vercel reported that within 24 hours of listing the model, nearly 13% of its paid teams had already used it, and described it as the fastest-adopted model it had seen on its own gateway. Also platform-reported.
But the number I find most persuasive is not the largest one. It is that gateways, orchestration frameworks and observability tooling — organizations with no shared roadmap and no coordination committee — each wired up the same primitive inside the same short window. Interfaces get adopted at that speed when they are obvious: when the thing they expose is already the thing everyone needed and could not name. A choice, a score, a probability. The naming was part of the contribution.
Everything above is why I think Jev is correct, and I am not going to follow it with a "but" that takes any of it back. Cheap judgment solved the problem it was aimed at: too many small judgments, each individually too expensive to make. That problem is now solved, and it stays solved.
Here is the problem standing beside it, which is a different problem rather than a criticism.
Making a judgment cheap requires making it a value, and a value has a domain. The domain is the options you passed in. That closure is exactly what makes the thing computable, and it is also why one particular kind of answer has nowhere to live: these two cannot be separated with what you have given me.
That answer is not a defect state and it is not a low probability. It is a different kind of finding, and it shows up precisely when two candidates both survive an honest reading of the question. In human affairs those are frequently the decisions that matter most: made once, not repeatable, signed by a named person, lived with for years. A per-call-priced interface has no field in which "this one deserves more of your attention than the last thousand" can be expressed. Every call costs the same, so every call looks the same size.
So the two ends of the same axis look like this. One end made judgment cheap, at scale, for the decisions where being wrong costs a re-run. It did that properly, and it is right. The other end is where the two sides really are level, and where the only useful outputs are: this is level; a confidence number that has actually been calibrated, so that a band means something; and an argument long enough to disagree with. An answer of that kind costs minutes. In a decision you will live with for years, the minutes are the cheapest thing in the room.
We are the second kind of judge, so it is fair to ask for our numbers instead of our adjectives.
They come from our own published, self-run evaluations, on named benchmarks, with failures disclosed rather than quietly retried. On JudgeBench: 620 judgments, with the 6 that failed on first verdict disclosed. Our accuracy came in at 92.5% against 92.2% for a plain direct baseline. That is a tie, and we report it as a tie; we claim no accuracy edge over a direct model call or over anyone else. On the same run, judgments we reported at 90% confidence or above were right 99.6% of the time, and the 80–90% band was right 94.0%. The number I would most want you to carry away is the least flattering one: on ContextualJudgeBench, run by us over the full official set, with 12 orders excluded after repeated platform failures and disclosed rather than imputed, the deliberately constructed near-ties sit at 46–60%.
We publish that last figure because it is the honest scale of the problem at this end of the axis. When the answer is genuinely close, the correct output is to say so — and then to spend the minutes, because that is what a decision of that weight is worth.
Jev showed that a judgment can be a value a program switches on, and made that value fast and nearly free — the right answer for the decisions that should never have been slow. At the other end of that same axis, where the two sides are level and a person still has to sign, we build a judge that spends the minutes and shows its work: Decider.