cd /news/artificial-intelligence/jev-ai-what-if-gordon-ramsay-didnt-c… · home › topics › artificial-intelligence › article
[ARTICLE · art-141168] src=pub.towardsai.net ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Jev AI: What If Gordon Ramsay Didn’t Cook? He Just Picked Who Did.

TypeSafe AI released Jev, a System 1 decision model that returns typed, probabilistic outputs without generating text, at $0.042 per million input tokens with output tokens free. Jev uses a Parallel Sampler to compute all output probabilities in a single pass, achieving 70ms to 500ms response times and a 0% schema mismatch rate, and reduces decision logic to three primitives: Choice, Score, and Noul. The model was developed by former OpenAI researcher Diogo Almeida, who worked on the instruction-following methods behind ChatGPT.

by read7 min views1 publishedSep 28, 2026

Imagine hiring Gordon Ramsay. Not to cook. Not even to judge the food.

Just to stand in the kitchen and decide which AI gets to cook.

Claude, you’re on pasta. GPT, steak. Gemini, dessert. And Gordon goes home.

It sounds like a ridiculous use of a very expensive person.

It also happens to be a surprisingly good way of thinking about Jev, a new AI model from TypeSafe AI. Except replace the kitchen with software.

And replace “pasta” with things like: retry, route, escalate, call this tool, don’t call this tool..* That sounds almost too simple.*

So I went looking for the complicated part.

Here’s something I started noticing when thinking about how AI gets wired into real applications.

We keep asking language models to make decisions that don’t actually need much language.

When an application needs to make a basic decision like routing a support ticket or deciding whether to run a database query developers wrap a massive, multi-billion-parameter LLM in a giant prompt:

“You are a helpful assistant. Please analyze this message and return ONLY a JSON object with keys ‘department’ and ‘urgency’…”

Then we cross our fingers 🤞

We pay for dozens of generated tokens we don’t need, wait 3 to 300+ seconds for autoregressive generation, and set up retry loops just in case the model hallucination inserts a stray trailing comma or a paragraph of apologetic prose.

We turned high-speed software into an anxious chatbot wrapper.

Former OpenAI researcher Diogo Almeida (who helped build the instruction-following methods behind ChatGPT) realized that if we want real software automation, we don’t need models that talk better but we need models that act like System One thinking.

In Daniel Kahneman’s framework, System 1 refers to fast, intuitive judgment, while System 2 refers to slower, deliberate reasoning. TypeSafe borrows that distinction as the inspiration for its “System One” model category.

Jev is built strictly for System 1. It takes unstructured program state in and returns typed, probabilistic decisions out with zero text generation.

How does Jev achieve response times between 70ms and 500ms (up to 200x faster than traditional LLMs)?

It strips away autoregressive text generation entirely.

Instead of generating token N+1 conditioned sequentially on token N, Jev uses a Parallel Sampler.

It evaluates the entire input state and computes all required output probabilities simultaneously in a single hardware-aware pass.

Because output choices are constrained to schemas defined in advance, type errors are prevented by the predefined output schema, so the model cannot return a value outside the allowed type (0% schema mismatch rate). That guarantee is about output type and structure, not semantic correctness.

And because it doesn’t generate tokens, output tokens are FREE. You only pay for input tokens at $0.042 per million tokens

Instead of open-ended prompt engineering, Jev reduces all software decision logic into three explicit primitives:

🟢 1. Choice (Categorical Routing)

Selects one option from a predefined list and returns the exact probability distribution across all candidates.

🟡 2. Score (Ordered Rubrics)

Rates state against a concrete, multi-tier scale with confidence metrics.

🔵 3. Noul (Binary Probabilistic Judgments)

Evaluates a specific yes/no statement and returns an exact probability score P(Yes) from 0 to 1.

At this point, I stopped trying to understand Jev only from diagrams and descriptions. I opened the Playground.

And I gave it a support message:

Subject: Charged twice again!!Hi! this is the SECOND month in a row I've been billed twice for the Pro plan.I already emailed last month and nobody replied.I run my whole business on this.If it's not refunded today I'm canceling and disputing the charge with my bank.

Then I asked three different questions about the same piece of state.

{  "topic": {    "type": "choice",    "options": ["billing", "bug", "account", "feature"],    "instructions": "What is the primary issue the customer is writing about?"  },  "severity": {    "type": "score",    "rubric": [      "routine, no rush",      "should be handled today",      "urgent, customer is frustrated",      "critical, customer is about to churn or dispute"    ],    "instructions": "How urgent and high-risk is this message?"  },  "escalate": {    "type": "noul",    "instructions": "Should this be escalated to a human agent immediately?"  }}

The Playground returned:

{  "topic": {    "type": "choice",    "choice": "billing",    "probabilities": {      "feature": 0,      "account": 0,      "bug": 0,      "billing": 1    },    "confidence": 1  },  "severity": {    "type": "score",    "score": 3,    "legend": {      "0": "routine, no rush",      "1": "should be handled today",      "2": "urgent, customer is frustrated",      "3": "critical, customer is about to churn or dispute"    },    "probabilities": {      "0": 0,      "1": 0,      "2": 0,      "3": 1    },    "confidence": 1  },  "escalate": {    "type": "noul",    "noul": 0.92  }}

Notice how your application backend can now branch directly on typed outputs without parsing a single word of text.

The most interesting number in the response wasn’t the 1.0.

It was the 0.92.

Jev had decided: escalate → 0.92

But what exactly does 0.92 mean?

A model can be right a lot of the time and still be terrible at knowing when it is likely to be wrong.

Imagine a model makes 100 decisions and gives all of them roughly 80% confidence.

If 80 of those decisions are correct, the number is doing something useful.

If only 55 are correct, then that 0.80 is mostly decorative.

This is the difference between accuracy and calibration.

Accuracy → How often was the model right?

Calibration → Does an 80% prediction actually behave like 80%?

And this is where RLCD, or Reinforcement Learning for Calibrated Decisions, comes in.

TypeSafe describes RLCD as its training approach for System One models, with the objective of producing probabilities that reflect uncertainty rather than simply optimizing for responses that humans prefer.

In its own description, a model should not just know how to make a decision; it should also communicate how likely that decision is to be correct. That matters enormously once the model is sitting inside code.

For example, an application might choose a policy like:

confidence > 0.90 → automate 0.60–0.90 → verify< 0.60 → ask a human

Those thresholds are an application-level policy, not something Jev decides for you.

A read-only classification might tolerate a lower threshold.

A payment, deletion, or other consequential action might require something much higher.

The current Jev documentation makes exactly this distinction: probability and confidence are intended to be control signals, while the application defines the thresholds and decides when to automate, verify, or escalate.

So RLCD isn’t just a fancy name for “the model gives confidence scores.”

The useful idea is: If software is going to act on a model’s judgment, uncertainty has to be part of the interface too.

And that makes Jev’s output a little more interesting than simply: billing

It becomes: billing + how confident the system is

To demonstrate what 70ms decision latency can unlock, TypeSafe showcases Jev across a few very different environments:

Looking at this comparison, the biggest shift isn’t just about saving money on tokens or cutting latencies down to 70ms. It’s a fundamental software design principle:

When building with traditional LLMs, developers accidentally hand the model total control. We ask GPT-4 to read an incident log, figure out what went wrong, write a explanation, format a JSON blob, and execute a tool call, hoping no step in that long chain fails or hallucinates.

System One models enforce a much cleaner contract inside code:

Maybe the interesting part isn’t making chatbots talk faster.

It’s figuring out which parts of an AI system actually need language in the first place.

Gordon Ramsay doesn’t need to cook every meal.

He just needs to point at the right stove.

Sometimes, AI doesn’t need to talk. It just needs to decide.

The examples, technical details, and benchmarks discussed in this article are based primarily on TypeSafe AI and Jev’s own documentation and published material.

Jev AI: What If Gordon Ramsay Didn’t Cook? He Just Picked Who Did. was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @typesafe ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/jev-ai-what-if-gordo…] indexed:0 read:7min 2026-09-28 · —