{"slug": "jev-ai-what-if-gordon-ramsay-didnt-cook-he-just-picked-who-did", "title": "Jev AI: What If Gordon Ramsay Didn’t Cook? He Just Picked Who Did.", "summary": "TypeSafe AI released Jev, a System 1 decision model that returns typed, probabilistic outputs without generating text, at $0.042 per million input tokens with output tokens free. Jev uses a Parallel Sampler to compute all output probabilities in a single pass, achieving 70ms to 500ms response times and a 0% schema mismatch rate, and reduces decision logic to three primitives: Choice, Score, and Noul. The model was developed by former OpenAI researcher Diogo Almeida, who worked on the instruction-following methods behind ChatGPT.", "body_md": "Imagine hiring Gordon Ramsay. Not to cook. Not even to judge the food.\n\nJust to stand in the kitchen and decide **which AI gets to cook.**\n\nClaude, you’re on *pasta*. GPT, *steak*. Gemini, *dessert*. And Gordon goes home.\n\nIt sounds like a ridiculous use of a very expensive person.\n\nIt also happens to be a surprisingly good way of thinking about **Jev**, a new AI model from [TypeSafe AI](https://typesafe.ai/). Except replace the kitchen with software.\n\nAnd replace “pasta” with things like: **retry, route, escalate, call this tool, don’t call this tool..*** That sounds almost too simple.*\n\nSo I went looking for the complicated part.\n\nHere’s something I started noticing when thinking about how AI gets wired into real applications.\n\nWe keep asking language models to make decisions that don’t actually need much language.\n\nWhen an application needs to make a basic decision like **routing a support ticket** or **deciding whether to run a database query developers wrap a massive, multi-billion-parameter LLM** in a giant prompt:\n\n*“You are a helpful assistant. Please analyze this message and return ONLY a JSON object with keys ‘department’ and ‘urgency’…”*\n\nThen we cross our fingers 🤞\n\nWe pay for dozens of generated tokens we don’t need, wait **3 to 300+ seconds** for autoregressive generation, and set up retry loops just in case the model hallucination inserts a stray trailing comma or a paragraph of apologetic prose.\n\nWe turned high-speed software into an anxious chatbot wrapper.\n\nFormer OpenAI researcher **Diogo Almeida** (who helped build the instruction-following methods behind ChatGPT) realized that if we want real software automation, we don’t need models that talk better but we need models that act like [**System One thinking**](https://docs.typesafe.ai/concepts/system-one).\n\nIn Daniel Kahneman’s framework, **System 1** refers to fast, intuitive judgment, while **System 2** refers to slower, deliberate reasoning. TypeSafe borrows that distinction as the inspiration for its “System One” model category.\n\nJev is built strictly for **System 1**. It takes **unstructured program state in** and returns **typed, probabilistic decisions out** with zero text generation.\n\nHow does Jev achieve response times between **70ms and 500ms** (up to 200x faster than traditional LLMs)?\n\n*It strips away autoregressive text generation entirely.*\n\nInstead of generating token N+1 conditioned sequentially on token N, Jev uses a **Parallel Sampler**.\n\nIt evaluates the entire input state and computes all required output probabilities simultaneously in a single hardware-aware pass.\n\nBecause output choices are constrained to schemas defined in advance, type errors are prevented by the predefined output schema, so the model **cannot return a value outside the allowed type** (0% schema mismatch rate). That guarantee is about **output type and structure**, not semantic correctness.\n\nAnd because it doesn’t generate tokens, **output tokens are FREE**. You only pay for input tokens at **$0.042 per million tokens**\n\nInstead of open-ended prompt engineering, Jev reduces all software decision logic into three explicit primitives:\n\n🟢 **1. Choice (Categorical Routing)**\n\nSelects one option from a predefined list and returns the exact probability distribution across all candidates.\n\n🟡 **2. Score (Ordered Rubrics)**\n\nRates state against a concrete, multi-tier scale with confidence metrics.\n\n🔵 **3. Noul (Binary Probabilistic Judgments)**\n\nEvaluates a specific yes/no statement and returns an exact probability score P(Yes) from 0 to 1.\n\nAt this point, I stopped trying to understand Jev only from diagrams and descriptions. I opened the [Playground](https://thejevai.com/playground).\n\nAnd I gave it a support message:\n\n```\nSubject: Charged twice again!!Hi! this is the SECOND month in a row I've been billed twice for the Pro plan.I already emailed last month and nobody replied.I run my whole business on this.If it's not refunded today I'm canceling and disputing the charge with my bank.\n```\n\nThen I asked three different questions about the **same piece of state**.\n\n```\n{  \"topic\": {    \"type\": \"choice\",    \"options\": [\"billing\", \"bug\", \"account\", \"feature\"],    \"instructions\": \"What is the primary issue the customer is writing about?\"  },  \"severity\": {    \"type\": \"score\",    \"rubric\": [      \"routine, no rush\",      \"should be handled today\",      \"urgent, customer is frustrated\",      \"critical, customer is about to churn or dispute\"    ],    \"instructions\": \"How urgent and high-risk is this message?\"  },  \"escalate\": {    \"type\": \"noul\",    \"instructions\": \"Should this be escalated to a human agent immediately?\"  }}\n```\n\nThe Playground returned:\n\n```\n{  \"topic\": {    \"type\": \"choice\",    \"choice\": \"billing\",    \"probabilities\": {      \"feature\": 0,      \"account\": 0,      \"bug\": 0,      \"billing\": 1    },    \"confidence\": 1  },  \"severity\": {    \"type\": \"score\",    \"score\": 3,    \"legend\": {      \"0\": \"routine, no rush\",      \"1\": \"should be handled today\",      \"2\": \"urgent, customer is frustrated\",      \"3\": \"critical, customer is about to churn or dispute\"    },    \"probabilities\": {      \"0\": 0,      \"1\": 0,      \"2\": 0,      \"3\": 1    },    \"confidence\": 1  },  \"escalate\": {    \"type\": \"noul\",    \"noul\": 0.92  }}\n```\n\nNotice how your application backend can now branch directly on typed outputs without parsing a single word of text.\n\nThe most interesting number in the response wasn’t the 1.0.\n\nIt was the 0.92.\n\nJev had decided: escalate → 0.92\n\nBut what exactly does **0.92** mean?\n\nA model can be right a lot of the time and still be terrible at knowing **when** it is likely to be wrong.\n\nImagine a model makes 100 decisions and gives all of them roughly 80% confidence.\n\nIf 80 of those decisions are correct, the number is doing something useful.\n\nIf only 55 are correct, then that 0.80 is mostly decorative.\n\nThis is the difference between **accuracy** and **calibration**.\n\n**Accuracy** → How often was the model right?\n\n**Calibration** → Does an 80% prediction actually behave like 80%?\n\nAnd this is where **RLCD**, or **Reinforcement Learning for Calibrated Decisions**, comes in.\n\nTypeSafe describes RLCD as its training approach for [System One models](https://docs.typesafe.ai/concepts/system-one), with the objective of producing probabilities that reflect uncertainty rather than simply optimizing for responses that humans prefer.\n\nIn its own description, **a model should not just know how to make a decision**; it should also communicate how likely that decision is to be correct. That matters enormously once the model is sitting inside code.\n\nFor example, an application might choose a policy like:\n\n*confidence > 0.90 → automate 0.60–0.90 → verify< 0.60 → ask a human*\n\nThose thresholds are an **application-level policy**, not something Jev decides for you.\n\nA read-only classification might tolerate a lower threshold.\n\nA payment, deletion, or other consequential action might require something much higher.\n\nThe current [Jev documentation](https://docs.typesafe.ai/introduction) makes exactly this distinction: probability and confidence are intended to be control signals, while the application defines the thresholds and decides when to **automate, verify, or escalate**.\n\nSo RLCD isn’t just a fancy name for “the model gives confidence scores.”\n\nThe useful idea is: **If software is going to act on a model’s judgment, uncertainty has to be part of the interface too.**\n\nAnd that makes Jev’s output a little more interesting than simply: **billing**\n\nIt becomes: **billing + how confident the system is**\n\nTo demonstrate what 70ms decision latency can unlock, TypeSafe showcases Jev across a few very different environments:\n\nLooking at this comparison, the biggest shift isn’t just about saving money on tokens or cutting latencies down to 70ms. It’s a fundamental software design principle:\n\nWhen building with traditional LLMs, developers accidentally hand the model total control. We ask GPT-4 to read an incident log, figure out what went wrong, write a explanation, format a JSON blob, and execute a tool call, hoping no step in that long chain fails or hallucinates.\n\nSystem One models enforce a much cleaner contract inside code:\n\nMaybe the interesting part isn’t making chatbots talk faster.\n\nIt’s figuring out which parts of an AI system actually need language in the first place.\n\nGordon Ramsay doesn’t need to cook every meal.\n\nHe just needs to point at the right stove.\n\n**Sometimes, AI doesn’t need to talk. It just needs to decide.**\n\nThe examples, technical details, and benchmarks discussed in this article are based primarily on TypeSafe AI and Jev’s own documentation and published material.\n\n[Jev AI: What If Gordon Ramsay Didn’t Cook? He Just Picked Who Did.](https://pub.towardsai.net/jev-ai-what-if-gordon-ramsay-didnt-cook-he-just-picked-who-did-c4ef6340ddbb) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/jev-ai-what-if-gordon-ramsay-didnt-cook-he-just-picked-who-did", "canonical_source": "https://pub.towardsai.net/jev-ai-what-if-gordon-ramsay-didnt-cook-he-just-picked-who-did-c4ef6340ddbb?source=rss----98111c9905da---4", "published_at": "2026-09-28 17:31:03+00:00", "updated_at": "2026-09-28 17:46:29.389541+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-tools", "ai-infrastructure"], "entities": ["TypeSafe AI", "Jev", "Diogo Almeida", "OpenAI", "ChatGPT", "Parallel Sampler", "System One", "Daniel Kahneman"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/jev-ai-what-if-gordon-ramsay-didnt-cook-he-just-picked-who-did", "markdown": "https://wpnews.pro/news/jev-ai-what-if-gordon-ramsay-didnt-cook-he-just-picked-who-did.md", "text": "https://wpnews.pro/news/jev-ai-what-if-gordon-ramsay-didnt-cook-he-just-picked-who-did.txt", "jsonld": "https://wpnews.pro/news/jev-ai-what-if-gordon-ramsay-didnt-cook-he-just-picked-who-did.jsonld"}}