Stop Paying Your LLM to Flip Coins TypeSafe launched its System One model Jev on September 15, 2026 after two years in stealth, and Convai Innovations open-sourced the competing Laya model under Apache 2.0 three days later, offering agents typed, constrained answers instead of free-form JSON. Both models answer only within developer-defined options across three question types, eliminating JSON parsing and retry loops for single-label decisions. TypeSafe advises filtering large, noisy input before querying Jev, which it names as one of the model's failure modes. Here’s an uncomfortable exercise. Open your agent’s traces and count the LLM calls that return a single word: “billing,” “yes,” “high,” “tool 3.” In most agent systems, it’s a lot of them. We’ve been paying a reasoning engine to think for three seconds and emit a paragraph of JSON. Then we write a parser to pull one label back out. It’s like hiring a lawyer to flip a coin. A new class of model is built for exactly this gap. TypeSafe’s Jev launched on September 15, 2026 https://typesafe.ai/blog/introducing-system-one-models-and-jev after two years in stealth. Three days later, Convai Innovations open-sourced Laya https://flowtivity.ai/blog/laya-open-source-jev-alternative/ under Apache 2.0. Both belong to what TypeSafe calls “System One,” named after Daniel Kahneman’s https://www.explainx.ai/blog/typesafe-ai-jev-system-one-models-launch-2026 fast, intuitive mode of thinking. This isn’t another model review. It’s about the architecture that emerges when you give your agents two brains instead of one. A System One model doesn’t write. You give it two things: It gives back typed answers with probabilities. There are only three question types: The key property is that the model can only answer within the options you define https://typesafe.ai/blog/introducing-system-one-models-and-jev , so it can’t invent a tool that doesn’t exist. There’s no JSON to parse and no retry loop when parsing fails. Two flavours exist today: The mental model I use: System One is a fuzzy if-statement. It’s the branch in your code that needs semantic understanding but not creativity. Now the patterns. This is the pattern everything else sits on: The model gives you refund requested: 0.92. Your code then checks whether the order is refundable, whether the user is authorised, and whether the amount is under the auto-approve limit. The model never touches the side effect. A developer testing Laya for fraud detection learned this the hard way https://dev.to/sharifulislamsourav/laya-a-decision-making-system-one-model-41fb . Their first instinct was to ask the model whether to approve, review, or hold an order. That’s a policy decision disguised as a semantic one. Ask the model “does this look like an account takeover?” and let code decide what to do with the answer. Rule of thumb: if a wrong answer is expensive or irreversible, the model supplies evidence and code makes the call. System One models evaluate all questions in parallel, so asking ten costs roughly the same time as asking one. Stop making sequential round-trips. Here’s one request for an inbound ticket: questions = { "department": Choice "billing", "bug", "feature request" , "urgency": Score "not urgent", "somewhat", "urgent", "critical" , "refund": Noul "Is the customer asking for a refund?" , "churn risk": Noul "Is the customer threatening to leave?" , "pii present": Noul "Does the message contain card or account numbers?" ,}result = system one state=ticket, questions=questions illustrative The “speculative” part is the interesting bit. Ask the questions you might need, and let code discard the irrelevant answers. That’s cheaper than a second round-trip when a branch turns out to need them. This is the flip side of fan-out: don’t feed the whole world to every question. A page-level check gets the headings. A clause check gets the clause. Less input means lower latency and, more importantly, less noise. TypeSafe is unusually upfront about Jev’s failure modes, and large, noisy input is one of them https://www.analyticsvidhya.com/?p=257705 . Their advice is to filter first. Small, precise state is the biggest accuracy lever you control. This is where System One models go from “cheaper classifier” to “architectural component.” Agents with 40 tools spend a surprising share of their reasoning budget deciding which tool to call. Sometimes they call one that doesn’t fit. Put a Choice question in front of tool selection, and include an explicit none of these option. The community is already doing this. One project routes agent skill selection through confidence-aware Jev decisions https://github.com/v-modal/awesome-jev-tools/wiki so that weak matches get declined instead of guessed. Declining is the feature. An agent that says “I don’t have a tool for this” beats one that confidently calls the wrong API. LLM guardrails usually live at the front door and nowhere else, because each check costs a full model call. At System One prices, you can check every hop: input, tool arguments, tool output, and final response. LangChain already ships experimental middleware for model routing and tool-risk gating https://www.width.ai/post/what-is-jev-ai-typesafe . A typical question set per tool call: Remember the sandwich, though. A guardrail score is an input to your permission system, not a replacement for it. As one community directory puts it, a model’s judgment doesn’t establish safety and doesn’t replace the host application’s permission checks. Every turn, your orchestrator makes a meta-decision: which model tier, how much reasoning effort, and which tools to expose. That’s a Choice question. One open-source router does exactly this https://github.com/hellogumbo/awesome-jev . A single ~350ms Jev call picks the tier, effort, tools and skill. It runs behind a hard deadline with a regex fallback. Copy that part. A router that can hang is worse than no router. A related variant is triage before escalation . Before dumping a 500-line stack trace into a frontier model, ask a System One model https://github.com/ismaelsoilet/jev-harness/wiki whether it’s a missing dependency, a flaky network, or a real bug. The cheap cases never reach the expensive brain. Some decisions are genuinely multi-dimensional, like “is this candidate a fit?” or “is this vendor risky?” Instead of asking one mushy question, score each criterion independently with Score questions. Then aggregate with weights in code. You get two wins: This is the pattern that makes the whole thing economical: The idea is local-first: run Laya, and fall back to Jev https://medium.com/data-science-in-your-pocket/laya-vs-typesafe-jev-ai-8b5dd9ce0176 only when confidence drops below your threshold. Most traffic should stop at the first tier. Two warnings, learned from the benchmarks rather than the marketing: Run the same small question over thousands of rows, such as log lines, support tickets, or contract clauses. Then aggregate with plain SQL. Someone has already built a DuckDB integration that exposes Choice, Noul and Score as SQL table functions. Semantic WHERE clauses are a genuinely new primitive. For large batches, respect the limits. Hosted Jev publishes caps of 250,000 tokens per second and 1,200 requests per minute. That’s where a local Laya instance earns its keep. System One models are not magic. The last two years of agent design assumed one brain doing everything. Good systems never worked that way. They’re layered: fast reflexes at the edges, slow deliberation in the middle. System One models give us the reflex layer. My test for what moves there: If you can list the valid answers in advance, and a wrong answer is cheap to catch, it’s a System One question. Everything else stays with the LLM. Open your traces. Count the one-word answers. That’s your migration backlog. Are you running System One models in production? I’d love to hear which patterns held up and which didn’t. Drop a comment. Official and model docs Explainers and comparisons Hands-on and community experiments Stop Paying Your LLM to Flip Coins https://pub.towardsai.net/stop-paying-your-llm-to-flip-coins-331370280b4b was originally published in Towards AI https://pub.towardsai.net on Medium, where people are continuing the conversation by highlighting and responding to this story.