{"slug": "stop-paying-your-llm-to-flip-coins", "title": "Stop Paying Your LLM to Flip Coins", "summary": "TypeSafe launched its System One model Jev on September 15, 2026 after two years in stealth, and Convai Innovations open-sourced the competing Laya model under Apache 2.0 three days later, offering agents typed, constrained answers instead of free-form JSON. Both models answer only within developer-defined options across three question types, eliminating JSON parsing and retry loops for single-label decisions. TypeSafe advises filtering large, noisy input before querying Jev, which it names as one of the model's failure modes.", "body_md": "Here’s an uncomfortable exercise. Open your agent’s traces and count the LLM calls that return a single word: “billing,” “yes,” “high,” “tool_3.”\n\nIn most agent systems, it’s a lot of them.\n\nWe’ve been paying a reasoning engine to think for three seconds and emit a paragraph of JSON. Then we write a parser to pull one label back out. It’s like hiring a lawyer to flip a coin.\n\nA new class of model is built for exactly this gap. TypeSafe’s **Jev** [launched on September 15, 2026](https://typesafe.ai/blog/introducing-system-one-models-and-jev) after two years in stealth. Three days later, Convai Innovations [open-sourced **Laya**](https://flowtivity.ai/blog/laya-open-source-jev-alternative/) under Apache 2.0. Both belong to what TypeSafe calls “System One,” [named after Daniel Kahneman’s](https://www.explainx.ai/blog/typesafe-ai-jev-system-one-models-launch-2026) fast, intuitive mode of thinking.\n\nThis isn’t another model review. It’s about the architecture that emerges when you give your agents two brains instead of one.\n\nA System One model doesn’t write. You give it two things:\n\nIt gives back typed answers with probabilities. There are only three question types:\n\nThe key property is that [the model can only answer within the options you define](https://typesafe.ai/blog/introducing-system-one-models-and-jev), so it can’t invent a tool that doesn’t exist. There’s no JSON to parse and no retry loop when parsing fails.\n\nTwo flavours exist today:\n\nThe mental model I use: **System One is a fuzzy if-statement.** It’s the branch in your code that needs semantic understanding but not creativity.\n\nNow the patterns.\n\nThis is the pattern everything else sits on:\n\nThe model gives you refund_requested: 0.92. Your code then checks whether the order is refundable, whether the user is authorised, and whether the amount is under the auto-approve limit. The model never touches the side effect.\n\nA developer testing Laya for fraud detection [learned this the hard way](https://dev.to/sharifulislamsourav/laya-a-decision-making-system-one-model-41fb). Their first instinct was to ask the model whether to *approve, review, or hold* an order. That’s a policy decision disguised as a semantic one. Ask the model *“does this look like an account takeover?”* and let code decide what to do with the answer.\n\n**Rule of thumb:** if a wrong answer is expensive or irreversible, the model supplies evidence and code makes the call.\n\nSystem One models evaluate all questions in parallel, so asking ten costs roughly the same time as asking one. Stop making sequential round-trips.\n\nHere’s one request for an inbound ticket:\n\n```\nquestions = {    \"department\":  Choice([\"billing\", \"bug\", \"feature_request\"]),    \"urgency\":     Score([\"not urgent\", \"somewhat\", \"urgent\", \"critical\"]),    \"refund\":      Noul(\"Is the customer asking for a refund?\"),    \"churn_risk\":  Noul(\"Is the customer threatening to leave?\"),    \"pii_present\": Noul(\"Does the message contain card or account numbers?\"),}result = system_one(state=ticket, questions=questions)  # illustrative\n```\n\nThe “speculative” part is the interesting bit. Ask the questions you *might* need, and let code discard the irrelevant answers. That’s cheaper than a second round-trip when a branch turns out to need them.\n\nThis is the flip side of fan-out: don’t feed the whole world to every question. A page-level check gets the headings. A clause check gets the clause.\n\nLess input means lower latency and, more importantly, less noise. TypeSafe is unusually upfront about Jev’s failure modes, and [large, noisy input is one of them](https://www.analyticsvidhya.com/?p=257705). Their advice is to filter first. Small, precise state is the biggest accuracy lever you control.\n\nThis is where System One models go from “cheaper classifier” to “architectural component.”\n\nAgents with 40 tools spend a surprising share of their reasoning budget deciding which tool to call. Sometimes they call one that doesn’t fit.\n\nPut a Choice question in front of tool selection, and include an explicit none_of_these option. The community is already doing this. One project [routes agent skill selection through confidence-aware Jev decisions](https://github.com/v-modal/awesome-jev-tools/wiki) so that weak matches get declined instead of guessed.\n\nDeclining is the feature. An agent that says “I don’t have a tool for this” beats one that confidently calls the wrong API.\n\nLLM guardrails usually live at the front door and nowhere else, because each check costs a full model call. At System One prices, you can check every hop: input, tool arguments, tool output, and final response.\n\nLangChain already ships [experimental middleware for model routing and tool-risk gating](https://www.width.ai/post/what-is-jev-ai-typesafe). A typical question set per tool call:\n\nRemember the sandwich, though. A guardrail score is an input to your permission system, not a replacement for it. As one community directory puts it, a model’s judgment doesn’t establish safety and doesn’t replace the host application’s permission checks.\n\nEvery turn, your orchestrator makes a meta-decision: which model tier, how much reasoning effort, and which tools to expose. That’s a Choice question.\n\nOne open-source router [does exactly this](https://github.com/hellogumbo/awesome-jev). A single ~350ms Jev call picks the tier, effort, tools and skill. It runs behind a hard deadline with a regex fallback. Copy that part. A router that can hang is worse than no router.\n\nA related variant is **triage before escalation**. Before dumping a 500-line stack trace into a frontier model, [ask a System One model](https://github.com/ismaelsoilet/jev-harness/wiki) whether it’s a missing dependency, a flaky network, or a real bug. The cheap cases never reach the expensive brain.\n\nSome decisions are genuinely multi-dimensional, like “is this candidate a fit?” or “is this vendor risky?” Instead of asking one mushy question, score each criterion independently with Score questions. Then aggregate with weights in code.\n\nYou get two wins:\n\nThis is the pattern that makes the whole thing economical:\n\nThe idea is local-first: [run Laya, and fall back to Jev](https://medium.com/data-science-in-your-pocket/laya-vs-typesafe-jev-ai-8b5dd9ce0176) only when confidence drops below your threshold. Most traffic should stop at the first tier.\n\nTwo warnings, learned from the benchmarks rather than the marketing:\n\nRun the same small question over thousands of rows, such as log lines, support tickets, or contract clauses. Then aggregate with plain SQL.\n\nSomeone has already built a DuckDB integration that exposes Choice, Noul and Score as SQL table functions. Semantic WHERE clauses are a genuinely new primitive.\n\nFor large batches, respect the limits. Hosted Jev publishes caps of 250,000 tokens per second and 1,200 requests per minute. That’s where a local Laya instance earns its keep.\n\nSystem One models are not magic.\n\nThe last two years of agent design assumed one brain doing everything. Good systems never worked that way. They’re layered: fast reflexes at the edges, slow deliberation in the middle.\n\nSystem One models give us the reflex layer. My test for what moves there:\n\n***If you can list the valid answers in advance, and a wrong answer is cheap to catch, it’s a System One question.***\n\nEverything else stays with the LLM.\n\nOpen your traces. Count the one-word answers. That’s your migration backlog.\n\n*Are you running System One models in production? I’d love to hear which patterns held up and which didn’t. Drop a comment.*\n\n**Official and model docs**\n\n**Explainers and comparisons**\n\n**Hands-on and community experiments**\n\n[Stop Paying Your LLM to Flip Coins](https://pub.towardsai.net/stop-paying-your-llm-to-flip-coins-331370280b4b) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/stop-paying-your-llm-to-flip-coins", "canonical_source": "https://pub.towardsai.net/stop-paying-your-llm-to-flip-coins-331370280b4b?source=rss----98111c9905da---4", "published_at": "2026-09-30 17:01:02+00:00", "updated_at": "2026-09-30 17:17:12.985816+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-tools", "ai-startups"], "entities": ["TypeSafe", "Jev", "Convai Innovations", "Laya", "Daniel Kahneman", "Apache 2.0"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/stop-paying-your-llm-to-flip-coins", "markdown": "https://wpnews.pro/news/stop-paying-your-llm-to-flip-coins.md", "text": "https://wpnews.pro/news/stop-paying-your-llm-to-flip-coins.txt", "jsonld": "https://wpnews.pro/news/stop-paying-your-llm-to-flip-coins.jsonld"}}