Memo: The Sentence Tax TypeSafe AI emerged from roughly two years of stealth on 15 September with Jev, a "System One model" that returns structured values with calibrated probabilities instead of prose, backed by a $40 million seed led by DCVC. Within sixteen days, eight vendors shipped competing decision models — including Nandakishor Mukkunnoth's 421-million-parameter Laya under Apache 2.0 on 18 September, OpenAI's Decisions API preview on GPT-6 Luna at DevDay on 29 September claiming 150 milliseconds versus 1.6 seconds for a standard call, Cloudflare's Clef and Clef-flash on 1 October returning a decision in 38.8 milliseconds at the median, and Amazon's Strands Decider 2B the same day — with Cloudflare adopting Jev's own request format. Vercel added Jev to its AI Gateway on day two and reported it reaching nearly 13% of paid teams within 24 hours, roughly double the first-day share. Memo: The Sentence Tax Agents spend most of their model calls answering closed questions – route this, classify that, is this action allowed, did the last step work. Every one of those has been answered by asking a language model to compose a sentence, which your code then parses back into a boolean. In sixteen days, eight vendors shipped models that skip the sentence entirely and return typed probabilities instead. The obvious objection is that this is a classifier with better marketing, and it is partly right. The more important development is where the moat went: the weights are free, and the platforms giving them away are capturing the training loop instead. On 15 September, TypeSafe AI came out of roughly two years of stealth with Jev and a $40 million seed led by DCVC. Founded by Diogo Almeida – a co-author of the InstructGPT work behind ChatGPT – with Erik Gafni and Sasha Sheng, the company calls Jev the first "System One model," borrowing Kahneman's fast-intuitive versus slow-deliberate distinction. The model itself is named after the Jevons paradox. Jev takes program state plus a set of typed questions defined in code, and returns structured values with calibrated probabilities. No prose, no reasoning trace, no code. Every field is produced in one pass rather than one token at a time, which is where the speed claim comes from. TypeSafe trains it with a method it calls Reinforcement Learning for Calibrated Decisions. What happened next is the actual story. On 18 September, a developer in Kerala named Nandakishor Mukkunnoth released Laya, a 421-million-parameter open-weight version under Apache 2.0. At DevDay on 29 September, OpenAI previewed a Decisions API built on GPT-6 Luna, its own slide claiming 150 milliseconds against 1.6 seconds for a standard call. Liquid AI announced d1 the same week, claiming to be first to beat Jev on Hugging Face's Decision Index. On 1 October, Cloudflare released Clef and Clef-flash – 27B and 9B, Apache 2.0, built on Qwen bases, served on Workers AI, with Clef-flash returning a decision in 38.8 milliseconds at the median. Amazon shipped Strands Decider 2B the same day on a Qwen3.5-2B base with a pointer head that scores predefined options directly, publishing weights, training data, and scripts. Perplexity launched its own Decisions API on an open-weight 27B model at four cents per million input tokens. Together has tev1; the community has Kev 9B, CLM-8B, and dozens of variants. Eight vendors, sixteen days, one primitive. And Cloudflare adopted Jev's own request format, so code written for the startup's API runs against the incumbent's model with minor changes. The tax Consider what an agent actually does across one task. It routes a request to a tool. It classifies an input. It checks whether a proposed action is permitted before executing it. It decides whether the previous step succeeded. It ranks candidate retrievals. None of these is a writing task. Each has a closed set of allowed answers. Under the architecture every agent framework inherited, each one is answered by asking a language model to compose a sentence, which your code then parses back into a value. Ask a frontier model whether an email is spam and it spends compute writing a sentence in order to say "yes." The model was built to be eloquent. You needed it to be decisive. That mismatch has a price, and the shape of the price is what makes it structural rather than annoying: on chat models, output tokens cost several times what input tokens cost, and decisions dressed as sentences are mostly output. You are paying a premium rate for the part of the response you throw away. Call it the sentence tax . It is not the cost of intelligence. It is the cost of making intelligence talk. The independent evidence that the tax is real is better than the vendor claims. Vercel added Jev to its AI Gateway on day two and reported it reaching nearly 13% of paid teams within 24 hours, roughly double the first-day share of the GPT-5.6 family. A Vercel engineer measured a safety classifier running five to eighteen times faster than the model it replaced. The CTO of Bryo AI found Gemini slightly more accurate on email classification but ten to twenty times more expensive, and valued Jev as the only option returning a real probability. "Isn't this just a classifier?" Any engineer who shipped text classification before 2023 will ask this immediately, and the question deserves a straight answer rather than a dodge, because the honest answer is: substantially yes, and the differences are the whole argument. One reviewer of Cloudflare's release put the skeptical case precisely: self-hosting Clef buys you a classifier, not a general model, and a fine-tuned small encoder will often match it on one narrow task for a fraction of the hardware. That is correct, and any team with a stable, high-volume, well-labelled single task should benchmark exactly that before buying anything. What the new models add is not classification. It is four things around it. They take arbitrary schemas at call time, so you get a decision on a new question without assembling a training set. They score every option of every question in a single forward pass, so a dozen judgments cost one request. They return calibrated probabilities rather than a label, which is what makes an escalation policy possible. And they collapse the engineering distance between "we should classify this" and "it is classified" from a sprint to an afternoon. So the sharper framing is that decision models trade peak per-task efficiency for breadth and time-to-first-decision. A fine-tuned encoder wins the narrow, stable, high-volume case. The decision model wins the long tail of judgments nobody would ever staff a training project for – which, in an agentic system, is most of them. The consequence nobody priced: guardrails move into the hot path The latency numbers change where governance can live. Guardrails have historically sat outside the request – a slow compliance check, a human in the loop, a post-hoc audit – because a synchronous check cost a full model call. Cloudflare's framing of Clef as agent guardrails is the tell: an agent can ask "should I take this action?" in tens of milliseconds before calling a tool. At a 38.8-millisecond median you can run that check on every tool call without a user noticing. That is a different safety posture from the one most AI programmes are architected around. The control moves from a brake you tap occasionally to a membrane the agent passes through continuously. It is also, notably, the thing this publication has argued was missing from runtime containment – a check that can sit in front of every action rather than reviewing the transcript afterwards. It just became affordable. Commoditized – but watch where the moat moved The tempting read is that three infrastructure giants gave away a startup's category in sixteen days because the primitive has no defensibility. Half right, and the other half is more interesting. Cloudflare did not merely open-source a model. It shipped a reinforcement-learning service around it: AI Gateway records your real requests as a dataset, Workers AI generates answers with the base model, Containers score and replay them, a Trainer updates the weights, and Workers AI serves the result as your own model. Read that pipeline again. The base model is free. The loop that turns your traffic into a model tuned to your decisions is not, and it runs on one vendor's infrastructure end to end. So the Apache 2.0 licence on the base is real and the escape hatch is real, but it is narrower than it looks. Open weights are an exit only if you can operate them – and what you would actually want to take with you is not the base model but the fine-tuned one, along with the dataset that produced it. The practical portability of your tuned weights matters more than the licence tag on the base. That reframes where defensibility sits for everyone in this stack. Not the model. The labelled decision data from your own domain, the evaluation harness that tells you when the model is wrong, the schema design that turns fuzzy business logic into typed questions, and the escalation policy. The platforms can copy the primitive in two weeks. They cannot copy your traffic – but they can host the loop that learns from it, which is exactly the position Cloudflare has taken. What to do Recompute your agent cost model with two tiers. If your forecast assumes every agent step consumes frontier tokens, it overstates the marginal cost of autonomy substantially. Model closed-set decisions at the cheap tier and frontier pricing only for genuine reasoning and escalations. Deployments shelved as too expensive to run at scale deserve re-examination with that arithmetic. Count your closed-set calls before buying anything. Instrument traces and count how many model calls answer route, classify, allow or deny, succeeded or failed, rank. In most agent codebases written this year it will be the majority. That count is your addressable saving, and you can compute it without a vendor conversation. Benchmark against a fine-tuned encoder, not just against your current LLM. For your highest-volume single decision, the old approach may still win on cost per decision. Run all three – your current model, a decision model, and a small fine-tuned classifier – on the same task. Decide per decision type rather than per vendor. Treat the escalation threshold as the real engineering. In a two-tier stack, your cost curve and your error rate are both set by one number: the confidence level at which the cheap model defers to the expensive one. Too low and you pay for frontier calls you did not need; too high and you ship the cheap model's mistakes. That threshold, tuned against your own data and re-tuned as traffic shifts, is where the work now lives – and it is the one artefact in this stack no vendor can hand you. Keep your schemas and labelled examples portable, and ask who owns the tuned weights. The request format is already standardising, which means lock-in will not arrive through the API. It will arrive through the training loop. Before adopting a platform's fine-tuning service, establish what you can export: the dataset, the tuned weights, or neither. Bottom line A decision is not a sentence, and for three years every agent has been paying to have one written anyway. Eight vendors fixed that in sixteen days, which tells you the primitive was never hard – it was simply unasked for while the industry was busy making models eloquent. The category is already commoditised at the model layer, the request format is converging on a startup's API, and the open weights are genuinely open. The forward call: within two quarters, decision-model support becomes a default feature of the major agent frameworks rather than a vendor choice, and competition moves entirely to the training loop – whose service captures your traffic, tunes on it, and serves the result. The tell to watch is an independent, non-vendor benchmark for calibration and decision accuracy. Every performance claim in this category today comes from a company selling a model, including the comparisons that look unfavourable to incumbents. The first credible third-party evaluation will reorder the table, and it will also tell you whether the primitive is as good as the launch posts say. Until then, the safe conclusion is the cheap one: find out how many of your model calls are answering yes-or-no questions in prose, and stop paying for the prose. Sources: TypeSafe AI's Jev launch of 15 September 2026, the $40M seed led by DCVC, roughly two years in stealth, founders Diogo Almeida, Erik Gafni and Sasha Sheng, the "System One model" naming after Kahneman and the model's naming after the Jevons paradox, the Reinforcement Learning for Calibrated Decisions training method, the parallel single-pass sampler, and the 255-option cardinality limit per choice, per TypeSafe's launch post and coverage from InfoQ, Dealroom, Truefoundry, Emergent and Latent Space. Vercel's AI Gateway adoption figures nearly 13% of paid teams within 24 hours, about double the GPT-5.6 family's first-day share , engineer Pranit Sharma's five-to-eighteen-times speedup on a safety classifier, and Bryo AI CTO Nikhil Mudholkar's accuracy and cost comparison per InfoQ. Laya 421M parameters, Apache 2.0, Nandakishor Mukkunnoth, 18 September per Startup Fortune and Dealroom. OpenAI's Decisions API preview at DevDay on 29 September, built on GPT-6 Luna, with the 150-millisecond versus 1.6-second comparison from OpenAI's own slide, per Dealroom and Startup Fortune. Liquid AI's d1 and its Hugging Face Decision Index claim per Dealroom. Cloudflare's Clef and Clef-flash 1 October, 27B and 9B, Apache 2.0, Qwen bases, Workers AI, 38.8-millisecond median for Clef-flash, Jev-compatible request format and the AI Gateway / Workers AI / Containers / Trainer reinforcement-learning pipeline per Cloudflare's announcement and coverage from DataNorth, Flavio Copes, Remio and Startup Fortune. Amazon Strands Decider 2B 1 October, Qwen3.5-2B base, pointer head, Apache 2.0 with training data and scripts, around 72% on the public JevBench v19 set per Startup Fortune and Digital Applied. Perplexity's Decisions API on pplx-decider-v1-27b at $0.04 per million input tokens, Together's tev1, and the community variants Kev 9B and CLM-8B per Dealroom and Digital Applied. The skeptical reading that self-hosting Clef yields a classifier rather than a general model, and that a fine-tuned small encoder may match it on a narrow task for a fraction of the hardware, per DataNorth. Cross-references to prior Signal Memo coverage: the unit was never the token, nobody can test what they're shipping, the compensation boundary. What is original to this memo: the "sentence tax" framing, the direct engagement with the classifier objection and the breadth-versus-peak-efficiency answer to it, the reading of Cloudflare's training pipeline as the relocation of the moat from weights to the tuning loop, and the operator prescriptions including the encoder baseline test.