Cloudflare Just Shipped Two Decision Models. I Raced Them Against Jev (On OpenRouter) A developer benchmarked Cloudflare's newly released open-source Clef decision models against TypeSafe's Jev via OpenRouter's decision endpoint, finding Clef tied Jev on accuracy (39/40, 30/30, 12/12) while Clef-flash made near-twin errors, confusing facts about similarly named companies. The tests also surfaced HTTP 429 'Capacity temporarily exceeded' errors from Workers AI on each Clef task, and the author notes Cloudflare's latency and accuracy claims come from its own hardware. A lot of AI calls these days don't really need to write anything. They just need to pick: which category, which passage, yes or no. But most of us still use a chat model for that. We pay it to generate a paragraph, then parse out the one word we wanted. Decision models skip the paragraph. You send the input and a list of allowed answers, and you get back a probability for every option. No text. One pass. I've been using one of them, Jev from TypeSafe ~typesafe/jev-latest on OpenRouter, currently Jev 1.13 , inside a small app that answers questions about long PDF manuals. Jev picks the best passages before the answer gets written. Then on October 1 Cloudflare announced Clef in a post called Introducing Clef: our open-source decision models, and new RL fine-tuning platform blog.cloudflare.com https://blog.cloudflare.com/clef-decision-models/ . Two models, actually: Both are Apache 2.0, with the weights on Hugging Face huggingface.co/Cloudflare/clef https://huggingface.co/Cloudflare/clef , and both run on Workers AI developers.cloudflare.com https://developers.cloudflare.com/changelog/post/2026-10-01-clef-workers-ai/ . Cloudflare calls Clef "fully Jev-API compatible", and both are on OpenRouter as cloudflare/clef and cloudflare/clef-flash openrouter.ai https://openrouter.ai/cloudflare/clef . So I tested both on their own, outside my app, and compared them with Jev on accuracy and speed. The announcement and the model card make three big claims: All three are Cloudflare's own numbers, on Cloudflare's own hardware. Traictory pointed that out too: latency on your own infrastructure is "the easiest number for a vendor to control" traictory.com https://traictory.com/news/2026-10-03-cloudflare-clef-decision-models . I wanted to see what you actually get when you call it through OpenRouter. All three models sit behind the same OpenRouter endpoint, POST /api/alpha/decisions OpenRouter's Jev docs https://openrouter.ai/docs/guides/community/jev . You send a state what the model looks at and questions what it has to decide : { "model": "cloudflare/clef", "state": {"message": "Tracking hasn't updated in five days."}, "questions": { "intent": { "type": "choice", "instructions": "Which intent best describes the customer's request in state?", "criteria": { "billing": "Charges, invoices, refunds, payment methods, pricing", "shipping": "Where a physical order is, delivery delays, returns of goods", "bug": "Something in the product is broken or behaving wrongly" } } } } Back comes the pick, a probability for every option, and a confidence number. The nice part: switching models means changing one string. Same answer from both. Note the confidence, though. We'll come back to it. The terminal exchanges in this post are reconstructed from the real session; the numbers are from the actual runs. I kept it small and boring on purpose. Three tasks: To keep the timing fair: I ran it twice. First Jev against Clef, then, a bit later, all three together. This is the second run, so all three models faced the same network at the same time: The Errors column is Clef: in each task, one call failed with HTTP 429 and "Capacity temporarily exceeded" from Workers AI, and it counts as wrong. Every Clef answer that actually came back was correct, except one. That one was the message every model got wrong: "Is there a way to log in with Google?" I labelled it feature , all three said account . Honestly, I'd label it differently myself today. So on answers, Clef ties with Jev. The first run said the same: 39/40, 30/30 and 12/12 for Clef, against 38/40, 30/30 and 12/12 for Jev. Clef-flash is a different story. Of its 5 wrong passages, 4 were the same mistake: right kind of fact, wrong company. Asked what Rowan Systems sells, it answered with what Rowan Labs sells. Exactly the near-twin trap the test was built to catch. The bigger models didn't fall for it once. Now speed. Jev was the fastest on every task, in both runs. Clef was 1.7–2.4× slower on small requests and 3.4× slower on the long one, in both runs. Clef-flash landed in between: 1.2–1.3× slower than Jev on small requests, 2.2× on the long one. That's the opposite of Cloudflare's chart. They measured Clef at 209 ms, Clef-flash at 38.8 ms and Jev at 524 ms. I never saw either Clef model beat Jev. My numbers are the full round trip from my machine through OpenRouter, which is what your app sees. Theirs are on their own setup. Both can be true. Only one of them is what you get today. The second run was slower across the board, Jev included, so compare models within a run, not across runs. Clef's listing says $0.24 per million input tokens and $0 for output , the same as on Workers AI openrouter.ai https://openrouter.ai/cloudflare/clef , developers.cloudflare.com https://developers.cloudflare.com/workers-ai/models/clef/ . Since a decision model doesn't write any output, that sounds cheap. Clef-flash is listed at $0.09 openrouter.ai https://openrouter.ai/cloudflare/clef-flash . Neither is cheap next to Jev. Jev is listed at $0.042 per million input tokens , also with free output openrouter.ai https://openrouter.ai/typesafe/jev-1.13 . The usage field in my responses matched all three prices. I'm not the only one who noticed: The Register and Hacker News called out the same 6× gap on launch week traictory.com https://traictory.com/news/2026-10-03-cloudflare-clef-decision-models . The Clef models count fewer tokens for the same text, but pay more per token: about 6× for Clef and 2× for Clef-flash. Per call, that made Clef 3.4–6× more expensive than Jev and Clef-flash 1.3–2.3× more. To be fair: both runs together, 410 counted calls plus warm-ups, cost about 5 cents. You won't feel this at small scale. You will at a million calls a day. This is the part that actually matters for my app. Jev doesn't just pick. It says how sure it is, and I use that. If the top passage comes back with low confidence, that's a hint the manual may not answer the question at all. Jev: 0.99 when right, 0.71 when wrong. A clear gap. Clef: 0.81 when right, 0.74 when wrong. Almost no gap. And on the long task it got every answer it returned right while saying it was about 41% sure. Clef-flash is interesting. On intents its gap was the biggest of all three: 0.81 vs. 0.43. But on the long task it flipped, and it was less sure when right than when wrong. A confidence number that's the same whether you're right or wrong isn't a confidence number. It's noise. That's odd, given Cloudflare trained it specifically for calibration. Maybe it shows on bigger and harder sets. On mine, it didn't. So even with equal accuracy, Clef isn't a drop-in swap for anything that acts on the confidence. And Clef-flash's confidence only works some of the time. state Clef is a real decision model. It picks as well as Jev on simple tasks, and open weights are a genuine plus. But on OpenRouter today it's slower, it gets slower as the input grows, it costs more per call, it sometimes runs out of capacity, and its confidence doesn't tell you much. Clef-flash is closer to Jev on speed and price. But it's still slower, still pricier, and it mixes up look-alike names. For my app, Jev stays. I'll run the same test again in a month.