# Cloudflare Just Shipped Two Decision Models. I Raced Them Against Jev (On OpenRouter)

> Source: <https://dev.to/therabbithole/cloudflare-just-shipped-two-decision-models-i-raced-them-against-jev-on-openrouter-3fdm>
> Published: 2026-10-08 11:09:41+00:00

A lot of AI calls these days don't really need to write anything. They just need to pick: which category, which passage, yes or no.

But most of us still use a chat model for that. We pay it to generate a paragraph, then parse out the one word we wanted.

Decision models skip the paragraph. You send the input and a list of allowed answers, and you get back a probability for every option. No text. One pass.

I've been using one of them, **Jev** from TypeSafe (`~typesafe/jev-latest` on OpenRouter, currently Jev 1.13), inside a small app that answers questions about long PDF manuals. Jev picks the best passages before the answer gets written.

Then on October 1 Cloudflare announced **Clef** in a post called *Introducing Clef: our open-source decision models, and new RL fine-tuning platform* ([blog.cloudflare.com](https://blog.cloudflare.com/clef-decision-models/)). Two models, actually:

Both are Apache 2.0, with the weights on Hugging Face ([huggingface.co/Cloudflare/clef](https://huggingface.co/Cloudflare/clef)), and both run on Workers AI ([developers.cloudflare.com](https://developers.cloudflare.com/changelog/post/2026-10-01-clef-workers-ai/)). Cloudflare calls Clef "fully Jev-API compatible", and both are on OpenRouter as `cloudflare/clef` and `cloudflare/clef-flash` ([openrouter.ai](https://openrouter.ai/cloudflare/clef)).

So I tested both on their own, outside my app, and compared them with Jev on accuracy and speed.

The announcement and the model card make three big claims:

All three are Cloudflare's own numbers, on Cloudflare's own hardware. Traictory pointed that out too: latency on your own infrastructure is "the easiest number for a vendor to control" ([traictory.com](https://traictory.com/news/2026-10-03-cloudflare-clef-decision-models)).

I wanted to see what you actually get when you call it through OpenRouter.

All three models sit behind the same OpenRouter endpoint, `POST /api/alpha/decisions` ([OpenRouter's Jev docs](https://openrouter.ai/docs/guides/community/jev)). You send a `state` (what the model looks at) and `questions` (what it has to decide):

```
{
  "model": "cloudflare/clef",
  "state": {"message": "Tracking hasn't updated in five days."},
  "questions": {
    "intent": {
      "type": "choice",
      "instructions": "Which intent best describes the customer's request in state?",
      "criteria": {
        "billing": "Charges, invoices, refunds, payment methods, pricing",
        "shipping": "Where a physical order is, delivery delays, returns of goods",
        "bug": "Something in the product is broken or behaving wrongly"
      }
    }
  }
}
```

Back comes the pick, a probability for every option, and a `confidence` number.

The nice part: switching models means changing one string.

Same answer from both. Note the confidence, though. We'll come back to it.

(The terminal exchanges in this post are reconstructed from the real session; the numbers are from the actual runs.)

I kept it small and boring on purpose. Three tasks:

To keep the timing fair:

I ran it twice. First Jev against Clef, then, a bit later, all three together. This is the second run, so all three models faced the same network at the same time:

The Errors column is Clef: in each task, one call failed with `HTTP 429` and "Capacity temporarily exceeded" from Workers AI, and it counts as wrong. Every Clef answer that actually came back was correct, except one.

That one was the message every model got wrong: "Is there a way to log in with Google?" I labelled it *feature*, all three said *account*. Honestly, I'd label it differently myself today.

So on answers, Clef ties with Jev. The first run said the same: 39/40, 30/30 and 12/12 for Clef, against 38/40, 30/30 and 12/12 for Jev.

Clef-flash is a different story.

Of its 5 wrong passages, 4 were the same mistake: **right kind of fact, wrong company.** Asked what Rowan Systems sells, it answered with what Rowan Labs sells. Exactly the near-twin trap the test was built to catch. The bigger models didn't fall for it once.

Now speed.

Jev was the fastest on every task, in both runs.

Clef was 1.7–2.4× slower on small requests and 3.4× slower on the long one, in both runs. Clef-flash landed in between: 1.2–1.3× slower than Jev on small requests, 2.2× on the long one.

That's the opposite of Cloudflare's chart. They measured Clef at 209 ms, Clef-flash at 38.8 ms and Jev at 524 ms. I never saw either Clef model beat Jev.

My numbers are the full round trip from my machine through OpenRouter, which is what your app sees. Theirs are on their own setup. Both can be true. Only one of them is what you get today.

(The second run was slower across the board, Jev included, so compare models within a run, not across runs.)

Clef's listing says **$0.24 per million input tokens and $0 for output**, the same as on Workers AI ([openrouter.ai](https://openrouter.ai/cloudflare/clef), [developers.cloudflare.com](https://developers.cloudflare.com/workers-ai/models/clef/)). Since a decision model doesn't write any output, that sounds cheap.

Clef-flash is listed at **$0.09** ([openrouter.ai](https://openrouter.ai/cloudflare/clef-flash)).

Neither is cheap next to Jev. Jev is listed at **$0.042 per million input tokens**, also with free output ([openrouter.ai](https://openrouter.ai/typesafe/jev-1.13)). The `usage` field in my responses matched all three prices. I'm not the only one who noticed: The Register and Hacker News called out the same 6× gap on launch week ([traictory.com](https://traictory.com/news/2026-10-03-cloudflare-clef-decision-models)).

The Clef models count fewer tokens for the same text, but pay more per token: about 6× for Clef and 2× for Clef-flash.

Per call, that made Clef 3.4–6× more expensive than Jev and Clef-flash 1.3–2.3× more.

To be fair: both runs together, 410 counted calls plus warm-ups, cost about 5 cents. You won't feel this at small scale. You will at a million calls a day.

This is the part that actually matters for my app.

Jev doesn't just pick. It says how sure it is, and I use that. If the top passage comes back with low confidence, that's a hint the manual may not answer the question at all.

Jev: 0.99 when right, 0.71 when wrong. A clear gap.

Clef: 0.81 when right, 0.74 when wrong. **Almost no gap.** And on the long task it got every answer it returned right while saying it was about 41% sure.

Clef-flash is interesting. On intents its gap was the biggest of all three: 0.81 vs. 0.43. But on the long task it flipped, and it was less sure when right than when wrong.

A confidence number that's the same whether you're right or wrong isn't a confidence number. It's noise.

That's odd, given Cloudflare trained it specifically for calibration. Maybe it shows on bigger and harder sets. On mine, it didn't.

So even with equal accuracy, Clef isn't a drop-in swap for anything that acts on the confidence. And Clef-flash's confidence only works some of the time.

`state` (
Clef is a real decision model. It picks as well as Jev on simple tasks, and open weights are a genuine plus.

But on OpenRouter today it's slower, it gets slower as the input grows, it costs more per call, it sometimes runs out of capacity, and its confidence doesn't tell you much.

Clef-flash is closer to Jev on speed and price. But it's still slower, still pricier, and it mixes up look-alike names.

For my app, Jev stays. I'll run the same test again in a month.
