# Open Decision Models Compared: Clef, Decider 2B and Jev

> Source: <https://www.digitalapplied.com/blog/open-decision-models-compared-clef-decider-jev>
> Published: 2026-10-01 00:00:00+00:00

On October 1, 2026 two companies released open-weight decision models. Cloudflare published Clef and Clef-flash under Apache 2.0 and hosted them on Workers AI. Amazon’s Strands Labs published Strands Decider 2B, also Apache 2.0, with its training data and scripts. Both answer typed questions about a block of state with a probability for every allowed option, and neither generates text. The class began with Jev from TypeSafe, launched September 15 as an API in early access; smaller open releases such as Laya, Kev 9B and Tev1 followed within weeks. This page puts the October 1 releases beside Jev.

**Editorial note:** Prepared October 3 as an October 1, 2026 dispatch. Every figure comes from the vendor’s own announcement, model card or documentation as read on October 3. Benchmark and latency rows are vendor-run and labelled as such.

1. 01Three open models in one dayClef at 27B, Clef-flash at 9B and Decider at 2B, all Apache 2.0, all built on Qwen bases with the text-generation head removed or bypassed.
2. 02Clef adds images and 64KCloudflare’s models read images and video and take 64K of context, against Jev’s text-only 32K. No other model in this comparison reads images.
3. 03Decider is the fully open oneAmazon ships the weights, the training data, the scripts and the version history. Cloudflare ships weights and a card.
4. 04The benchmarks are Cloudflare’sEvery score comparing Clef to Jev was run by Cloudflare. Jev leads on two of the six rows shown, so read them as a vendor table, not a verdict.

## 01 — The classWhat these models do

A decision model takes a block of state, such as a support ticket or an agent’s proposed tool call, and a schema of questions with fixed answer types: yes or no, a choice from a list, a score on a scale. It returns a probability for each option in one forward pass. There is no generated text to parse and no way for the output to fall outside the schema. Amazon’s post describes the trade plainly: the model gives up open-ended output, and in return is faster and more capable at a given size, always answers from the offered options, and reports how sure it is. The same post is equally plain about the cost: a single parallel pass is much worse at hard problems than a reasoning model, and the lack of text makes the class useless for coding, chat or summarisation.

We explained the format and the first API routes in [decision models: AI that returns probabilities, not text](https://www.digitalapplied.com/blog/decision-models-probabilities-not-text-api). What changed on October 1 is that the class now has weights you can download, inspect and fine-tune.

## 02 — The ledgerThe comparison

Two tables, because five columns do not fit the page. The first covers what the model is; the second covers where it runs and what it costs. Jev is included as the closed reference the open models are measured against, and Together’s Tev1 as the earlier open experiment that took a different route.

| Sources: Cloudflare blog and Workers AI model pages, Hugging Face model cards for Cloudflare/clef, Cloudflare/clef-flash and StrandsAgents/strands-decider-2B-hobson-v19, the Strands announcement, Together’s tev1 repository and TypeSafe’s launch post. Read October 3, 2026. |  |  |  | 
|---|---|---|---|
| Model | Weights and licence | Size and base | Context and inputs | 
|---|---|---|---|
| Clef (Cloudflare) | Open, Apache 2.0, Hugging Face | 27.4B, post-trained from Qwen3.8-27B | 64K. Text, JSON, images and video | 
| Clef-flash (Cloudflare) | Open, Apache 2.0, Hugging Face | 9.4B, post-trained from Qwen3.5-9B | 64K. Text, JSON, images and video | 
| Strands Decider 2B (Amazon) | Open, Apache 2.0, Hugging Face and GitHub, with training data and scripts | 1.9B per the repository, Qwen3.5-2B-Base torso, LM head replaced by a pointer head | 4,096-token serving window in the repository. Text | 
| Tev1 4B experimental (Together) | Open weights; code MIT, weights and datasets under their own terms | 4B, supervised fine-tune of Qwen3.5-4B | 32K (32,768 tokens) on its OpenRouter listing. Text, 2 to 24 options | 
| Jev (TypeSafe), the closed reference | Closed. API only, early access | Not disclosed | 32K, per Cloudflare’s comparison. Text | 

| Prices from the Workers AI model pages, the OpenRouter listings and TypeSafe’s launch post. Latency figures are each vendor’s own measurement under its own conditions and are not comparable across rows. Read October 3, 2026. |  |  |  | 
|---|---|---|---|
| Model | Where it runs | Price | Vendor-reported latency | 
|---|---|---|---|
| Clef | Workers AI, or self-hosted | $0.24 per million input tokens on Workers AI | Median 209 ms, p95 239 ms (Cloudflare-run) | 
| Clef-flash | Workers AI, or self-hosted | $0.09 per million input tokens on Workers AI | Median 39 ms, p95 122 ms (Cloudflare-run) | 
| Strands Decider 2B | Local CPU or GPU via the strands-decider CLI | Free; your hardware | Median about 115 ms on an RTX 3090, about 153 ms on an M3 MacBook for small tasks (Amazon-run) | 
| Tev1 4B experimental | Together endpoint, OpenRouter, or self-hosted | $0.042 per million input, free output, on OpenRouter | Not published | 
| Jev | TypeSafe API, OpenRouter route | $0.042 per million input, output free (TypeSafe) | 70 to 500 ms end to end (TypeSafe); median 524 ms (Cloudflare-run) | 

Three rows need a note. Cloudflare’s 64K context figure comes from its blog, and the Workers AI page lists 65,536 tokens; the two agree. Decider’s repository gives a 4,096-token serving window (its JevBench run was preregistered at 3,072) rather than a model context limit, and Amazon’s post says latency grows roughly linearly with task size. Tev1 keeps Qwen’s language-model head and returns one answer letter, with token log-probabilities, rather than scoring every option through a dedicated head, which puts it at the edge of the class; Together labels it experimental.

## 03 — Vendor-runBenchmarks, as the vendor ran them

Cloudflare published a table of ten benchmarks across six models, all scored by Cloudflare. Six rows are reproduced below for the three models this page is about. Cloudflare marks Clef as the leader on the [Jev Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index), a Hugging Face space maintained by a third party, and says Clef beat Jev on three of TypeSafe’s own four workflow evaluations. We found no independent replication by October 3.

| Source: Cloudflare, “Clef decision models” blog post, October 1, 2026. Every score in every column is Cloudflare’s own run, including the Jev column. Higher is better on every row. |  |  |  | 
|---|---|---|---|
| Benchmark | Clef | Clef-flash | Jev | 
|---|---|---|---|
| BFCL, case exact | 98.47 | 98.76 | 95.75 | 
| API-Bank, accuracy | 91.93 | 93.11 | 88.19 | 
| BANKING77, macro-F1 | 94.20 | 90.93 | 79.74 | 
| When2Call, accuracy | 72.37 | 65.58 | 80.97 | 
| BRIGHT, nDCG@10 | 45.91 | 39.26 | 47.52 | 
| PhishNChips, accuracy | 79.60 | 75.05 | 62.55 | 

#### Median decision latency, milliseconds, lower is better

Cloudflare’s own measurement across its 43-benchmark run, published October 1, 2026. Laya (5.8 ms median, fastest in Cloudflare’s run) is omitted; Cloudflare says it trades off quality. Decider 2B is excluded because Amazon measured it on different hardware; Amazon reports about 115 ms on an RTX 3090.
Amazon publishes less. Its post reports that Decider 2B ranks third of 33 models in the 2B class on the public JevBench set, and first of 30 once models just over 2B are excluded, citing a third-party leaderboard, and that it answers all of JevBench’s easy tasks correctly. It gives a Brier score trend across versions rather than a single calibration number. As with Cloudflare’s table, these are the vendor’s figures. Our [audit of what vendor benchmark rows disclose](https://www.digitalapplied.com/blog/vendor-benchmark-reproducibility-audit-2026) is the checklist for reading them.

## 04 — What changedWhat open weights change

Three things, in order of how soon a team will use them. First, the model can run where the data is. Decider at 2B runs on a laptop CPU; Clef-flash at 9B runs on one consumer GPU; Clef at 27B needs server hardware or Workers AI. A guardrail that must see every tool call before it executes can now sit in the process rather than across a network hop, which is the use Amazon demonstrates with its Strands intervention example.

Second, the model can be fine-tuned on the decisions a company has already made. Cloudflare says its internal teams want Clef tuned on years of labelled trust-and-safety and support decisions, and it is offering that as a service first and a self-serve product later. Amazon’s release includes the training recipe, so the same thing can be done without a vendor.

Third, the format is converging. Cloudflare built Clef to be compatible with Jev’s API shape, and Decider’s CLI uses the same three question types: a yes-or-no, a choice and a score. A team that writes its questions once can swap the model underneath, which is the condition our [second-source playbook](https://www.digitalapplied.com/blog/ai-vendor-resilience-open-weight-second-source-2026-playbook) asks for before any model goes into a hot path.

Jev remains the reference and remains an API. TypeSafe has not published weights, a parameter count or a training recipe, and its own launch post concedes it cannot prove its pricing is not subsidised and that only the long term will show it is sustainable. Its documented choice cardinality of up to 255 options is matched by Clef’s Workers AI schema, which accepts 2 to 255 options per choice.

## 05 — Practical implicationsWhich one to try first

Whichever model is chosen, the evaluation is the same: a few hundred decisions your team has already made, scored for accuracy and for whether the probabilities mean what they say. Our [AI transformation](https://www.digitalapplied.com/services/ai-transformation) work builds that evaluation set first, because a decision model is only as good as the decisions it is checked against.

## 06 — MethodMethod and as-of date

A vendor-page comparison, not a test. Nothing on this page was run by Digital Applied.

- What was collected
- Licence, parameter count, base model, context window, accepted inputs, hosting, price and vendor-reported latency for the October 1 open-weight releases and Together’s Tev1, plus Jev as the closed reference.
- Sources
- Cloudflare’s announcement and Workers AI model pages; the Hugging Face API records for each model, which supply the licence tag, base-model tag and safetensors parameter count; Amazon’s Strands announcement and the strands-decider repository; Together’s tev1 repository and OpenRouter listing; TypeSafe’s launch post.
- As-of date
- October 3, 2026. Hugging Face records show the Cloudflare cards created September 30 and last modified October 1, and the Strands card created September 30 and modified October 1; the vendor announcements are dated October 1.
- Units
- Parameter counts are safetensors totals from the Hugging Face API, rounded to one decimal in billions. Prices are per million input tokens as the vendor lists them. Latency is milliseconds as the vendor reports it, with the measuring vendor named.
- Exclusions
- Kev 9B (Apache 2.0, Hugging Face, September 20) and Laya (Apache 2.0, Hugging Face, September 18) are independent open-weight decision models outside this vendor comparison, and DiffusionGemma Jev has no model page linked from Cloudflare’s table. Kev 9B and DiffusionGemma Jev appear in the latency chart; Laya does not. Closed API routes other than Jev are covered in the September 28 post.
- Limitations
- No benchmark was reproduced. Cloudflare’s Jev scores are Cloudflare’s runs against a competitor. Amazon’s ranking cites a third-party leaderboard whose method we did not audit.
- Refresh
- Add a row when a vendor publishes weights for a decision model. Re-read the Workers AI price pages monthly. Correct any figure in place with a dated note rather than silently editing it.

### Download one, and score it on decisions you already made

The open releases make the class testable without a waitlist. Pick the row that matches the state you need to judge, run it on a few hundred past decisions, and keep the vendor’s benchmark table for context rather than as the answer.
