Cloudflare tries to outplay Jev with open-weight Clef models Cloudflare announced on Thursday the Clef family of open-weight decision models, Clef and Clef-flash, which it claims outperform TypeSafe's Jev on three of four TypeSafe benchmarks while also handling images and video. Clef is built on post-trained, frozen Qwen3.8-27B and Clef-flash on Qwen3.5-9B, is priced at $0.24 per million tokens versus Jev's $0.042/M, and is available hosted on Workers AI or downloadable from Hugging Face under Apache-2.0, though Cloudflare AI Platform group product manager Michelle Chen confirmed the training datasets are not public. Clef-flash requires a GPU with at least 41 GB of VRAM and Clef 85 GB, and the API is fully Jev compatible as a drop-in replacement. Cloudflare tries to outplay Jev with open-weight Clef models Source: The Register https://www.theregister.com Sure, it costs more, but it can handle images and video and it's available on Hugging Face if you have the hardware horsepower to run it locally Two weeks after the Jev model took the AI world by storm, Cloudflare has released its own pair of "Clef" decision models that it claims are smarter and faster than Jev, while also being open weight and runnable locally. The Clef family of decision models was announced by Cloudflare on Thursday, and consists of two models: Clef and Clef-flash, a slightly smaller and faster version of the model. For all intents and purposes, they work the same way as TypeSafe’s Jev, in that they can answer three types of bounded, structured questions: Yes/no, multiple choice, and rankings. It’s there that the bigger differences emerge, though, as Clef isn’t only built differently but, if Cloudflare’s benchmark https://www.machinebrief.com/glossary/benchmark claims hold up, also appears more capable than Jev and several other decision models in a number of tests. For starters, Clef has an LLM backbone. According to Cloudflare, Clef uses specially post-trained, frozen versions of Qwen3.8-27B and Qwen3.5-9B for Clef and Clef-flash, respectively, with the Qwen backbone performing a prefill-only pass during inference https://www.machinebrief.com/glossary/inference . Clef is still fast - faster than Jev, to be fair - and scores choices in parallel after that prefill-only pass. It’s not clear what Jev’s underlying architecture is, as TypeSafe has kept that a secret. As for its speed and capability, Clef moves fast. Cloudflare ran it against Jev and some other open decision models using the Jev Decision Index available on Hugging Face https://www.machinebrief.com/glossary/hugging-face , and the company’s own ranking suggests Clef is slightly slower than other open models, but more accurate, with Clef-flash just as accurate as most of the others, but far faster. To be fair to the competition, Cloudflare self-reported its own scores against the benchmark, and they have yet to be reproduced for ranking on the official Decision Index. Cloudflare also ran Clef against TypeSafe’s own benchmarks, and claimed it beat Jev in three out of four areas, only losing out on agent trace observability. Even if it were a bit slower or less accurate, Clef has another major leg up on Jev: It’s not limited to classifying text - it can also handle images and video. Additionally, Clef supports a 64k context window https://www.machinebrief.com/glossary/context-window . Jev can also handle up to 64k tokens across a request, although its state plus longest individual question is limited to 32k. Clef is available directly from Cloudflare hosted on Workers AI, which the company said makes the models even faster because “we’re able to take advantage of our GPUs at the edge, leading to low network latency and faster decisions.” For those that would prefer not to pay the token cost Clef costs $0.24 per million tokens - nearly six times the price of Jev at $0.042/M , Clef can also be downloaded from Hugging Face, and is open weight under the same Apache-2.0 terms as Qwen. While described as “open source” in the announcement, Cloudflare AI Platform group product manager Michelle Chen confirmed to The Register that its training https://www.machinebrief.com/glossary/training datasets aren’t public. As for whether your hardware can run it, Chen told us that Clef-flash will run on any GPU with at least 41 GB of VRAM, while Clef requires 85 GB of VRAM on a GPU for it to function. “This is assuming single concurrency and a 64k context window,” Chen added. Don’t worry about having to rebuild your decision model architecture for Clef either - its API is fully Jev compatible, allowing it to serve as a drop-in replacement if you want to give it a shot, either locally or using Cloudflare’s hosting option. ® Get AI news in your inbox Daily digest of what matters in AI.