cd /news/ai-agents/cloudflare-clef-decision-models-for-… · home › topics › ai-agents › article
[ARTICLE · art-149061] src=byteiota.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Cloudflare Clef: Decision Models for Faster AI Agents

Cloudflare released Clef and Clef-flash, open-weight decision models that return typed, structured answers in 38.8ms and 209ms respectively, on October 1 under Apache 2.0 on Hugging Face and Cloudflare Workers AI. Clef-flash scores 98.76% on tool-call selection and Clef scores 94.20% macro F1 on BANKING77 versus Jev's 79.74%, but Jev wins judgment tasks with 80.97% on When2Call versus Clef's 72.37%, and Jev is cheaper at $0.042 per million tokens versus Clef-flash's $0.09 and Clef's $0.24. Cloudflare has not published training data or training scripts, making these open-weight rather than fully open-source models.

read5 min views1 publishedOct 11, 2026
Cloudflare Clef: Decision Models for Faster AI Agents
Image: Byteiota (auto-discovered)

Your AI agent is paying frontier model prices to answer “is this ticket urgent?” That yes/no question demands a full round trip through Claude or GPT — parsing free-form output, retrying when the response drifts, and burning tokens at $15 per million for what amounts to a boolean. Cloudflare dropped Clef and Clef-flash on October 1 to fix exactly that. They’re open-weight decision models that return typed, structured answers in 38 milliseconds. Jev, the current market leader, takes 524. That’s a 13x gap.

What Clef Actually Is #

Clef is not a chat model. It doesn’t generate text. It’s a decision model — a specialized architecture that takes a state (a support ticket, a JSON blob, an agent trace) and a set of structured questions, then returns typed answers with probability scores. True/false, single-choice, rating. That’s it.

Under the hood, both models use Qwen backbones. Clef is built on Qwen 3.8-27B (209ms median latency, stronger accuracy on judgment tasks). Clef-flash uses Qwen 3.5-9B (38.8ms, optimized for speed). The architecture is non-autoregressive — instead of generating tokens sequentially, Clef scores all valid choices in parallel during a prefill-only pass, then reads answers directly from internal representations. No decoding step. No parsing. No drift.

Both models are on Hugging Face under Apache 2.0. They run on Cloudflare Workers AI for managed deployment or self-hosted on your own hardware. One caveat worth flagging: Cloudflare has not published training data or training scripts. These are open-weight models, not fully open-source — a distinction that matters if reproducibility or auditability is a hard requirement.

Where Clef Wins — and Where It Doesn’t #

The benchmark story is useful precisely because it’s nuanced. Clef dominates on classification tasks. On BANKING77 intent recognition, Clef scores 94.20% macro F1 against Jev’s 79.74%. On tool-call selection for agent workflows, Clef-flash hits 98.76% — the best result in the field. Latency: 38.8ms (Clef-flash) vs 524ms (Jev).

But Jev wins on judgment tasks — the harder question of “should the agent act at all?” On the When2Call benchmark, Jev scores 80.97% versus Clef’s 72.37%. On GPQA Diamond, 78.3 to 48.0. On MMLU-Pro, 82.7 to 65.9. Reaching for Clef on safety and approval decisions because it’s faster is a mistake that surfaces in production error logs.

Model Latency Price/M tokens Open Weight When2Call BANKING77 F1
Clef-flash 38.8ms $0.09 Yes (Apache 2.0) — best
Clef 209ms $0.24 Yes (Apache 2.0) 72.37% 94.20%
Jev 524ms $0.042 No 80.97% 79.74%
Strands Decider 2B 115ms Free* Yes (full OSS) ~50% —

The Price Reality Check #

Clef-flash costs $0.09 per million tokens. Jev costs $0.042. Full Clef is $0.24 — nearly 6x Jev’s price. For 1 million classification decisions on 2,000-token inputs (2 billion tokens total), you’re looking at $84 on Jev versus $480 on Clef.

That looks bad until you factor in latency. If your agent is making routing decisions in a real-time user-facing flow, 38ms vs 524ms is the entire perceived UX. At high decision volume where latency compounds — live support queues, high-frequency agent orchestration, auto-scaling pipelines — the 13x speed advantage justifies the cost premium. Below roughly 100,000 decisions per day, a well-prompted GPT-4o-mini or a fine-tuned BERT classifier costs less and performs comparably.

Using It: The Workers AI API #

The integration is straightforward. One POST request, a JSON schema describing your questions, typed output back. Clef is also API-compatible with Jev, so existing Jev pipelines migrate without rewriting the integration layer:

curl https://api.cloudflare.com/client/v4/accounts/$ACCOUNT_ID/ai/run/@cf/cloudflare/clef \
  -X POST \
  -H "Authorization: Bearer $CF_TOKEN" \
  -d '{
    "model": "clef",
    "state": "Checkout has been failing for every customer for the last hour.",
    "questions": {
      "urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
      "team": {
        "type": "choice",
        "instructions": "Which team should handle this request?",
        "criteria": {
          "billing": "Payments, invoices, and refunds",
          "technical": "Outages, errors, and configuration",
          "sales": "Plans and upgrades"
        }
      }
    }
  }'

Clef adds two capabilities Jev lacks: a vision encoder for image classification, and a 64k context window versus Jev’s 32k. The RL fine-tuning service is also worth noting — Cloudflare captures your decision data via AI Gateway, retrains Clef on your specific domain, and lets you redeploy the fine-tuned version via Workers AI’s Bring Your Own Model feature.

Three Cases Where You Shouldn’t Use Clef #

Cloudflare’s announcement naturally emphasizes the wins. Here’s a more honest picture of when to skip it:

Low decision volume. Below roughly 100,000 routing decisions per day, the latency advantage doesn’t compound meaningfully in your cost model. A fine-tuned BERT classifier or a cheap prompted LLM performs well enough for a fraction of the token price.

Safety and approval decisions. The When2Call gap — 72.37% (Clef) vs 80.97% (Jev) — is not a vanity metric. It measures whether a model correctly declines to act in ambiguous situations. Routing “should we approve this transaction?” or “is this content safe to publish?” through Clef introduces meaningful risk.

Full reproducibility requirements. If you’re in healthcare, finance, or a regulated environment, the absence of published training data and training scripts means you cannot independently audit the model’s behavior. Amazon’s Strands Decider 2B publishes weights, training data, and scripts — though at roughly 50% accuracy on hard classification tasks, it’s effectively a coin flip for anything beyond simple routing.

Decision Models Are Infrastructure Now #

Cloudflare Clef is the right tool for classification-heavy agent workflows at scale where latency is a first-class constraint. The 13x speed advantage over Jev is real. So is the price premium and the judgment-task gap. The practical architecture splits responsibility: Clef-flash for routing and classification, Jev or a larger model for approval decisions and safety checks.

The broader signal is harder to ignore. TypeSafe AI’s Jev launched September 15. Cloudflare Clef launched October 1. Amazon Strands Decider followed shortly after. OpenAI has its own Decisions API. This is no longer an experimental category. Decision models are becoming standard infrastructure for production AI agents — the same way message queues and caches formalized in earlier infrastructure generations. Pick the right one for the right job.

── more in #ai-agents 4 stories · sorted by recency
── more on @cloudflare 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cloudflare-clef-deci…] indexed:0 read:5min 2026-10-11 · —