# Clef 27B: Cloudflare's Multimodal Decision Model Explained and Tested

> Source: <https://www.mindstudio.ai/blog/clef-27b-multimodal-capabilities-test/>
> Published: 2026-10-02 00:00:00+00:00

# Clef 27B: Cloudflare's Multimodal Decision Model Explained and Tested

Cloudflare's Clef 27B reads images, video and charts to return calibrated probabilities instead of text. Here's how it works and what it got right.

## What is Clef 27B?

Clef 27B is a 27 billion parameter decision model from Cloudflare, released under the Apache 2.0 license. It doesn’t generate text. Instead, you feed it a situation (text, image, video frames, or JSON), give it a question with a defined set of possible answers, and it returns calibrated probabilities for each option in a single forward pass. There’s no token-by-token generation, no waiting for a response to stream out. It scores every option at once and hands back a probability distribution.

That makes Clef part of a smaller, less-discussed category of AI models built specifically for structured decision-making rather than conversation. What sets it apart from other decision models in that category is that it’s multimodal: it can look at an image, watch a short video, or read a diagram, and still answer in the same probability-driven format.

## TL;DR

- **Clef 27B is a decision model, not a chatbot.** It takes a situation and a set of possible answers, then returns a probability for each answer in one pass, with no text generation involved.
- **Multimodality is the headline feature.** Clef can process images, video frames, and JSON alongside text, which separates it from earlier decision models limited to text-only inputs.
- **It caught a social engineering prompt in testing.** Given a fabricated “emergency access” request from a supposed manager, Clef flagged a high probability of social engineering and recommended escalation rather than granting access.
- **It read emotional and narrative cues from a single still image.** In a relationship-dynamics test, it assigned confidence scores to who was being genuine, who wasn’t, and how the situation was likely to play out.
- **It interpreted an 8-second AI-generated video with no dialogue.** It inferred probable meaning (a ritual, not danger) from visual context alone.
- **It handled non-English script without translation.** Given a screenshot of an Uzbek-language news site, it identified the language, the topic, and the tone with high confidence.
- **It read a university-level chemistry diagram accurately.** On a titration curve, it correctly identified the acid type, equivalence points, and other details with near-perfect confidence scores.

## Other agents ship a demo. Remy ships an app.

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

## How does a decision model like Clef actually work?

A decision model skips the generative step entirely. Where a chatbot predicts the next word over and over until it produces a sentence, a decision model is structured around a fixed set of possible outcomes. You define the question (“which team should this ticket go to?”) and the options (“technical,” “billing,” “sales”), and the model scores each option’s likelihood based on the input it’s given.

This has two practical effects. First, it’s fast: no generation overhead means no waiting for tokens to stream. Second, it’s structured: the output is a clean probability distribution rather than free text that needs parsing. In one demonstration, a support ticket describing an API outage was run through Clef with two separate questions, “which team should handle this” and “is this urgent,” asked at the same time. Clef returned answers for both in parallel within a single forward pass: a technical team routing at 100% confidence and urgency marked “yes” at 100%. That parallel, multi-question handling in one pass is the core mechanical difference between a decision model and a conversational one.

## What makes Clef’s multimodality different from other decision models?

Most decision models in this category have been text-only: feed them a scenario description, get back probabilities. Clef extends the same architecture to images, video, and JSON, while keeping the output format identical. That means the same probability-based answers you’d get from a text scenario are now available for a photo, a video clip, a screenshot of foreign-language text, or a scientific diagram.

In testing, this showed up across several distinct scenarios:

- **A relationship-dynamics test using a single image.** Given a photo depicting a woman between two men (an infatuated boss and a younger colleague), Clef was asked to assess the emotional dynamics and predict outcomes. It returned scores like 94.5% confidence that the woman was primarily stressed about work rather than the men, 92.7% confidence the boss wasn’t genuinely in love, and a prediction that the situation would end messily but that she’d come out of it intact.
- **An 8-second AI-generated video with no dialogue** , showing a man alone in a frozen forest holding a candle. Clef read the scene as a ritual (72.2% confidence) rather than a memorial, assigned only a 16% probability to danger, and predicted the man survives.
- **A screenshot of a Uzbek-language news site.** Clef identified the language itself at 95.9% confidence, picked out a specific headline about the Ukraine conflict at 97.8% confidence, correctly called the tone neutral (93.1%), and classified the main topic as politics, all without translation.
- **A diprotic acid titration curve** , a university-level chemistry diagram. Clef correctly identified the diagram type (99.1%), the acid type (97.8%), and the number of equivalence points (96.7%).

Across these tests, the model wasn’t just labeling images, it was answering specific, structured questions about them and attaching a confidence score to each answer.

## Is Clef 27B good at catching risky or adversarial inputs?

Decision models are frequently pitched for security and policy-enforcement use cases, where you want a fast, auditable probability rather than a conversational judgment call. In one test, Clef was given a scripted social engineering scenario: an “employee” requesting emergency access, citing an unverifiable manager approval. Clef assigned only a 5.9% probability to granting access and a 76% probability that the request was social engineering, recommending escalation to security.

This result is consistent with how the model performed against other decision models tested previously by the same reviewer, where larger models tended to catch the manipulation attempt while smaller ones granted access. The pattern suggests that, at least in these scenarios, model scale correlates with better judgment on adversarial inputs, though that’s a trend observed across a handful of tests rather than a guarantee.

## How does Clef compare to other decision models on benchmarks?

Benchmark comparisons covering tasks like tool retrieval, API routing, and clinical classification showed Clef ahead of competing decision models in most categories, including a model previously treated as a strong baseline in this space. The one category where the comparison model held an edge was deciding when to escalate a situation to a human reviewer, arguably the most subjective and judgment-heavy task in the set. A smaller decision model in the same comparison consistently scored at the bottom, reinforcing the idea that it’s well suited to lightweight use cases but not a direct competitor to larger models like Clef.

## Is Clef 27B worth using?

For anyone building systems that need fast, structured, auditable decisions rather than open-ended text, Clef is worth a look. It’s open source under Apache 2.0, which removes licensing friction for commercial use, and its no-generation design means lower latency for routing, classification, and triage tasks. The multimodal capability is the real differentiator: being able to feed it a photo, a video clip, or a diagram and get the same structured probability output as you would from a text prompt opens it up to use cases well beyond typical text classification, including visual inspection, content moderation, and document triage involving charts or foreign-language text.

The tradeoff is that it’s a 27 billion parameter model, large enough to require real GPU resources rather than running comfortably on a laptop. For teams that need decision-grade outputs at scale and already have the infrastructure to run a model of this size, it looks like one of the more capable options currently available in the open-source decision model space.

## Frequently Asked Questions

### What is a decision model, and how is it different from a chatbot?

A decision model takes a defined scenario and a set of possible answers, then returns a probability for each answer in a single pass, without generating free text. A chatbot predicts text token by token to produce an open-ended response. Decision models are built for speed and structured output, not conversation.

### Can Clef 27B process video, or just still images?

Yes. Clef can process video by analyzing frames, as demonstrated with an 8-second clip where it inferred emotional context and likely outcomes without any dialogue or text description provided.

### Does Clef 27B need translation to understand non-English text?

No. In testing, it correctly identified the language, topic, and tone of an Uzbek-language news screenshot without any translation step, reading directly from the image.

### Is Clef 27B open source?

Yes, it’s released under the Apache 2.0 license, which allows commercial use and modification without the licensing restrictions that come with some other model releases.

### What are good use cases for a model like Clef 27B?

Based on demonstrated tests, it suits tasks like security and access-request screening, support ticket routing, visual or video-based risk assessment, multilingual content classification, and reading structured diagrams such as charts or scientific graphs.
