How to Run Cloudflare's Clef 27B Decision Model Locally Cloudflare released Clef 27B, a 27 billion parameter multimodal decision model under the Apache 2.0 license that returns calibrated probabilities for a fixed set of answer options in a single forward pass rather than generating text. Local testing measured VRAM usage under 54GB for the fully loaded model including KV cache, and in a staged social engineering scenario Clef returned a 5.9% probability of granting access and a 76% probability of social engineering with a recommendation to escalate to security. Benchmark results favored Clef across most categories including tool retrieval, API routing, and clinical classification, though escalation-to-human decisions still favored the comparison model referred to as Jev. How to Run Cloudflare's Clef 27B Decision Model Locally A practical look at Cloudflare's open-source Clef 27B multimodal decision model: what it does, VRAM needs, and how it performs in local tests. What is Clef 27B? Clef is a 27 billion parameter multimodal decision model released by Cloudflare under the Apache 2.0 license. It is not a chatbot and does not generate text. Instead, you feed it a situation as text, an image, a video, or JSON , give it a typed question with a fixed set of possible answers, and it returns calibrated probabilities for each option in a single forward pass. There is no token-by-token generation, which means no streaming delay and no risk of the model wandering off script with a rambling answer. This puts Clef in the same category as other “decision models” like Cloudflare’s earlier releases often compared against models referred to informally as CLM, Jev, Kev, and Lia in community testing . What separates Clef from that group is multimodality: it can look at an image or watch video frames and reason about them the same way it reasons about text, still returning structured probability outputs rather than prose. TL;DR - Clef 27B is a decision model, not a chatbot. It takes a situation and a question with defined answer options, then outputs calibrated probabilities for each option instead of generating free text. - It runs multimodal input natively. Text, images, video frames, and JSON can all feed into the same decision pipeline, with no separate captioning or OCR step required. - Local testing showed VRAM usage under 54GB for the fully loaded 27 billion parameter model with KV cache included, putting it within reach of a single high-memory GPU rather than a multi-GPU rig. - It caught a social engineering prompt injection in a staged access-request scenario, correctly flagging low probability of granting access and high probability of escalation, matching the behavior of larger, more cautious models in similar tests. - Benchmark results favor Clef across most categories including tool retrieval, API routing, and clinical classification, with one notable exception: deciding when to escalate a situation to a human still favors the comparison model referred to as Jev. - Multilingual and cross-domain image reading worked well , correctly parsing an Uzbek-language news screenshot and a university-level chemistry titration curve with high confidence scores. - Apache 2.0 licensing means full open access to weights and usage with no restrictive terms, which matters for anyone building decision infrastructure they want to self-host. How does a decision model differ from a chatbot? A chatbot like most general-purpose LLMs generates an open-ended text response one token at a time. You ask a question, it writes an answer, and you parse that answer yourself if you need structured data out of it. A decision model works backward from that. You predefine the question and the finite set of possible answers up front. The model’s job is to score each option with a calibrated probability and hand that back in one pass, with no generation step. This is useful anywhere you need fast, structured, auditable decisions rather than conversational text: fraud flags, access control, triage, routing a support ticket to the right team, or classifying a document. In testing, Clef was given a staged social engineering scenario: an “employee” message requesting emergency access approval that is unverifiable and bypasses normal process. Clef returned a 5.9% chance of granting access and a 76% probability of social engineering, with a recommendation to escalate to security. That mirrors the behavior of larger, more cautious decision models in the same test, while smaller models in earlier comparisons reportedly got fooled and leaned toward granting access. Clef can also answer multiple unrelated questions about the same input simultaneously. In a demonstrated example, a support ticket describing an API outage was fed into the model alongside two separate questions which team should handle it, and how urgent is it . Both answers came back in the same forward pass: a technical team assignment at 100% confidence and “urgent: yes” at 100% confidence. That parallel answering is the practical payoff of skipping text generation entirely. What does running Clef 27B locally actually require? Clef is a 27 billion parameter model, which puts it well above lightweight local models but still within single-GPU territory if you have enough VRAM. In a local run with the model fully loaded including KV cache overhead, VRAM consumption stayed under 54GB. That lines up with what you’d expect from a dense 27B model at typical precision, and it means a single high-memory GPU the kind available from most GPU rental providers, or a workstation-class card with 48GB or more can handle it without sharding across multiple devices. Because Clef does not generate long sequences of output tokens, the inference overhead per query is low once the model is loaded. The expensive part is holding the full 27B parameters and the vision/video encoding pipeline in memory, not the decoding loop. If you’re planning capacity, treat Clef similarly to other dense 27B-class open models: budget for a single GPU in the 48-80GB VRAM range depending on quantization choices, and expect fast turnaround per query since there’s no multi-token generation to wait on. How well does it handle images and video? Multimodality is the feature that separates Clef from earlier text-only decision models, and it held up across several unscripted tests. Everyone else built a construction worker. We built the contractor. One file at a time. UI, API, database, deploy. Given a single AI-generated image depicting a workplace relationship triangle, Clef was asked to infer emotional dynamics and predict an outcome. It returned structured probabilities: 94.5% confidence that the central figure was stressed about work rather than romance, 92.7% confidence that one party “wasn’t in love” and was using the relationship, and a prediction that the situation ends messily but she comes out fine. That’s a reasonably coherent read from one still frame with no additional context. With an 8-second AI-generated video clip of a man alone in a snowy forest holding a candle, Clef was asked to assess his emotional state, danger level, and likely outcome. It read the scene as a ritual 72.2% confidence , identified the candle as a likely memorial object 60.2% , estimated only a 16% chance of danger, and predicted he survives. For a short, dialogue-free clip, that’s a fairly grounded interpretation rather than a guess. Clef also handled a screenshot of an Uzbek-language news site, correctly identifying the language 95.9% confidence , the main topic as a geopolitical conflict reference in the page’s ticker 97.8% , a neutral tone 93.1% , and the general subject as politics. It underperformed only on a specific acronym reference, scoring that interpretation at 39.6%, described as the one clear miss in that test. On a more technical image, a diprotic acid titration curve, Clef correctly identified the diagram type 99.1% , the acid type 97.8% , and the number of equivalence points 96.7% , reading quantitative detail off a scientific chart rather than just recognizing it as “a graph.” Is Clef 27B worth running over other decision models? Based on the benchmark comparisons referenced during testing, Clef leads in most categories, including tool retrieval, API-style routing decisions, and clinical classification tasks, often by a wide margin over smaller models in the same comparison set. One category went the other way: deciding when to escalate a situation to a human judgment call, where the comparison model called Jev reportedly still held the edge. That’s a notable exception precisely because knowing when not to trust an automated decision is arguably the hardest judgment call on the list. Whether it’s “worth it” depends on what you’re building. If your use case is single-modality and text-only, a smaller, cheaper decision model may be sufficient and easier to host. If you need a single pipeline that can ingest screenshots, video frames, or multilingual documents alongside text and return structured, parallel answers, Clef’s multimodal support and Apache 2.0 licensing make it a reasonable default to evaluate first. Frequently Asked Questions What makes Clef different from a typical LLM chatbot? Clef doesn’t generate free-form text. You give it a situation and a predefined question with fixed answer options, and it returns calibrated probabilities for each option in a single pass, with no token-by-token generation. How much VRAM does Clef 27B need to run locally? In testing, the fully loaded 27 billion parameter model consumed under 54GB of VRAM including KV cache, which fits on a single high-memory GPU without needing multi-GPU sharding. Can Clef actually understand images and video, or just text? It processes images and video frames natively alongside text and JSON input, and testing showed it correctly interpreting photos, short video clips, foreign-language screenshots, and technical diagrams. Is Clef 27B free to use commercially? Yes. It’s released under the Apache 2.0 license, which is fully open source and permits commercial use without the restrictive terms some other open-weight models carry. Remy is new. The platform isn't. Remy is the latest expression of years of platform work. Not a hastily wrapped LLM. Does Clef outperform other decision models on every task? Not quite. It led most benchmark categories tested, including tool retrieval and clinical classification, but one comparison model still performed better on deciding when to escalate a situation to human review.