# Cloudflare puts audio and video into its decision model with Clef-omni

> Source: <https://runtimewire.com/article/cloudflare-clef-omni-multimodal-decision-model>
> Published: 2026-10-09 18:54:58+00:00

# Cloudflare puts audio and video into its decision model with Clef-omni

**The model takes text, images, audio and video in one request; Cloudflare also cut Clef-flash pricing while reducing its hosted context window.**

        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
        · Published 

Primary source: [Cloudflare Blog](https://blog.cloudflare.com/clef-faster-cheaper-multimodal/)

## Why it matters

Cloudflare is making schema-bound decisions work across more types of input while competing on hosted price. The tradeoff is visible: Clef-flash costs less, but its hosted context window is now smaller.

Cloudflare engineers [Michelle Chen, Alex Reneau and Kevin Jain](https://blog.cloudflare.com/clef-faster-cheaper-multimodal/?ref=runtimewire) have extended the company's Clef decision-model line to accept audio and video alongside text and images. [Clef-omni](https://runtimewire.com/models/huggingface/cloudflare-clef-omni-4da0d1b416ca1b82) is designed to return structured decisions from those inputs in one API call, avoiding separate transcription and media-processing steps, according to Cloudflare's October 9th announcement.

The release arrived eight days after Cloudflare introduced Clef and [Clef-flash](https://runtimewire.com/models/huggingface/cloudflare-clef-flash-5253e3667af2eb4d), its first open-weight models for schema-bound decisions. In its first announcement, the company said the initial Clef effort went from a Friday-evening decision to enter the category to a Thursday launch, with training over the weekend. The fast follow-up extends that short-cycle product push: along with Clef-omni, Cloudflare lowered Clef-flash's hosted price and improved the serving speed of Clef.

The move puts the team behind Clef into a direct contest with TypeSafe's Jev in a narrow but useful part of the agent stack: deciding among predefined actions, categories or values. Instead of asking a general-purpose model to generate text and then parsing that text into an application decision, Clef returns probabilities for options supplied in a schema. That format can route a support request, classify a security event or select an action for an agent.

### One call for mixed media

Clef-omni accepts WAV or MP3 audio and MP4 or WebM video, in addition to text and images, Cloudflare says. Its [Hugging Face model card](https://huggingface.co/Cloudflare/clef-omni?ref=runtimewire) describes a 30B-A3B mixture-of-experts model post-trained from [Qwen3-Omni-30B-A3B-Instruct](https://runtimewire.com/models/huggingface/qwen-qwen3-omni-30b-a3b-instruct-4f7ce1648f9f9fd4). The model uses the foundation's comprehension backbone and audio and vision encoders, while its speech-output components are not loaded for inference.

That architecture reflects the product's specific task. Clef-omni processes the media and scores allowed answers rather than generating a transcript or caption and asking another model to interpret it. The model card describes a joint schema head that scores the options across questions together. Cloudflare says the training approach freezes the Qwen backbone, adds low-rank adapters and tunes for calibrated schema decisions.

Cloudflare reports median response times of about 130 milliseconds for text, about 150 milliseconds for image inputs and a few hundred milliseconds for audio. It says a 21-second video with sound takes about 1.5 seconds to score. These are company-reported figures; the published material does not provide independent testing of accuracy or latency on customer workloads.

The benchmark table also shows why the model's multimodal breadth should not be read as a universal quality win. Clef-omni posts strong numbers on some listed tests, including 94.8 macro-F1 on BANKING77 and 97.7 on CLINC150+OOS, while scoring below the existing Clef model on several others. On Cloudflare's TypeSafe workflow tests, Clef-omni scores 60.2 for invoice-processing exact actions, compared with 64.7 for Clef and 61.8 for Jev. The comparisons are Cloudflare's own evaluations, not an independent audit.

### Lower price, smaller hosted window

Cloudflare cut Clef-flash's hosted input price from $0.09 to $0.038 per million tokens. The [current Workers AI pricing page](https://developers.cloudflare.com/workers-ai/platform/pricing/?ref=runtimewire) lists Clef at $0.24 per million input tokens and Clef-omni at $0.15. Cloudflare says Clef-omni is now cheaper than Jev, though the comparison depends on the applicable Jev price and workload.

The lower Clef-flash price comes with a meaningful limit: Cloudflare reduced the hosted model's context window from 64,000 tokens to 24,000. The company says 0.24% of its Clef-flash requests exceeded 24,000 input tokens. Its published model weights remain unchanged, and Cloudflare says self-hosters can use a 256,000-token context window. That makes the tradeoff unusually explicit: hosted users get a cheaper option if their prompts fit, while teams with longer inputs may need to move to Clef or manage their own deployment.

Cloudflare also says its hosted Clef model is up to 2.0 times faster after serving-layer optimizations; it released no new Clef weights for that change. The updates point to a business strategy built around the full Workers AI stack: publish weights for developers who want to run models themselves, offer hosted inference for teams that prioritize convenience, and make smaller, cheaper decisions practical inside repeated agent workflows. The company has previously said it uses Clef internally for domain classification in its threat-intelligence work, giving the product a concrete operational use case alongside its model benchmarks.

The question for adopters is whether a single multimodal decision call removes enough pipeline complexity to offset the cost and accuracy of processing media in one model. Cloudflare has published the architecture, the weights and its own evaluation results. Independent tests across real audio and video workflows will determine how often the one-call design holds up outside the demo and benchmark setting.
