cd /news/artificial-intelligence/cloudflare-puts-audio-and-video-into… · home › topics › artificial-intelligence › article
[ARTICLE · art-148479] src=runtimewire.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Cloudflare puts audio and video into its decision model with Clef-omni

Cloudflare engineers Michelle Chen, Alex Reneau and Kevin Jain released Clef-omni on October 9, a 30B-A3B mixture-of-experts model post-trained from Qwen3-Omni-30B-A3B-Instruct that accepts WAV or MP3 audio and MP4 or WebM video alongside text and images and returns schema-bound decisions in one API call. Cloudflare reported median response times of about 130 milliseconds for text, about 150 milliseconds for images and a few hundred milliseconds for audio, with a 21-second video with sound scoring in about 1.5 seconds, while also lowering Clef-flash's hosted price and shrinking its hosted context window. Clef-omni scored 94.8 macro-F1 on BANKING77 and 97.7 on CLINC150+OOS but 60.2 on Cloudflare's TypeSafe invoice-processing exact-action test, below Clef's 64.7 and Jev's 61.8, in company-run evaluations.

by read4 min views2 publishedOct 9, 2026
Cloudflare puts audio and video into its decision model with Clef-omni
Image: Runtimewire (auto-discovered)

The model takes text, images, audio and video in one request; Cloudflare also cut Clef-flash pricing while reducing its hosted context window.

        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
        · Published 

Primary source: [Cloudflare Blog](https://blog.cloudflare.com/clef-faster-cheaper-multimodal/)

Why it matters #

Cloudflare is making schema-bound decisions work across more types of input while competing on hosted price. The tradeoff is visible: Clef-flash costs less, but its hosted context window is now smaller.

Cloudflare engineers Michelle Chen, Alex Reneau and Kevin Jain have extended the company's Clef decision-model line to accept audio and video alongside text and images. Clef-omni is designed to return structured decisions from those inputs in one API call, avoiding separate transcription and media-processing steps, according to Cloudflare's October 9th announcement.

The release arrived eight days after Cloudflare introduced Clef and Clef-flash, its first open-weight models for schema-bound decisions. In its first announcement, the company said the initial Clef effort went from a Friday-evening decision to enter the category to a Thursday launch, with training over the weekend. The fast follow-up extends that short-cycle product push: along with Clef-omni, Cloudflare lowered Clef-flash's hosted price and improved the serving speed of Clef.

The move puts the team behind Clef into a direct contest with TypeSafe's Jev in a narrow but useful part of the agent stack: deciding among predefined actions, categories or values. Instead of asking a general-purpose model to generate text and then parsing that text into an application decision, Clef returns probabilities for options supplied in a schema. That format can route a support request, classify a security event or select an action for an agent.

One call for mixed media

Clef-omni accepts WAV or MP3 audio and MP4 or WebM video, in addition to text and images, Cloudflare says. Its Hugging Face model card describes a 30B-A3B mixture-of-experts model post-trained from Qwen3-Omni-30B-A3B-Instruct. The model uses the foundation's comprehension backbone and audio and vision encoders, while its speech-output components are not loaded for inference.

That architecture reflects the product's specific task. Clef-omni processes the media and scores allowed answers rather than generating a transcript or caption and asking another model to interpret it. The model card describes a joint schema head that scores the options across questions together. Cloudflare says the training approach freezes the Qwen backbone, adds low-rank adapters and tunes for calibrated schema decisions.

Cloudflare reports median response times of about 130 milliseconds for text, about 150 milliseconds for image inputs and a few hundred milliseconds for audio. It says a 21-second video with sound takes about 1.5 seconds to score. These are company-reported figures; the published material does not provide independent testing of accuracy or latency on customer workloads.

The benchmark table also shows why the model's multimodal breadth should not be read as a universal quality win. Clef-omni posts strong numbers on some listed tests, including 94.8 macro-F1 on BANKING77 and 97.7 on CLINC150+OOS, while scoring below the existing Clef model on several others. On Cloudflare's TypeSafe workflow tests, Clef-omni scores 60.2 for invoice-processing exact actions, compared with 64.7 for Clef and 61.8 for Jev. The comparisons are Cloudflare's own evaluations, not an independent audit.

Lower price, smaller hosted window

Cloudflare cut Clef-flash's hosted input price from $0.09 to $0.038 per million tokens. The current Workers AI pricing page lists Clef at $0.24 per million input tokens and Clef-omni at $0.15. Cloudflare says Clef-omni is now cheaper than Jev, though the comparison depends on the applicable Jev price and workload.

The lower Clef-flash price comes with a meaningful limit: Cloudflare reduced the hosted model's context window from 64,000 tokens to 24,000. The company says 0.24% of its Clef-flash requests exceeded 24,000 input tokens. Its published model weights remain unchanged, and Cloudflare says self-hosters can use a 256,000-token context window. That makes the tradeoff unusually explicit: hosted users get a cheaper option if their prompts fit, while teams with longer inputs may need to move to Clef or manage their own deployment.

Cloudflare also says its hosted Clef model is up to 2.0 times faster after serving-layer optimizations; it released no new Clef weights for that change. The updates point to a business strategy built around the full Workers AI stack: publish weights for developers who want to run models themselves, offer hosted inference for teams that prioritize convenience, and make smaller, cheaper decisions practical inside repeated agent workflows. The company has previously said it uses Clef internally for domain classification in its threat-intelligence work, giving the product a concrete operational use case alongside its model benchmarks.

The question for adopters is whether a single multimodal decision call removes enough pipeline complexity to offset the cost and accuracy of processing media in one model. Cloudflare has published the architecture, the weights and its own evaluation results. Independent tests across real audio and video workflows will determine how often the one-call design holds up outside the demo and benchmark setting.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @cloudflare 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cloudflare-puts-audi…] indexed:0 read:4min 2026-10-09 · —