# Qwen3.8-Omni-Flash: text, image, audio and video in one request, 1 million token context, audio input over 98 percent cheaper

> Source: <https://qwen.ai/blog?id=qwen3.8-omni-flash>
> Published: 2026-09-18 15:14:04+00:00

Released September 18. 288 points on Hacker News. It takes text, images, audio and video in and writes text out. Context is 1 million tokens: up to 991K in and 131K out. Price on the international API: $0.15 per million input tokens, $0.47 per million output, $0.016 on cache hits. Qwen says an hour of audio input now costs more than 98 percent less than on its last omni model, Qwen3.5-Omni-Plus. It reports a 25 percent average gain across 29 evals against that model, including big jumps on agent benchmarks like WildClawBench-MM (up 36.5 points) and AgenticVBench (up 22.3). For video it samples the parts that matter instead of reading every frame. Function calling and web search are on. Reasoning is on by default and can be turned off. It is API only through QwenCloud, Alibaba Cloud Model Studio and Qwen Studio. No open weights. Qwen did release Apache-2.0 plugins, Qwen-MM-Plugins, so agent harnesses can use it for audio and video. Why it matters: cheap audio in a 1M window makes 'listen to every call' a normal workload. What to watch: the benchmarks are Qwen's own.
