{"slug": "alibaba-ships-qwen3-8-omni-flash-to-watch-listen-and-call-tools", "title": "Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools", "summary": "Alibaba's Qwen team released Qwen3.8-Omni-Flash through the Qianwen AI Platform on September 18th local time, a model that accepts text, images, audio and video inputs across a 1 million-token context window but natively outputs only text. According to Alibaba Cloud's model documentation, the model is available via the Chat Completions and Responses APIs with function calling and web search, and is priced at $0.15 per million input tokens, $0.016 per million cache-hit input tokens and $0.47 per million output tokens. Alibaba positions Qwen3.8-Omni-Flash as an orchestration layer for agent developers, pairing it with Qwen-MM-Plugins for multimodal agent workflows and Qwen-Live Harness for continuous audiovisual interaction, with availability in Beijing, Singapore, Hong Kong, Tokyo, Frankfurt and Virginia.", "body_md": "# Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools\n\n**The text-output model accepts text, images, audio and video across a 1M-token context window, with function calling and web search for work beyond analysis.**\n\n        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)\n        · Published \n\nPrimary source: [Qwen](https://qwen.ai/blog?id=qwen3.8-omni-flash)\n\n## Why it matters\n\nAlibaba is positioning Qwen3.8-Omni-Flash as a multimodal analysis and orchestration layer for agent developers, pairing a 1M-token context window with function calling, web search and discounted cached inputs.\n\nAlibaba's Qwen team released Qwen3.8-Omni-Flash through the [Qianwen AI Platform](https://www.qianwenai.com/?ref=runtimewire) on September 18th local time, giving developers one model for reading text, images, audio and video, planning multistep jobs and calling tools to finish those tasks.\n\nThe important distinction sits between the inputs and the output. According to [Alibaba Cloud's model documentation](https://www.alibabacloud.com/help/en/model-studio/omni?ref=runtimewire), Qwen3.8-Omni-Flash accepts all four input types through the Chat Completions and Responses APIs, with a 1 million-token context window. Its native output is text. Finished media and other work beyond analysis require external tools coordinated through the model's function-calling capabilities.\n\nThat makes the release a bet on orchestration rather than a single model that directly produces every medium it understands. Alibaba is pairing the model with Qwen-MM-Plugins for multimodal agent workflows and [Qwen-Live Harness](https://qwen.ai/blog?id=qwen3.8-omni-flash&ref=runtimewire) for continuous audiovisual interaction.\n\nAlibaba's [launch announcement](https://qwen.ai/blog?id=qwen3.8-omni-flash&ref=runtimewire) remains marked \"draft,\" although the production documentation was updated on September 18th and lists Qwen3.8-Omni-Flash as available in Beijing, Singapore, Hong Kong, Tokyo, Frankfurt and Virginia.\n\n### An omni model that delegates\n\nAlibaba describes workflows spanning meeting analysis, video editing, translation, film commentary, music-video creation and real-time conversations. The common requirement is a model that can inspect long stretches of audiovisual material, decide what needs to happen and send instructions to specialized software.\n\nThe model can analyze scenes, speakers and events, then invoke custom tools through function calling. For a startup building media software, the appeal is practical: an application can keep its existing production stack and put Qwen above it as the planner.\n\nThe open-source [Qwen-MM-Plugins repository](https://github.com/QwenLM/Qwen-MM-Plugins?ref=runtimewire) presents a tool layer for multimodal agents. Those tools are separate from the model's documented native output. Alibaba's current model documentation describes Qwen3.8-Omni-Flash as a text-output system for text, image, audio and [video understanding](https://runtimewire.com/models/fal/video-understanding), with function calling and web search.\n\nThe surrounding engineering work remains substantial. Long media files still leave developers with engineering work around storage, processing and token budgets. A model can accept video without making video-agent infrastructure simple.\n\n### Alibaba prices the model for repeated media work\n\nAlibaba Cloud's [published pricing](https://www.alibabacloud.com/help/en/model-studio/model-pricing?ref=runtimewire) lists international Qwen3.8-Omni-Flash access at $0.15 per million input tokens, $0.016 per million cache-hit input tokens and $0.47 per million output tokens. Alibaba applies modality-specific token-conversion rules, while publishing one input-token rate for the model rather than separate rates for text and media.\n\nCaching matters more for audiovisual agents than it does for a short chatbot exchange. Editing or analyzing a long recording often requires repeated requests over the same source. Charging cache hits at about one-tenth of the standard international input rate gives developers an economic reason to keep that media context inside a continuing session rather than resend the full representation for every step.\n\nAlibaba says its estimated API cost per hour of audio input fell by more than 98%, while audiovisual input fell by more than 93%. Those percentages come from Alibaba's own calculation: it priced two minutes of source material, multiplied the result by 30 and used 720p video sampled at one frame per second for audiovisual inputs. [Independent testing has not validated the comparison](https://www.aistackcurrent.com/models/qwen3-8-omni-family/?ref=runtimewire).\n\nAlibaba also claims an average improvement of more than 25% over Qwen3.5-Omni-Plus across 29 evaluations. It reports gains of 36.5 points on WildClawBench-MM and 22.3 points on AgenticVBench, alongside a 69.6 score on UniClawBench. Alibaba says audio performance exceeded [Gemini 3.8 Flash](https://runtimewire.com/models/google/gemini-3.8-flash) overall and audiovisual performance came close to Google's model. These are selected, company-reported results from [Qwen's launch announcement](https://qwen.ai/blog?id=qwen3.8-omni-flash&ref=runtimewire), not independent benchmarks.\n\nAlibaba has not published the parameter count, training-compute budget or training-data composition for Qwen3.8-Omni-Flash in its [current model documentation](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-8-omni-flash?ref=runtimewire). The release therefore gives developers firm API specifications and pricing, while offering much less visibility into the model underneath them.\n\n### Qwen is becoming Alibaba's full application stack\n\nQwen3.8-Omni-Flash arrives during a rapid expansion of Alibaba's model and agent portfolio. On August 12th, Alibaba Cloud's Qwen team [released Qwen3.8-2.4T-A95B](https://runtimewire.com/article/qwen3-8-gpt-5-5-pro-reasoning-prefill-test), an open-weight mixture-of-experts model with 2.4 trillion total parameters and 95 billion active during inference. Earlier in September, Alibaba [updated Qwen3.8-Max for coding and enterprise agents](https://runtimewire.com/article/alibaba-qwen3-8-max-0902-coding-enterprise-agents).\n\nThe omni release fills a different position. The large open model gives sophisticated teams control over deployment. [Qwen3.8-Max](https://runtimewire.com/models/qwen/qwen3.8-max) targets demanding text and agent tasks. Qwen3.8-Omni-Flash is the lower-cost sensory layer for applications that must process meetings, recordings, screens and video before taking action.\n\nAlibaba is also tightening the connection between its models and its cloud distribution. RuntimeWire reported in August that [QwenCloud packages models, agent tools and compatible APIs](https://runtimewire.com/article/alibaba-qwencloud-eddie-wu-ai-native-cloud-strategy) as part of Alibaba's roughly $53 billion, three-year AI and cloud infrastructure plan. Qwen3.8-Omni-Flash gives that platform another workload with heavy storage, inference and repeated-context demands, all of which lead back to Alibaba Cloud services.\n\nFor builders, the useful part of the release is the separation of responsibilities. Qwen3.8-Omni-Flash handles perception, reasoning and tool selection. External software manipulates files or performs other requested work. Applications own the final workflow and user experience.\n\nThe remaining test is reliability over long jobs. A 1 million-token context window can hold substantial media representations, yet successful production work also requires accurate timestamps, stable tool calls, recoverable failures and consistent results across many steps. Alibaba's benchmarks describe pieces of that problem. Developers will determine whether the assembled system can complete the whole job.", "url": "https://wpnews.pro/news/alibaba-ships-qwen3-8-omni-flash-to-watch-listen-and-call-tools", "canonical_source": "https://runtimewire.com/article/alibaba-qwen3-8-omni-flash-audio-video-agents", "published_at": "2026-09-18 02:02:35+00:00", "updated_at": "2026-09-18 02:24:00.107859+00:00", "lang": "en", "topics": ["large-language-models", "ai-agents", "ai-products", "ai-tools", "ai-infrastructure"], "entities": ["Alibaba", "Qwen", "Qwen3.8-Omni-Flash", "Qianwen AI Platform", "Alibaba Cloud", "Qwen-MM-Plugins", "Qwen-Live Harness"], "alternates": {"html": "https://wpnews.pro/news/alibaba-ships-qwen3-8-omni-flash-to-watch-listen-and-call-tools", "markdown": "https://wpnews.pro/news/alibaba-ships-qwen3-8-omni-flash-to-watch-listen-and-call-tools.md", "text": "https://wpnews.pro/news/alibaba-ships-qwen3-8-omni-flash-to-watch-listen-and-call-tools.txt", "jsonld": "https://wpnews.pro/news/alibaba-ships-qwen3-8-omni-flash-to-watch-listen-and-call-tools.jsonld"}}