Alibaba switches on Qwen3.8-Flash API at $0.16 per million input tokens Alibaba's Qwen organization activated the Qwen3.8-Flash API on QwenCloud on August 26, 2026, pricing it at $0.16 per million input tokens and $0.47 per million output tokens with a default 1-million-token context and OpenAI and Anthropic API compatibility. The same-day launch follows the release of the open-weight Qwen3.8-Flash-Next model, and Alibaba reports 8.6 times the prefill throughput of Qwen3.7-Plus in an experimental benchmark with a 90% prefix-cache hit rate. Alibaba switches on Qwen3.8-Flash API at $0.16 per million input tokens The same-day follow-up puts Qwen's new model on QwenCloud with a default 1-million-token context and OpenAI and Anthropic API compatibility. By RuntimeWire Staff /author/runtimewire-staff ยท Published Primary source: Qwen / Alibaba https://x.com/Alibaba Qwen/status/2092636376990990503 Why it matters The live API turns the same-day Flash-Next model release into a product developers can price and integrate. DeepSeek already offers comparable context length and API compatibility, leaving QwenCloud to compete on inference cost, managed tools and production performance. QwenCloud https://www.qwencloud.com/models/qwen3.8-flash?ref=runtimewire now serves Qwen3.8-Flash at $0.16 per million input tokens and $0.47 per million output tokens, with a default 1-million-token context and compatibility with OpenAI and Anthropic API formats. QwenCloud's pricing documentation https://docs.qwencloud.com/developer-guides/getting-started/pricing?ref=runtimewire lists the published token rates. The activation is a same-day follow-up to Qwen's technical announcement https://qwen.ai/blog?id=qwen3.8-flash-next&ref=runtimewire , which released the Qwen3.8-Flash-Next weights and listed the production API as "coming soon." The change since that announcement is commercial availability: developers can now send requests to the managed service at published rates. Alibaba's https://home.alibabagroup.com/en-US/about-alibaba?ref=runtimewire Qwen organization announced the API launch on August 26, 2026. What developers can buy QwenCloud's platform introduction https://docs.qwencloud.com/developer-guides/getting-started/introduction?ref=runtimewire describes its managed API service. Qwen says the production version includes official built-in tools, and QwenCloud's platform materials describe web search support https://docs.qwencloud.com/developer-guides/text-generation/web-search?ref=runtimewire . The hosted model packages the open-weight Qwen3.8-Flash-Next release into a production service with the longer context enabled by default. Qwen's repository https://github.com/QwenLM/Qwen3.8-Flash-Next?ref=runtimewire documents support for OpenAI and Anthropic API specifications, reducing the integration work for teams already using those request formats. The rate card At the published token rates, a workload processing 1 million input tokens and generating 100,000 output tokens would incur $0.207 in model charges. QwenCloud also supports context caching, and its pricing documentation https://docs.qwencloud.com/developer-guides/getting-started/pricing?ref=runtimewire lists separate cached-input pricing. The rate card gives developers a concrete basis for comparing managed inference against self-hosting the open weights. It does not account for tool charges, storage, engineering labor or the infrastructure required to operate the model independently. What carries over from Flash-Next Qwen3.8-Flash-Next is an open-weight multimodal mixture-of-experts model and an early preview of the architecture intended for Qwen4. Qwen describes a 125-billion-parameter main network, another 51 billion parameters in n-gram embeddings and roughly 6 billion parameters activated for each token. In an experimental benchmark using a 1-million-token context and a 90% prefix-cache hit rate, Alibaba reports https://qwen.ai/blog?id=qwen3.8-flash-next&ref=runtimewire 8.6 times the prefill throughput of Qwen3.7-Plus /models/qwen/qwen3.7-plus . The result is a company benchmark under a specified cache condition and does not establish general production performance. Qwen also says training used about one-ninth as much compute as Qwen3.7-Plus. Long context is already a competitive baseline DeepSeek's official V4 announcement https://deepseek.com/en/news/v4-preview/?ref=runtimewire says V4 Flash and V4 Pro also support 1-million-token contexts. Its API documentation https://api-docs.deepseek.com/quick start/pricing/?ref=runtimewire provides OpenAI ChatCompletions and Anthropic-compatible interfaces. Qwen's context length and request formats therefore match capabilities already available from another low-cost model provider. QwenCloud adds its published $0.16 input-token rate and managed tools to that baseline. Developers using OpenAI or Anthropic request formats can evaluate Alibaba's managed inference without first rewriting their integration or operating model-serving infrastructure. Compatibility does not guarantee identical behavior across providers, though it lowers the initial engineering cost of testing QwenCloud against an existing workload. Alibaba has not disclosed standalone customer counts, API volume or revenue for QwenCloud. The commercial test is whether its token prices, long-context capacity, built-in tools and familiar API protocols can move production workloads onto Alibaba's infrastructure.