# Alibaba switches on Qwen3.8-Flash API at $0.16 per million input tokens

> Source: <https://runtimewire.com/article/alibaba-qwen3-8-flash-api-qwencloud-pricing>
> Published: 2026-08-26 16:05:54+00:00

# Alibaba switches on Qwen3.8-Flash API at $0.16 per million input tokens

**The same-day follow-up puts Qwen's new model on QwenCloud with a default 1-million-token context and OpenAI and Anthropic API compatibility.**

By [RuntimeWire Staff](/author/runtimewire-staff)
· Published

Primary source: [Qwen / Alibaba](https://x.com/Alibaba_Qwen/status/2092636376990990503)

## Why it matters

The live API turns the same-day Flash-Next model release into a product developers can price and integrate. DeepSeek already offers comparable context length and API compatibility, leaving QwenCloud to compete on inference cost, managed tools and production performance.

[QwenCloud](https://www.qwencloud.com/models/qwen3.8-flash?ref=runtimewire) now serves Qwen3.8-Flash at $0.16 per million input tokens and $0.47 per million output tokens, with a default 1-million-token context and compatibility with OpenAI and Anthropic API formats. QwenCloud's [pricing documentation](https://docs.qwencloud.com/developer-guides/getting-started/pricing?ref=runtimewire) lists the published token rates.

The activation is a same-day follow-up to [Qwen's technical announcement](https://qwen.ai/blog?id=qwen3.8-flash-next&ref=runtimewire), which released the Qwen3.8-Flash-Next weights and listed the production API as "coming soon." The change since that announcement is commercial availability: developers can now send requests to the managed service at published rates.

[Alibaba's](https://home.alibabagroup.com/en-US/about-alibaba?ref=runtimewire) Qwen organization announced the API launch on August 26, 2026.

### What developers can buy

[QwenCloud's platform introduction](https://docs.qwencloud.com/developer-guides/getting-started/introduction?ref=runtimewire) describes its managed API service. Qwen says the production version includes official built-in tools, and QwenCloud's platform materials describe [web search support](https://docs.qwencloud.com/developer-guides/text-generation/web-search?ref=runtimewire).

The hosted model packages the open-weight Qwen3.8-Flash-Next release into a production service with the longer context enabled by default. Qwen's [repository](https://github.com/QwenLM/Qwen3.8-Flash-Next?ref=runtimewire) documents support for OpenAI and Anthropic API specifications, reducing the integration work for teams already using those request formats.

### The rate card

At the published token rates, a workload processing 1 million input tokens and generating 100,000 output tokens would incur $0.207 in model charges. QwenCloud also supports context caching, and its [pricing documentation](https://docs.qwencloud.com/developer-guides/getting-started/pricing?ref=runtimewire) lists separate cached-input pricing.

The rate card gives developers a concrete basis for comparing managed inference against self-hosting the open weights. It does not account for tool charges, storage, engineering labor or the infrastructure required to operate the model independently.

### What carries over from Flash-Next

Qwen3.8-Flash-Next is an open-weight multimodal mixture-of-experts model and an early preview of the architecture intended for Qwen4. Qwen describes a 125-billion-parameter main network, another 51 billion parameters in n-gram embeddings and roughly 6 billion parameters activated for each token.

In an experimental benchmark using a 1-million-token context and a 90% prefix-cache hit rate, [Alibaba reports](https://qwen.ai/blog?id=qwen3.8-flash-next&ref=runtimewire) 8.6 times the prefill throughput of [Qwen3.7-Plus](/models/qwen/qwen3.7-plus). The result is a company benchmark under a specified cache condition and does not establish general production performance. Qwen also says training used about one-ninth as much compute as Qwen3.7-Plus.

### Long context is already a competitive baseline

DeepSeek's official [V4 announcement](https://deepseek.com/en/news/v4-preview/?ref=runtimewire) says V4 Flash and V4 Pro also support 1-million-token contexts. Its [API documentation](https://api-docs.deepseek.com/quick_start/pricing/?ref=runtimewire) provides OpenAI ChatCompletions and Anthropic-compatible interfaces. Qwen's context length and request formats therefore match capabilities already available from another low-cost model provider. QwenCloud adds its published $0.16 input-token rate and managed tools to that baseline.

Developers using OpenAI or Anthropic request formats can evaluate Alibaba's managed inference without first rewriting their integration or operating model-serving infrastructure. Compatibility does not guarantee identical behavior across providers, though it lowers the initial engineering cost of testing QwenCloud against an existing workload.

Alibaba has not disclosed standalone customer counts, API volume or revenue for QwenCloud. The commercial test is whether its token prices, long-context capacity, built-in tools and familiar API protocols can move production workloads onto Alibaba's infrastructure.
