cd /news/large-language-models/new-model-available-qwen3-8-flash · home › topics › large-language-models › article
[ARTICLE · art-124200] src=zenmux.ai ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

New Model Available: Qwen3.8-Flash

Alibaba's Qwen team released Qwen3.8-Flash, a multimodal model with a million-token context window, native support for OpenAI and Anthropic API protocols, and integration with developer tools like Claude Code and Codex. Priced at $0.16 per million input tokens and $0.47 per million output tokens, it targets coding assistance, agentic workflows, and visual understanding tasks.

read1 min views20 publishedAug 27, 2026
New Model Available: Qwen3.8-Flash
Image: Zenmux (auto-discovered)

Qwen3.8-Flash is the latest multimodal model from the Qwen family, combining powerful reasoning and generation with remarkable speed. It natively supports a million-token context window, allowing it to process lengthy documents, entire codebases, and complex conversations in a single pass. It shines in coding assistance, agentic workflows, and visual understanding — whether it's fixing code autonomously, operating desktop applications, or analyzing charts and long videos. Fully compatible with both OpenAI and Anthropic API protocols, it integrates seamlessly with popular developer tools like Claude Code and Codex, making it easy to build high-concurrency applications and intelligent workflows. With strong performance and highly competitive inference costs, Qwen3.8-Flash is an ideal choice for developers and businesses seeking the best of both worlds in AI applications.

Back to Models

Providers #

Route requests across multiple providers. Copy a provider slug to set your preference.

$0.16

/ M tokens $0.47

/ M tokens Read:

0.016/ M tokens

Write:

0.2/ M tokens1M12.1s33.7tps

Uptime #

24hours Direct request success rate on AI Gateway and per-provider.

Throughput #

24hours P50 throughput on live AI Gateway traffic, in tokens per second (TPS).

Latency #

24hours P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds.

Activity #

Token volume and request traffic to this model over time.

Apps #

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for. View All

More models from Qwen

── more in #large-language-models 4 stories · sorted by recency
── more on @alibaba 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/new-model-available-…] indexed:0 read:1min 2026-08-27 · —