{"slug": "qwen3-30b-a3b-cheapest-deepinfra-0-12-m-input", "title": "Qwen3 30B A3B - cheapest: DeepInfra $0.12/M input", "summary": "DeepInfra offers the cheapest API pricing for Qwen3 30B A3B at $0.12 per million input tokens, with Alibaba at $0.13 and NextBit at $0.14, according to live pricing data refreshed about 19 hours ago via OpenRouter. The model, released in April 2025 under Apache 2.0, has 30.5B total parameters with 3B active per forward pass (MoE), a 128k context length, and embeds text watermarks plus C2PA provenance metadata. Its general score of 72 yields 600 points per dollar per million input tokens at the cheapest price.", "body_md": "# Qwen3 30B A3B\n\nMoE enthusiast30.5B total params, 3B active per forward pass (MoE). Min memory based on full weight loading ~15GB.\n\nAI-generated content marks\n\nThis model embeds **text watermarks** in generated text and adds **C2PA provenance metadata** to supported files such as .png, .jpg, and .svg.\nMarks can be lost through editing, screenshots, or format conversion, so their absence does not prove a file is human-made.\n\n- 30.5B\n- 128k\n- apache 2.0\n- Apr 2025\n\n## Scores\n\n## Score per dollar\n\n600 pts per $/M input\n\ngeneral_score (72) divided by cheapest input price\n($0.12/M).\nHigher is better value. [See live pricing](/models/qwen3-30b-a3b/pricing).\n\n## Run it locally\n\nPer-quant memory needs and a static \"can you run it?\" reference - no rig entry required\n\n### Can you run it? - reference rigs\n\n| Rig | Q4_K_M | Q8_0 |\n|---|---|---|\n| NVIDIA Jetson Orin NX 16GB | tight |\n|\n\n[no -> cloud](#cloud-pricing)Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.\n\n## Download options\n\n## Or run it in the cloud\n\nLive per-provider pricing, throughput and uptime - refreshed about 19 hours ago via OpenRouter. Click a column to sort.\n\nsome pricing may be stale - last verified 2026-08-18\n\n| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |\n|---|---|---|---|---|---|---|---|---|\n|\nDeepInfra\n|\nAPI | 0.12 | 0.50 | - | - | - | 100.00% | best uptime |\n|\nAlibaba\n|\nAPI | 0.13 | 0.52 | - | - | - | 100.00% | |\n|\nNextBit\nstale\n|\nAPI | 0.14 | 0.55 | - | - | - | 100.00% |\n\nDefault order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are \"-\". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.\n\n[Detailed API pricing page + JSON endpoint →](/models/qwen3-30b-a3b/pricing)\n\n[See who runs Alibaba in production →](/adoption/alibaba)\n\n## Inference cost over time\n\nData accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.", "url": "https://wpnews.pro/news/qwen3-30b-a3b-cheapest-deepinfra-0-12-m-input", "canonical_source": "https://tokenstead.ai/models/qwen3-30b-a3b", "published_at": "2026-08-18 22:21:35+00:00", "updated_at": "2026-08-18 22:42:53.210554+00:00", "lang": "en", "topics": ["large-language-models", "generative-ai", "ai-products"], "entities": ["DeepInfra", "Alibaba", "NextBit", "Qwen3 30B A3B", "OpenRouter", "C2PA"], "alternates": {"html": "https://wpnews.pro/news/qwen3-30b-a3b-cheapest-deepinfra-0-12-m-input", "markdown": "https://wpnews.pro/news/qwen3-30b-a3b-cheapest-deepinfra-0-12-m-input.md", "text": "https://wpnews.pro/news/qwen3-30b-a3b-cheapest-deepinfra-0-12-m-input.txt", "jsonld": "https://wpnews.pro/news/qwen3-30b-a3b-cheapest-deepinfra-0-12-m-input.jsonld"}}