# Qwen3 30B A3B - cheapest: DeepInfra $0.12/M input

> Source: <https://tokenstead.ai/models/qwen3-30b-a3b>
> Published: 2026-08-18 22:21:35+00:00

# Qwen3 30B A3B

MoE enthusiast30.5B total params, 3B active per forward pass (MoE). Min memory based on full weight loading ~15GB.

AI-generated content marks

This model embeds **text watermarks** in generated text and adds **C2PA provenance metadata** to supported files such as .png, .jpg, and .svg.
Marks can be lost through editing, screenshots, or format conversion, so their absence does not prove a file is human-made.

- 30.5B
- 128k
- apache 2.0
- Apr 2025

## Scores

## Score per dollar

600 pts per $/M input

general_score (72) divided by cheapest input price
($0.12/M).
Higher is better value. [See live pricing](/models/qwen3-30b-a3b/pricing).

## Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

### Can you run it? - reference rigs

| Rig | Q4_K_M | Q8_0 |
|---|---|---|
| NVIDIA Jetson Orin NX 16GB | tight |
|

[no -> cloud](#cloud-pricing)Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

## Download options

## Or run it in the cloud

Live per-provider pricing, throughput and uptime - refreshed about 19 hours ago via OpenRouter. Click a column to sort.

some pricing may be stale - last verified 2026-08-18

| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
|
DeepInfra
|
API | 0.12 | 0.50 | - | - | - | 100.00% | best uptime |
|
Alibaba
|
API | 0.13 | 0.52 | - | - | - | 100.00% | |
|
NextBit
stale
|
API | 0.14 | 0.55 | - | - | - | 100.00% |

Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.

[Detailed API pricing page + JSON endpoint →](/models/qwen3-30b-a3b/pricing)

[See who runs Alibaba in production →](/adoption/alibaba)

## Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.
