MoE enthusiast25.2B total, 3.8B active per token (MoE: 128 experts, top-8 + 1 shared). Multimodal (text + image). The small-MoE sweet spot - near-4B decode speed with 25B-class quality. NVFP4 (~14GB) needs NVIDIA Blackwell (vLLM) or Apple Silicon (MLX); Q4_K_M (~13GB weights, ~18GB runtime) for Ollama/llama.cpp.
AI-generated content marks
The provider reports that this model does not add embedded watermarks or provenance metadata to generated output.
- 25.2B
- 256k
- other
- Apr 2026
Scores #
Score per dollar #
1400 pts per $/M input
general_score (84) divided by cheapest input price
($0.06/M).
Higher is better value. [See live pricing](/models/gemma-4-26b-a4b/pricing).
Run it locally #
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
Can you run it? - reference rigs
| Rig | NVFP4 | Q4_K_M | Q8_0 | BF16 |
|---|---|---|---|---|
| NVIDIA Jetson Orin NX 16GB | ||||
no -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudFit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options #
Or run it in the cloud #
Live per-provider pricing, throughput and uptime - refreshed about 19 hours ago via OpenRouter. Click a column to sort.
some pricing may be stale - last verified 2026-08-18
| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
| Wafer | ||||||||
| stale | ||||||||
| API | 0.13 | 0.40 | 0.050 | - | - | 99.97% | best uptime | |
| Cloudflare | ||||||||
| API | 0.10 | 0.30 | - | - | - | 99.95% | ||
| NextBit | ||||||||
| API | 0.10 | 0.40 | 0.050 | - | - | 99.88% | ||
| Novita | ||||||||
| API | 0.13 | 0.40 | - | - | - | 99.85% | ||
| DeepInfra | ||||||||
| API | 0.07 | 0.34 | - | - | - | 99.70% | ||
| DekaLLM | ||||||||
| stale | ||||||||
| API | 0.06 | 0.33 | 0.050 | - | - | 99.61% | cheapest | |
| Venice | ||||||||
| API | 0.13 | 0.40 | 0.050 | - | - | 99.53% | ||
| API | 0.15 | 0.60 | - | - | - | 99.50% | ||
| Parasail | ||||||||
| API | 0.13 | 0.40 | 0.050 | - | - | 99.42% | ||
| SiliconFlow | ||||||||
| API | 0.12 | 0.40 | - | - | - | 99.31% | ||
| Ionstream | ||||||||
| API | 0.13 | 0.40 | 0.050 | - | - | 99.19% |
Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.
[Detailed API pricing page + JSON endpoint →](/models/gemma-4-26b-a4b/pricing)
[See who runs Google in production →](/adoption/google)
Inference cost over time #
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.