# Gemma 4 26B A4B - cheapest: DekaLLM $0.06/M input

> Source: <https://tokenstead.ai/models/gemma-4-26b-a4b>
> Published: 2026-08-18 22:21:35+00:00

# Gemma 4 26B A4B

MoE enthusiast25.2B total, 3.8B active per token (MoE: 128 experts, top-8 + 1 shared). Multimodal (text + image). The small-MoE sweet spot - near-4B decode speed with 25B-class quality. NVFP4 (~14GB) needs NVIDIA Blackwell (vLLM) or Apple Silicon (MLX); Q4_K_M (~13GB weights, ~18GB runtime) for Ollama/llama.cpp.

AI-generated content marks

The provider reports that this model does not add embedded watermarks or provenance metadata to generated output.

- 25.2B
- 256k
- other
- Apr 2026

## Scores

## Score per dollar

1400 pts per $/M input

general_score (84) divided by cheapest input price
($0.06/M).
Higher is better value. [See live pricing](/models/gemma-4-26b-a4b/pricing).

## Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

### Can you run it? - reference rigs

| Rig | NVFP4 | Q4_K_M | Q8_0 | BF16 |
|---|---|---|---|---|
| NVIDIA Jetson Orin NX 16GB |
|

[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

## Download options

## Or run it in the cloud

Live per-provider pricing, throughput and uptime - refreshed about 19 hours ago via OpenRouter. Click a column to sort.

some pricing may be stale - last verified 2026-08-18

| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
|
Wafer
stale
|
API | 0.13 | 0.40 | 0.050 | - | - | 99.97% | best uptime |
|
Cloudflare
|
API | 0.10 | 0.30 | - | - | - | 99.95% | |
|
NextBit
|
API | 0.10 | 0.40 | 0.050 | - | - | 99.88% | |
|
Novita
|
API | 0.13 | 0.40 | - | - | - | 99.85% | |
|
DeepInfra
|
API | 0.07 | 0.34 | - | - | - | 99.70% | |
|
DekaLLM
stale
|
API | 0.06 | 0.33 | 0.050 | - | - | 99.61% | cheapest |
|
Venice
|
API | 0.13 | 0.40 | 0.050 | - | - | 99.53% | |
|
Google
|
API | 0.15 | 0.60 | - | - | - | 99.50% | |
|
Parasail
|
API | 0.13 | 0.40 | 0.050 | - | - | 99.42% | |
|
SiliconFlow
|
API | 0.12 | 0.40 | - | - | - | 99.31% | |
|
Ionstream
|
API | 0.13 | 0.40 | 0.050 | - | - | 99.19% |

Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.

[Detailed API pricing page + JSON endpoint →](/models/gemma-4-26b-a4b/pricing)

[See who runs Google in production →](/adoption/google)

## Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.
