Gemma 4 26B A4B - cheapest: DekaLLM $0.06/M input Google's Gemma 4 26B A4B, a 25.2B-parameter multimodal mixture-of-experts model with 3.8B active parameters and 128 experts, is now available with DekaLLM offering the cheapest API pricing at $0.06 per million input tokens and $0.33 per million output tokens. The model, which supports 256k context and runs on NVIDIA Blackwell or Apple Silicon, scores 84 on general quality, yielding 1400 points per dollar per million input tokens. Google lists the model at $0.15 per million input tokens and $0.60 per million output tokens. Gemma 4 26B A4B MoE enthusiast25.2B total, 3.8B active per token MoE: 128 experts, top-8 + 1 shared . Multimodal text + image . The small-MoE sweet spot - near-4B decode speed with 25B-class quality. NVFP4 ~14GB needs NVIDIA Blackwell vLLM or Apple Silicon MLX ; Q4 K M ~13GB weights, ~18GB runtime for Ollama/llama.cpp. AI-generated content marks The provider reports that this model does not add embedded watermarks or provenance metadata to generated output. - 25.2B - 256k - other - Apr 2026 Scores Score per dollar 1400 pts per $/M input general score 84 divided by cheapest input price $0.06/M . Higher is better value. See live pricing /models/gemma-4-26b-a4b/pricing . Run it locally Per-quant memory needs and a static "can you run it?" reference - no rig entry required Can you run it? - reference rigs | Rig | NVFP4 | Q4 K M | Q8 0 | BF16 | |---|---|---|---|---| | NVIDIA Jetson Orin NX 16GB | | no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast =20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark. Download options Or run it in the cloud Live per-provider pricing, throughput and uptime - refreshed about 19 hours ago via OpenRouter. Click a column to sort. some pricing may be stale - last verified 2026-08-18 | Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value | |---|---|---|---|---|---|---|---|---| | Wafer stale | API | 0.13 | 0.40 | 0.050 | - | - | 99.97% | best uptime | | Cloudflare | API | 0.10 | 0.30 | - | - | - | 99.95% | | | NextBit | API | 0.10 | 0.40 | 0.050 | - | - | 99.88% | | | Novita | API | 0.13 | 0.40 | - | - | - | 99.85% | | | DeepInfra | API | 0.07 | 0.34 | - | - | - | 99.70% | | | DekaLLM stale | API | 0.06 | 0.33 | 0.050 | - | - | 99.61% | cheapest | | Venice | API | 0.13 | 0.40 | 0.050 | - | - | 99.53% | | | Google | API | 0.15 | 0.60 | - | - | - | 99.50% | | | Parasail | API | 0.13 | 0.40 | 0.050 | - | - | 99.42% | | | SiliconFlow | API | 0.12 | 0.40 | - | - | - | 99.31% | | | Ionstream | API | 0.13 | 0.40 | 0.050 | - | - | 99.19% | Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing. Detailed API pricing page + JSON endpoint → /models/gemma-4-26b-a4b/pricing See who runs Google in production → /adoption/google Inference cost over time Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.