{"slug": "gemma-4-26b-a4b-cheapest-dekallm-0-06-m-input", "title": "Gemma 4 26B A4B - cheapest: DekaLLM $0.06/M input", "summary": "Google's Gemma 4 26B A4B, a 25.2B-parameter multimodal mixture-of-experts model with 3.8B active parameters and 128 experts, is now available with DekaLLM offering the cheapest API pricing at $0.06 per million input tokens and $0.33 per million output tokens. The model, which supports 256k context and runs on NVIDIA Blackwell or Apple Silicon, scores 84 on general quality, yielding 1400 points per dollar per million input tokens. Google lists the model at $0.15 per million input tokens and $0.60 per million output tokens.", "body_md": "# Gemma 4 26B A4B\n\nMoE enthusiast25.2B total, 3.8B active per token (MoE: 128 experts, top-8 + 1 shared). Multimodal (text + image). The small-MoE sweet spot - near-4B decode speed with 25B-class quality. NVFP4 (~14GB) needs NVIDIA Blackwell (vLLM) or Apple Silicon (MLX); Q4_K_M (~13GB weights, ~18GB runtime) for Ollama/llama.cpp.\n\nAI-generated content marks\n\nThe provider reports that this model does not add embedded watermarks or provenance metadata to generated output.\n\n- 25.2B\n- 256k\n- other\n- Apr 2026\n\n## Scores\n\n## Score per dollar\n\n1400 pts per $/M input\n\ngeneral_score (84) divided by cheapest input price\n($0.06/M).\nHigher is better value. [See live pricing](/models/gemma-4-26b-a4b/pricing).\n\n## Run it locally\n\nPer-quant memory needs and a static \"can you run it?\" reference - no rig entry required\n\n### Can you run it? - reference rigs\n\n| Rig | NVFP4 | Q4_K_M | Q8_0 | BF16 |\n|---|---|---|---|---|\n| NVIDIA Jetson Orin NX 16GB |\n|\n\n[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.\n\n## Download options\n\n## Or run it in the cloud\n\nLive per-provider pricing, throughput and uptime - refreshed about 19 hours ago via OpenRouter. Click a column to sort.\n\nsome pricing may be stale - last verified 2026-08-18\n\n| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |\n|---|---|---|---|---|---|---|---|---|\n|\nWafer\nstale\n|\nAPI | 0.13 | 0.40 | 0.050 | - | - | 99.97% | best uptime |\n|\nCloudflare\n|\nAPI | 0.10 | 0.30 | - | - | - | 99.95% | |\n|\nNextBit\n|\nAPI | 0.10 | 0.40 | 0.050 | - | - | 99.88% | |\n|\nNovita\n|\nAPI | 0.13 | 0.40 | - | - | - | 99.85% | |\n|\nDeepInfra\n|\nAPI | 0.07 | 0.34 | - | - | - | 99.70% | |\n|\nDekaLLM\nstale\n|\nAPI | 0.06 | 0.33 | 0.050 | - | - | 99.61% | cheapest |\n|\nVenice\n|\nAPI | 0.13 | 0.40 | 0.050 | - | - | 99.53% | |\n|\nGoogle\n|\nAPI | 0.15 | 0.60 | - | - | - | 99.50% | |\n|\nParasail\n|\nAPI | 0.13 | 0.40 | 0.050 | - | - | 99.42% | |\n|\nSiliconFlow\n|\nAPI | 0.12 | 0.40 | - | - | - | 99.31% | |\n|\nIonstream\n|\nAPI | 0.13 | 0.40 | 0.050 | - | - | 99.19% |\n\nDefault order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are \"-\". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.\n\n[Detailed API pricing page + JSON endpoint →](/models/gemma-4-26b-a4b/pricing)\n\n[See who runs Google in production →](/adoption/google)\n\n## Inference cost over time\n\nData accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.", "url": "https://wpnews.pro/news/gemma-4-26b-a4b-cheapest-dekallm-0-06-m-input", "canonical_source": "https://tokenstead.ai/models/gemma-4-26b-a4b", "published_at": "2026-08-18 22:21:35+00:00", "updated_at": "2026-08-18 22:42:57.125037+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "generative-ai", "ai-products"], "entities": ["Google", "Gemma 4 26B A4B", "DekaLLM", "OpenRouter", "NVIDIA", "Apple", "Cloudflare", "DeepInfra"], "alternates": {"html": "https://wpnews.pro/news/gemma-4-26b-a4b-cheapest-dekallm-0-06-m-input", "markdown": "https://wpnews.pro/news/gemma-4-26b-a4b-cheapest-dekallm-0-06-m-input.md", "text": "https://wpnews.pro/news/gemma-4-26b-a4b-cheapest-dekallm-0-06-m-input.txt", "jsonld": "https://wpnews.pro/news/gemma-4-26b-a4b-cheapest-dekallm-0-06-m-input.jsonld"}}