{"slug": "anyone-going-to-repurpose-a-bc250-crypto-gpu-for-local-ai", "title": "Anyone going to repurpose a bc250 crypto GPU for local AI", "summary": "A community kernel patch by S. Duggan re-enables 16 factory-masked Compute Units on the AMD GFX1013 die of the bc250 crypto GPU, boosting median generation speed by 1.32× and median prefill speed by 1.50× across 11 LLM models in llama.cpp b9265 at 4K context. The unlocked 40-CU configuration raised generation speeds from 104 to 129 tok/s on Granite 4.0-H Tiny, 66 to 87.5 tok/s on GPT-OSS 20B MXFP4, 59.5 to 78.7 tok/s on Qwen3.5 35B-A3B MoE, and 9.6 to 14.9 tok/s on QwQ-32B IQ2_M, with power draw rising from 101 W to 116 W median.", "body_md": "**TL;DR:** The GFX1013 die physically has 40 Compute Units; the factory firmware masks 16 of them. A community kernel patch by S. Duggan re-enables the remaining CUs. After an independent FP32 sanity check (100M error-free multiply-adds), a controlled paired A/B re-run on 11 models followed. Key results at 4K context, llama.cpp b9265, verified-stock 24 CU → patched 40 CU:\n\n- **Median generation speed-up: 1.32×** (all 11 deltas positive, p < 0.01)\n- **Median prefill speed-up: 1.50×** (prefill gains more because it’s more compute-bound)\n- **Granite 4.0-H Tiny (hybrid Mamba): 104 → 129 tok/s** (+24%)\n- **GPT-OSS 20B MXFP4: 66 → 87.5 tok/s** (+32%)\n- **Qwen3.5 35B-A3B MoE: 59.5 → 78.7 tok/s** (+32%)\n- **QwQ-32B IQ2_M (reasoning dense): 9.6 → 14.9 tok/s** (+55%)\n\nThe speed-up is larger for prefill than generation, consistent with generation being memory-bandwidth-bound on this hardware (adding compute units helps less than reducing weight bytes per token). A roofline measurement confirms: peak streaming bandwidth 357 GB/s, peak FP32 3901 GFLOP/s, ridge point 10.9 FLOP/byte — all LLM decode quantizations sit left of the ridge (bandwidth-bound regime). The unlock is bounded by power/cooling: the 40-CU arm ran at 116 W median vs 101 W at 24 CU, holding the clock ~3–4% lower under the oberon self-throttling governor. A better-cooled unit would likely gain more.\n\nThe long-context picture is strong. The Qwen3.5 35B-A3B MoE **runs to 64K filled** (16.6 tok/s, n=3) — provided `ttm.pages_limit` is genuinely at the full 16 GiB (§3.3). The dense qwen3.5:9b also runs to 64K (16.4 tok/s) and retrieves all 5 needles at 128K; its smaller footprint leaves more headroom alongside the always-on services, so it’s the safer production default — but the MoE is now a genuine long-context option, not short-context-only. And the **40-CU unlock carries this into long context**: re-run under the 14.7 GiB working-set guard, all four heavy MoEs climb the full 4K→64K ladder at 40-CU with 2/2 needle retrieval at every tier — the flagship Qwen3.5-35B-A3B at 74→48 tok/s — so the active-parameter speed advantage survives the long-context regime under the unlock, not just at stock 24-CU (§B9.2).", "url": "https://wpnews.pro/news/anyone-going-to-repurpose-a-bc250-crypto-gpu-for-local-ai", "canonical_source": "https://forum.level1techs.com/t/anyone-going-to-repurpose-a-bc250-crypto-gpu-for-local-ai/255207#post_1", "published_at": "2026-09-06 21:56:27+00:00", "updated_at": "2026-09-06 23:30:55.318016+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-infrastructure"], "entities": ["AMD", "GFX1013", "bc250", "S. Duggan", "llama.cpp", "Granite 4.0-H Tiny", "GPT-OSS 20B", "Qwen3.5 35B-A3B"], "alternates": {"html": "https://wpnews.pro/news/anyone-going-to-repurpose-a-bc250-crypto-gpu-for-local-ai", "markdown": "https://wpnews.pro/news/anyone-going-to-repurpose-a-bc250-crypto-gpu-for-local-ai.md", "text": "https://wpnews.pro/news/anyone-going-to-repurpose-a-bc250-crypto-gpu-for-local-ai.txt", "jsonld": "https://wpnews.pro/news/anyone-going-to-repurpose-a-bc250-crypto-gpu-for-local-ai.jsonld"}}