{"slug": "qwen3-8-flash-next-is-on-tensorfold-with-speed-boosts", "title": "Qwen3.8-Flash-Next Is on TensorFold with Speed Boosts", "summary": "TensorFold 0.3.6.2 delivered decode speeds over 62 tokens per second on a single stream and 119 tokens per second across five concurrent streams running Qwen3.8-Flash-Next on a single Nvidia DGX Spark, according to benchmark results posted on X. The new inference engine also reached 2,500 tokens per second prefill with a 256k default context window and a KV cache pool of roughly 1.3M, outperforming prior vLLM (MiaAI-Lab) setups in time-to-first-token and multi-user tests. Community builders reported running comparable performance on consumer hardware including an RTX 3090 with 64GB of RAM and 100GB of NVMe for about $3,000 in compute.", "body_md": "Qwen3.8-Flash-Next for a single DGX Spark got a serious upgrade with TensorFold🔥\nThis is a completely new recipe, optimized and tuned for TesnorFold! Expect further improvements!\n- KV cache pool is ~1.3M\n- Default context 256k, with 5 concurrent.\n- Faster everything compared \n\n# TensorFold Boosts Qwen3.8-Flash-Next Speed on Single DGX Spark\n\nLast updated Sep 29, 2026\n\nTensorFold's new inference engine delivers decode speeds over 62 tokens per second on a single stream and 119 across five, with prefill at 2,500 tokens per second and a 256k context window. It outperforms prior vLLM setups in time-to-first-token and multi-user tests, while supporting vision and video inputs on DGX Spark's unified memory. Community builders praise its accessibility, running strong performance on consumer hardware like RTX 3090 with 64GB RAM, and look forward to tweaks for models like GLM 5.3 Flash.\n\nThis story is a summary of posts on X and may evolve over time. Grok can make mistakes, verify its outputs.\n\n## Related Trending Stories on X\n\nYou can now run an LLM which scores higher than Sonnet-5 medium, and GPT-6-Sol medium for 3000$ in compute \n- 65 decode tok/s\n- 2000 prefill tok/s\n- 200k kv cache\n1. RTX 3090 / 4090 / 5090 / Intel Arc B70 / AMD\n2. 64 GB of RAM\n3. 100 GB of NVMe\nRecipe today\n\nEdit: In the previous post I was running old vllm version.\nI ran the tests again with the following setup:\n𝐐𝐰𝐞𝐧𝟑.𝟖 𝐅𝐥𝐚𝐬𝐡 𝐍𝐞𝐱𝐭, 𝐨𝐧𝐞 𝐃𝐆𝐗 𝐒𝐩𝐚𝐫𝐤, 𝟐𝟔𝟐𝐤 𝐜𝐨𝐧𝐭𝐞𝐱𝐭: 𝐓𝐞𝐧𝐬𝐨𝐫𝐅𝐨𝐥𝐝 𝟎.𝟑.𝟔.𝟐 𝐯𝐬 𝐯𝐋𝐋𝐌 (𝐌𝐢𝐚𝐀𝐈-𝐋𝐚𝐛 𝐬𝐞𝐭𝐮𝐩). \n\nwhat an insane performance leap and great accomplishment. Congratulations to Mia and \n\n[@ashxhart](https://x.com/ashxhart)can’t wait to see the next recipes\nQwen3.8-Flash-Next for a single DGX Spark got a serious upgrade with TensorFold🔥\nThis is a completely new recipe, optimized and tuned for TesnorFold! Expect further improvements!\n- KV cache pool is ~1.3M\n- Default context 256k, with 5 concurrent.\n- Faster everything compared", "url": "https://wpnews.pro/news/qwen3-8-flash-next-is-on-tensorfold-with-speed-boosts", "canonical_source": "https://x.com/i/trending/2104894082678214786", "published_at": "2026-09-29 11:44:45+00:00", "updated_at": "2026-09-29 11:47:40.716174+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-tools", "mlops", "ai-chips"], "entities": ["TensorFold", "Qwen3.8-Flash-Next", "Nvidia DGX Spark", "vLLM", "MiaAI-Lab", "RTX 3090", "GLM 5.3 Flash", "Grok"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/qwen3-8-flash-next-is-on-tensorfold-with-speed-boosts", "markdown": "https://wpnews.pro/news/qwen3-8-flash-next-is-on-tensorfold-with-speed-boosts.md", "text": "https://wpnews.pro/news/qwen3-8-flash-next-is-on-tensorfold-with-speed-boosts.txt", "jsonld": "https://wpnews.pro/news/qwen3-8-flash-next-is-on-tensorfold-with-speed-boosts.jsonld"}}