Qwen3.8-Flash-Next Is on TensorFold with Speed Boosts TensorFold 0.3.6.2 delivered decode speeds over 62 tokens per second on a single stream and 119 tokens per second across five concurrent streams running Qwen3.8-Flash-Next on a single Nvidia DGX Spark, according to benchmark results posted on X. The new inference engine also reached 2,500 tokens per second prefill with a 256k default context window and a KV cache pool of roughly 1.3M, outperforming prior vLLM (MiaAI-Lab) setups in time-to-first-token and multi-user tests. Community builders reported running comparable performance on consumer hardware including an RTX 3090 with 64GB of RAM and 100GB of NVMe for about $3,000 in compute. Qwen3.8-Flash-Next for a single DGX Spark got a serious upgrade with TensorFold๐Ÿ”ฅ This is a completely new recipe, optimized and tuned for TesnorFold Expect further improvements - KV cache pool is ~1.3M - Default context 256k, with 5 concurrent. - Faster everything compared TensorFold Boosts Qwen3.8-Flash-Next Speed on Single DGX Spark Last updated Sep 29, 2026 TensorFold's new inference engine delivers decode speeds over 62 tokens per second on a single stream and 119 across five, with prefill at 2,500 tokens per second and a 256k context window. It outperforms prior vLLM setups in time-to-first-token and multi-user tests, while supporting vision and video inputs on DGX Spark's unified memory. Community builders praise its accessibility, running strong performance on consumer hardware like RTX 3090 with 64GB RAM, and look forward to tweaks for models like GLM 5.3 Flash. This story is a summary of posts on X and may evolve over time. Grok can make mistakes, verify its outputs. Related Trending Stories on X You can now run an LLM which scores higher than Sonnet-5 medium, and GPT-6-Sol medium for 3000$ in compute - 65 decode tok/s - 2000 prefill tok/s - 200k kv cache 1. RTX 3090 / 4090 / 5090 / Intel Arc B70 / AMD 2. 64 GB of RAM 3. 100 GB of NVMe Recipe today Edit: In the previous post I was running old vllm version. I ran the tests again with the following setup: ๐๐ฐ๐ž๐ง๐Ÿ‘.๐Ÿ– ๐…๐ฅ๐š๐ฌ๐ก ๐๐ž๐ฑ๐ญ, ๐จ๐ง๐ž ๐ƒ๐†๐— ๐’๐ฉ๐š๐ซ๐ค, ๐Ÿ๐Ÿ”๐Ÿ๐ค ๐œ๐จ๐ง๐ญ๐ž๐ฑ๐ญ: ๐“๐ž๐ง๐ฌ๐จ๐ซ๐…๐จ๐ฅ๐ ๐ŸŽ.๐Ÿ‘.๐Ÿ”.๐Ÿ ๐ฏ๐ฌ ๐ฏ๐‹๐‹๐Œ ๐Œ๐ข๐š๐€๐ˆ-๐‹๐š๐› ๐ฌ๐ž๐ญ๐ฎ๐ฉ . what an insane performance leap and great accomplishment. Congratulations to Mia and @ashxhart https://x.com/ashxhart canโ€™t wait to see the next recipes Qwen3.8-Flash-Next for a single DGX Spark got a serious upgrade with TensorFold๐Ÿ”ฅ This is a completely new recipe, optimized and tuned for TesnorFold Expect further improvements - KV cache pool is ~1.3M - Default context 256k, with 5 concurrent. - Faster everything compared