cd /news/large-language-models/qwen3-8-flash-next-is-on-tensorfold-โ€ฆ ยท home โ€บ topics โ€บ large-language-models โ€บ article
[ARTICLE ยท art-141663] src=x.com โ†— pub= topic=large-language-models verified=true sentiment=โ†‘ positive

Qwen3.8-Flash-Next Is on TensorFold with Speed Boosts

TensorFold 0.3.6.2 delivered decode speeds over 62 tokens per second on a single stream and 119 tokens per second across five concurrent streams running Qwen3.8-Flash-Next on a single Nvidia DGX Spark, according to benchmark results posted on X. The new inference engine also reached 2,500 tokens per second prefill with a 256k default context window and a KV cache pool of roughly 1.3M, outperforming prior vLLM (MiaAI-Lab) setups in time-to-first-token and multi-user tests. Community builders reported running comparable performance on consumer hardware including an RTX 3090 with 64GB of RAM and 100GB of NVMe for about $3,000 in compute.

by read2 min views3 publishedSep 29, 2026
Qwen3.8-Flash-Next Is on TensorFold with Speed Boosts
Image: source

Qwen3.8-Flash-Next for a single DGX Spark got a serious upgrade with TensorFold๐Ÿ”ฅ This is a completely new recipe, optimized and tuned for TesnorFold! Expect further improvements!

  • KV cache pool is ~1.3M
  • Default context 256k, with 5 concurrent.
  • Faster everything compared

Last updated Sep 29, 2026

TensorFold's new inference engine delivers decode speeds over 62 tokens per second on a single stream and 119 across five, with prefill at 2,500 tokens per second and a 256k context window. It outperforms prior vLLM setups in time-to-first-token and multi-user tests, while supporting vision and video inputs on DGX Spark's unified memory. Community builders praise its accessibility, running strong performance on consumer hardware like RTX 3090 with 64GB RAM, and look forward to tweaks for models like GLM 5.3 Flash.

This story is a summary of posts on X and may evolve over time. Grok can make mistakes, verify its outputs.

You can now run an LLM which scores higher than Sonnet-5 medium, and GPT-6-Sol medium for 3000$ in compute

  • 65 decode tok/s
  • 2000 prefill tok/s
  • 200k kv cache
  1. RTX 3090 / 4090 / 5090 / Intel Arc B70 / AMD
  2. 64 GB of RAM
  3. 100 GB of NVMe Recipe today

Edit: In the previous post I was running old vllm version. I ran the tests again with the following setup: ๐๐ฐ๐ž๐ง๐Ÿ‘.๐Ÿ– ๐…๐ฅ๐š๐ฌ๐ก ๐๐ž๐ฑ๐ญ, ๐จ๐ง๐ž ๐ƒ๐†๐— ๐’๐ฉ๐š๐ซ๐ค, ๐Ÿ๐Ÿ”๐Ÿ๐ค ๐œ๐จ๐ง๐ญ๐ž๐ฑ๐ญ: ๐“๐ž๐ง๐ฌ๐จ๐ซ๐…๐จ๐ฅ๐ ๐ŸŽ.๐Ÿ‘.๐Ÿ”.๐Ÿ ๐ฏ๐ฌ ๐ฏ๐‹๐‹๐Œ (๐Œ๐ข๐š๐€๐ˆ-๐‹๐š๐› ๐ฌ๐ž๐ญ๐ฎ๐ฉ).

what an insane performance leap and great accomplishment. Congratulations to Mia and

@ashxhartcanโ€™t wait to see the next recipes Qwen3.8-Flash-Next for a single DGX Spark got a serious upgrade with TensorFold๐Ÿ”ฅ This is a completely new recipe, optimized and tuned for TesnorFold! Expect further improvements!

  • KV cache pool is ~1.3M
  • Default context 256k, with 5 concurrent.
  • Faster everything compared
โ”€โ”€ more in #large-language-models 4 stories ยท sorted by recency
โ”€โ”€ more on @tensorfold 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain โ€” perfect for shipping the agent you just read about.

$git push zahid main
โ†’ Live at https://your-agent.zahid.host โœ“
Get free account โ†’ Pricing
from โ‚ฌ0/mo ยท no card required
LIVE [news/qwen3-8-flash-next-iโ€ฆ] indexed:0 read:2min 2026-09-29 ยท โ€”