Run frontier models on gaming GPUs FreeToken, a new inference engine from FlashML, lets users run frontier models on gaming GPUs at interactive speeds, with Qwen3.6 35B running on an 8GB RTX 4060 laptop at 39 tokens per second, DeepSeek-V4-Flash 284B on an RTX 5090 desktop at 22-25 tokens per second, and GLM-5.2 753B on an RTX PRO 6000 workstation at 15 tokens per second. The tool claims 3-4x faster decode and 6-30x faster prefill compared to Ollama, using bandwidth-adaptive CPU-GPU execution and semantic-aware caching, and is available for free on Windows and Linux. Your gaming PC can now serve frontier models at interactive speed using official checkpoints without extreme quantization Qwen3.6 35B → 8GB RTX 4060 laptop @ 39 tok/s DeepSeek-V4-Flash 284B → RTX 5090 desktop @ 22-25 tok/s GLM-5.2 753B → RTX PRO 6000 workstation @ 15 tok/s Shuo Yang on X: "Download: https://t.co/McSu3Uc1Cf Code: https://t.co/fe3lnZU7xO Reply with your GPU + RAM, and I'll tell you the biggest frontier model your machine can run 👇" - FreeToken is fast. Comparing to Ollama, we have 3–4× faster decode, and 6–30× faster prefill How? We introduce bandwidth-adaptive CPU–GPU execution + semantic-aware caching across agent turns. More details in the technical report: arxiv.org/abs/2608.16157 http://arxiv.org/abs/2608.16157 - FreeToken provides native GUI. No GGUF conversion. No building from source. One-click install on Windows and Linux. FreeToken-desktop ships with agent harnesses built in — pick a model, pick an app, go. - Download: flashml.ai http://flashml.ai Code: github.com/FlashML-org/Fr… http://github.com/FlashML-org/FreeToken Reply with your GPU + RAM, and I'll tell you the biggest frontier model your machine can run 👇flashml.aiFreeToken — Bring Frontier to EdgeDownload and run large language models on your own machine. Free for Windows & Linux. - 256gb ddr4 ecc xeon e5 2680 v4 and 2x rtx306012gb and 2x rtx50608gb and via rpc rtx 3080 16gb - Amazing Rtx 4080 16GB, 32GB RAM.