Your gaming PC can now serve frontier models at interactive speed using official checkpoints without extreme quantization! Qwen3.6 35B → 8GB RTX 4060 laptop @ 39 tok/s DeepSeek-V4-Flash 284B → RTX 5090 desktop @ 22-25 tok/s GLM-5.2 753B → RTX PRO 6000 workstation @ 15 tok/s
- FreeToken is fast. Comparing to Ollama, we have 3–4× faster decode, and 6–30× faster prefill How? We introduce bandwidth-adaptive CPU–GPU execution + semantic-aware caching across agent turns. More details in the technical report: arxiv.org/abs/2608.16157 - FreeToken provides native GUI. No GGUF conversion. No building from source. One-click install on Windows and Linux. FreeToken-desktop ships with agent harnesses built in — pick a model, pick an app, go.
- Download: flashml.aiCode:github.com/FlashML-org/Fr…Reply with your GPU + RAM, and I'll tell you the biggest frontier model your machine can run 👇flashml.aiFreeToken — Bring Frontier to EdgeDownload and run large language models on your own machine. Free for Windows & Linux. - 256gb ddr4 ecc xeon e5 2680 v4 and 2x rtx306012gb and 2x rtx50608gb and via rpc rtx 3080 16gb
- Amazing! Rtx 4080 16GB, 32GB RAM.