Run 290B+ frontier MoE models locally on your gaming PC FlashML released FreeToken, an edge-native Mixture-of-Experts (MoE) serving engine that runs 290B+ parameter frontier MoE models locally on consumer gaming PCs at interactive speeds. The engine supports heterogeneous edge resources, semantic-aware caching, and elastic memory management, with compatibility for NVIDIA RTX 30, RTX 40, and RTX 50 series GPUs. FreeToken is available for download on Windows and Linux at flashml.ai, with installation via uv or pip. | Download https://www.flashml.ai/ | | https://arxiv.org/abs/2608.16157 Paper | https://join.slack.com/t/flashml/shared invite/zt-3zpdh5j10-9dwTXrgLiqpVxizhA9KVbA Developer Slack | https://discord.gg/xzwSnMdsX Community Discord | https://github.com/FlashML-org/FreeToken/blob/main/assets/freetoken-wechatgroup.png Community WeChat Unlock datacenter-class intelligence on the hardware you already own — Run 290B+ frontier MoE models locally on your gaming PC at blistering interactive speeds. FreeToken is an edge-native Mixture-of-Experts MoE serving engine designed for running frontier-scale open-weight models on personal and consumer hardware. It treats heterogeneous edge resources—GPUs, CPUs, host memory, and interconnects—as a unified, elastic inference platform. Its core features include: - Fast Edge-Native Runtime : Provides efficient MoE serving with bandwidth-adaptive CPU–GPU co-execution $q^\star$ policy , full-layer double-buffered prefill streaming, global LRU expert caching, graph-compatible execution, and the FTW fast weight format. - Semantic-Aware Caching : Features semantic anchor checkpoints for recurrent state and KV caches, allowing agentic context edits e.g., tool calls, thinking blocks to avoid redundant context recomputation. - Elastic Memory Management : Supports dynamic, runtime VRAM re-allocation between expert caches and KV memory without engine restarts or weight reloading. - Broad MoE & Ecosystem Support : Supports frontier open-weight MoE models e.g., DeepSeek-V4-Flash, Qwen3.6-35B-A3B, GLM-5.2 across various parameter scales and quantization formats e.g., MXFP4, NVFP4, FP8, BF16 , with Anthropic/OpenAI-compatible APIs for seamless integration with real-world coding and tool-calling agents e.g., Codex, Claude Code, OpenCode, OpenClaw, DeepSeek Harness . - Diverse Consumer Hardware : Scales across consumer laptops, gaming desktops, and workstation GPUs, with native support for NVIDIA RTX 30, RTX 40, and RTX 50 series GPUs. Download FreeToken for Windows or Linux at flashml.ai https://www.flashml.ai/ . It sets the engine up for you and gives you a GUI for running models, chatting, and tuning the engine. Install FreeToken with uv https://docs.astral.sh/uv/ recommended or pip: uv pip install "freetoken accel " Or build from source: git clone https://github.com/FlashML-org/FreeToken.git && cd FreeToken uv venv && source .venv/bin/activate uv pip install -e ". accel " For More details: If you use FreeToken for your research, please cite our paper https://arxiv.org/abs/2608.16157 : @article{yang2026freetoken, title={FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution}, author={Yang, Shuo and Fan, Xiaoze and Pan, Melissa and Xi, Haocheng and Wang, Zhe and Sun, Shanlin and Keutzer, Kurt and Han, Song and Zaharia, Matei and Xu, Chenfeng and Stoica, Ion}, journal={arXiv preprint arXiv:2608.16157}, year={2026} } FreeToken was deeply inspired by mini-sglang https://github.com/sgl-project/mini-sglang , and learned the design and reused code from the following projects: SGLang https://github.com/sgl-project/sglang , vLLM https://github.com/vllm-project/vllm , FlashInfer https://github.com/flashinfer-ai/flashinfer , flash-linear-attention https://github.com/fla-org/flash-linear-attention , LightLLM https://github.com/ModelTC/lightllm and llama.cpp https://github.com/ggml-org/llama.cpp .