cd /news/large-language-models/let-s-dig-into-freetoken-an-llm-engi… · home topics large-language-models article
[ARTICLE · art-110421] src=forum.level1techs.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Let's dig into FreeToken : An LLM Engine to Peg the BW of All-The-Things

Researchers from UT Austin and UC Berkeley released FreeToken, an edge-native Mixture-of-Experts (MoE) serving engine that adapts model execution to available hardware bandwidth, enabling a 753B GLM-5.2 model to run on a single workstation GPU and a 284B model on a gaming desktop. The engine supports over 20 MoE models and real coding and tool-using agents on hardware ranging from an 8GB laptop GPU to a workstation GPU, and is currently Nvidia-only.

read1 min views5 publishedAug 25, 2026

UT Austin and UC Berkeley have teamed up to find the bottlenecks of running LLM on consumer hardware, and have released FreeToken.

FreeToken supports more than 20 MoE models and real coding and tool-using agents across hardware ranging from an 8GB laptop GPU to a single workstation GPU. More importantly, it changes what these machines can practically serve, from a 35B model on a laptop to a 284B model on a gaming desktop and the 753B GLM-5.2 on a single workstation GPU.

Source: [[2608.16157] FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution](https://arxiv.org/abs/2608.16157)

From my initial reading, it seems to be a way to analyze a machine’s capabilities and adjust a model’s layers accordingly to simultaneously maximize all available bandwidth between CPU, Memory, PCIe lanes, and GPU.

Markettechposts has made a fantastic little widget to visualize how it breaks a model up according to your hardware specs:

Well worth the read for the details: Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU - MarkTechPost

Seems to be only for Nvidia at this time. I’m trying to find out what work is being done for AMD.

Lastly, give a Hat Tip to DigitalSpaceport as I found this through casual background viewing. This is worth a watch (WRX80, 3945WX, DDR4 3200 hindered at 89 GB/s + Nvidia 3090 - running DeepSeek Flash V4).

FreeToken Deepseek V4 Flash on a Single 3090 Local AI Testing I have a workstation laptop with an A5500 (aka 3080 Ti, but with 16GB ram) and an AM5 with an 4070 Super (12GB) to try. I’ll post details as I experiment.

── more in #large-language-models 4 stories · sorted by recency
── more on @ut austin 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/let-s-dig-into-freet…] indexed:0 read:1min 2026-08-25 ·