Chutes AI trains an 8B model on distributed gaming GPUs, runs it on a phone CPU Chutes AI trained an 8 billion parameter model on 240 RTX 5090 consumer GPUs spread across 30 hosts in 13 countries for approximately $6,500, or about $11 per billion tokens processed, Jon Durbin said in a presentation of the Parallax system at the Exploit Summit in Montreal on September 28-29, 2026. The finished 8B model runs on a phone CPU at approximately 59.6 tokens per second with no cloud connection, and a 40.75 billion parameter variant ran at 26.9 tokens per second using 12GB of peak memory. Chutes AI, which operates as Bittensor subnet SN64, says it uses the libp2p peer-to-peer framework for syncing between machines and plans to publish a tech report and the full model. Photo: KyoRa Kee / Pexels Chutes AI trains an 8B model on distributed gaming GPUs, runs it on a phone CPU Jon Durbin showed off Parallax at the Exploit Summit, a decentralized training system built on 240 consumer graphics cards across 13 countries Training an AI model usually conjures images of warehouse-sized data centers humming with specialized chips. Chutes AI just did it with the same graphics cards people buy to play video games. At the Exploit Summit in Montreal, held September 28-29, 2026, Jon Durbin presented an 8 billion parameter model trained on distributed consumer GPUs. The finished model runs on a phone’s CPU. What Chutes AI actually built The system is called Parallax, and it is Chutes AI’s approach to decentralized model training. Instead of renting one giant cluster, Parallax stitches together hardware scattered across the globe. For this run, the setup used 240 RTX 5090 GPUs spread across 30 hosts in 13 countries. The price tag is the headline number. Chutes AI put the training cost at approximately $6,500, or about $11 per billion tokens processed. The 8B model reportedly hits approximately 59.6 tokens per second on mobile CPUs, generating text on a phone processor with no cloud connection required. AI, tech, and the markets they move—in one daily briefing. Daily. Free. Join 34,000+ readers across crypto, finance, and policy. Chutes AI also showed a larger variant with 40.75 billion parameters. That version ran at 26.9 tokens per second while using 12GB of peak memory. The privacy and resilience pitch Local inference was a central selling point of the demo. Because the models run directly on the device, no prompts were sent off the user’s hardware. Chutes AI says it tackles distributed training challenges with architectures like libp2p, a peer-to-peer networking framework used for efficient syncing between machines. Where Bittensor fits in Chutes AI operates as Bittensor https://cryptobriefing.com/markets/bittensor/ subnet SN64, focused on serverless decentralized compute for open-source AI models. SN64’s job is providing compute, and Parallax extends that mission from running models to training them. Chutes AI previously ran 20B-scale experiments across a range of GPUs, including H100s and RTX 6000 models. The current cluster was built entirely on RTX 5090 consumer cards. The company plans to publish a comprehensive tech report and the full model. Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .