cd /news/artificial-intelligence/chutes-dominates-onchain-inference-r… · home topics artificial-intelligence article
[ARTICLE · art-123568] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Chutes dominates onchain inference requests as Surplus sees rapid growth

Decentralized AI inference platform Chutes, built on Bittensor's Subnet 64, has processed over 34 trillion tokens cumulatively with daily peaks exceeding 120 billion tokens, while Surplus Intelligence, launched in late May 2026, has grown to 2.5 million requests and 107 billion input tokens by late July 2026, a roughly 50-fold increase in two months. Both platforms offer lower prices than centralized providers, with Surplus reporting user savings of 40-60%, signaling a shift toward cheaper, onchain AI compute.

read3 min views1 publishedSep 8, 2026
Chutes dominates onchain inference requests as Surplus sees rapid growth
Image: Cryptobriefing (auto-discovered)

Decentralized AI inference platforms are processing trillions of tokens at a fraction of centralized pricing, reshaping how AI workloads get served.

The AI inference market is quietly being rewired from underneath. Two decentralized platforms, Chutes and Surplus Intelligence, are pulling an increasing share of onchain inference requests by doing something deceptively simple: making it cheaper to run AI models.

Chutes, built on Bittensor’s Subnet 64, has processed over 34 trillion tokens cumulatively, with daily peaks exceeding 120 billion tokens. Surplus Intelligence, which launched just two months ago, has already spiked to 2.5 million requests and 107 billion input tokens. Together, they represent a growing thesis that centralized AI providers are charging too much for what is, at its core, commodity compute.

How Chutes became the default #

Chutes launched publicly in late January 2025 as a decentralized serverless inference platform. Independent miners with spare graphics cards serve open-source AI models on demand, and users pay per token processed.

The pricing tells the story. Chutes charges single-digit cents per million tokens for certain models, a fraction of what centralized alternatives typically command for equivalent workloads.

After monetization kicked in post-launch, daily token throughput climbed to that 120 billion figure. The platform uses permissionless GPU miners who are scored on capacity, speed, and availability, creating a competitive marketplace where the best-performing hardware wins more jobs.

Security gets handled through Trusted Execution Environments, or TEEs. These are hardware-level isolation zones that keep data private even from the miners processing it. Hardware verification runs through something called GraVal technology, which confirms that miners are actually running the GPUs they claim to be running.

All settlements happen onchain in USDC, with no invoicing cycles or payment terms.

Surplus Intelligence’s explosive two months #

Surplus Intelligence debuted in late May 2026 as a marketplace for unused AI inference capacity, an order book where buyers and sellers of compute meet directly.

In late May 2026, Surplus handled 47,000 requests and 1.7 billion input tokens. By late July 2026, those figures had ballooned to 2.5 million requests and 107 billion tokens — roughly a 50-fold increase in two months.

Surplus reports user savings of 40-60% compared to traditional inference providers. The marketplace model enables dynamic price discovery, meaning prices adjust based on actual supply and demand. Like Chutes, Surplus settles in USDC onchain.

Where the two platforms differ is their approach to supply. Chutes recruits GPU miners into a permissionless network scored by performance metrics. Surplus operates more like a secondary market, connecting entities that have excess inference capacity with those who need it.

What this means for the AI compute market #

For the Bittensor ecosystem specifically, Chutes’ position on Subnet 64 reinforces TAO’s utility narrative. When a subnet is processing 120 billion tokens daily, the token economics become grounded in real throughput rather than speculation about future adoption. The risk with decentralized infrastructure is reliability. Centralized providers offer SLAs, dedicated support, and guaranteed uptime. Permissionless miner networks offer lower prices and censorship resistance, but a miner going offline mid-inference is a different failure mode than an AWS region going down.

These platforms also primarily serve open-source models. If your workload requires proprietary frontier models from OpenAI or Anthropic, decentralized inference isn’t an option yet.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @chutes 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/chutes-dominates-onc…] indexed:0 read:3min 2026-09-08 ·