{"slug": "chutes-ai-and-harvard-release-public-dataset-of-6-12-billion-llm-requests", "title": "Chutes AI and Harvard release public dataset of 6.12 billion LLM requests", "summary": "Chutes AI, a decentralized inference platform on Bittensor's Subnet 64, and Harvard University researchers released a public dataset of 6,122,413,756 LLM requests spanning 9,174 models from 314,970 anonymized users between April 11, 2025, and April 12, 2026. The dataset, available via GitHub and a Harvard S3 bucket, contains only serving metadata — request timing, token counts, latency, and time-to-first-token — with no prompts or responses, and shows that 99% of repeat requests occur within a 15-minute window. Researchers say prefix-aware routing can reach near-optimal cache-hit rates with slight server load imbalance, while output lengths fell from hundreds of tokens to fewer than 100, which they attribute to a shift toward agentic or machine-driven queries.", "body_md": "Photo: Tima Miroshnichenko / Pexels\n\n# Chutes AI and Harvard release public dataset of 6.12 billion LLM requests\n\nThe year-long dataset spanning 9,174 models reveals that 99% of requests are repeats within 15 minutes, a finding with major implications for how AI infrastructure gets built.\n\nIf you’ve ever wondered what billions of AI requests actually look like under the hood, now you can find out. Chutes, a decentralized AI inference platform running on [Bittensor](https://cryptobriefing.com/markets/bittensor/)’s Subnet 64, has teamed up with Harvard University researchers to release what may be the largest public dataset of LLM serving metadata ever assembled.\n\nThe numbers are staggering: 6,122,413,756 requests across 9,174 models, generated by 314,970 anonymized users over a full year of production traffic. The dataset spans from April 11, 2025, to April 12, 2026, and is now freely available through GitHub and a Harvard S3 bucket.\n\n## What’s actually in this thing\n\nThe dataset captures metadata, not the actual conversations people had with AI models. It includes request timing, token counts, latency measurements, and time-to-first-token (TTFT) metrics, but zero prompts or responses. Over the year-long period, Chutes processed roughly 35.8 trillion input tokens and 2.52 trillion output tokens.\n\nUser identifiers in the dataset rotate every three months, adding another layer of anonymization. The accompanying documentation, co-authored by researchers from Harvard, the University of Chicago, and Chutes, provides the kind of detailed methodology notes that make the data actually usable for academic work.\n\n## The 99% repeat problem\n\nPerhaps the most consequential finding buried in this dataset: 99% of repeat requests occur within a 15-minute window. The research team found that prefix-aware routing strategies can achieve nearly optimal cache-hit rates with only slight load imbalances across servers.\n\n### AI, tech, and the markets they move—in one daily briefing.\n\nDaily. Free. Join 34,000+ readers across crypto, finance, and policy.\n\nOutput lengths have been declining over the dataset’s timespan, dropping from hundreds of tokens per response to fewer than 100. The researchers interpret this as evidence of a shift toward agentic or machine-driven queries, where automated systems are making quick, targeted requests rather than humans asking for lengthy explanations.\n\n## The Bittensor connection\n\nChutes operates as part of Bittensor’s decentralized network, specifically Subnet 64, running open-source LLMs on a distributed GPU infrastructure. Payments on the platform flow through TAO, Bittensor’s native token.\n\nThe collaboration between Chutes and Harvard included an opt-in period from March to July 2026 where researchers received a 25% discount for contributing their usage data to the dataset.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/chutes-ai-and-harvard-release-public-dataset-of-6-12-billion-llm-requests", "canonical_source": "https://cryptobriefing.com/chutes-harvard-6-billion-llm-requests-dataset/", "published_at": "2026-09-21 21:39:53+00:00", "updated_at": "2026-09-21 21:53:16.149126+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-research", "ai-agents"], "entities": ["Chutes AI", "Harvard University", "Bittensor", "Subnet 64", "University of Chicago", "TAO", "GitHub"], "alternates": {"html": "https://wpnews.pro/news/chutes-ai-and-harvard-release-public-dataset-of-6-12-billion-llm-requests", "markdown": "https://wpnews.pro/news/chutes-ai-and-harvard-release-public-dataset-of-6-12-billion-llm-requests.md", "text": "https://wpnews.pro/news/chutes-ai-and-harvard-release-public-dataset-of-6-12-billion-llm-requests.txt", "jsonld": "https://wpnews.pro/news/chutes-ai-and-harvard-release-public-dataset-of-6-12-billion-llm-requests.jsonld"}}