{"slug": "mainframes-became-personal-so-will-your-data-center", "title": "Mainframes became personal. So will your data center.", "summary": "Local AI models can now answer 89% of everyday chat and reasoning queries as well as frontier cloud models, according to a November 2025 study by Stanford University and Together AI. The best local model's win/tie rate against frontier models rose from 23.2% in 2023 to 71.3% in 2025, and intelligence-per-watt increased 5.3x, driven by a 3.1x gain from better models and a 1.7x gain from better chips. The study concludes that local models plus a router can handle the supermajority of everyday knowledge work, cutting energy use by 80%, compute by 77%, and cost by 74% versus an all-cloud baseline.", "body_md": "Local AI models can already answer 89% of everyday chat & reasoning queries as well as a frontier cloud model, a result that holds broadly across more than a million real queries & 20+ local models tested.[1](#fn:1)\n\nThat means we’re generating more intelligence per watt of electricity : more work from the same number of electrons.[1](#fn:1)\n\nTracking computing efficiency over time is not new. Koomey’s law found that computing power per watt doubled roughly every 1.5 years for decades, a trend that shrank the power of a mainframe into a laptop’s chassis.[2](#fn:2)[3](#fn:3)\n\nJust as performance-per-watt guided the mainframe-to-PC transition, intelligence-per-watt will guide AI’s transition to the edge.\n\n— Jon Saad-Falcon, Avanika Narayan, et al., “Intelligence per Watt”\n\nGPUs follow a more languid curve today, doubling efficiency roughly every 2.7 years over the last 15 years, not every 1.5 years.[4](#fn:4)\n\nThe impact is no less impressive : the best local model’s win/tie rate against a frontier model, rose from 23.2% in 2023 to 71.3% in 2025, adding roughly 20 percentage points a year. In 2026, a local model selected for the task by another local AI bumps that number to nearly 90%.[5](#fn:5)\n\nEfficiency improved alongside the trend line : intelligence-per-watt rose 5.3x over the same period, split into a 3.1x gain from better models & a 1.7x gain from better chips. Computers are faster, models are smarter. And the end user benefits.\n\n[5](#fn:5)Cloud remains essential for long multi-step reasoning, the hardest technical domains, & workloads where scale & parallelization matter. Cloud hardware still holds an edge over local hardware in many use cases. Cloud inference delivers a 40% energy efficiency gain relative to local models. 6 Datacenters batch queries, a trick local hardware serving one user at a time cannot use, yet.\n\nBut for much of everyday knowledge work, there is no reason to send the query to the data center at all. The combination of local models plus a router is sufficient for the supermajority of work.[7](#fn:7)\n\nIt also cuts energy 80%, compute 77%, & cost 74% against an all-cloud baseline.\n\nMainframes became personal. So will your data center.\n\n-\nJon Saad-Falcon, Avanika Narayan, et al., “Intelligence per Watt: Measuring Intelligence Efficiency of Local AI,” Stanford University & Together AI, November 2025.\n\n[arXiv:2511.07885](https://arxiv.org/abs/2511.07885). See also the[Stanford Hazy Research overview](https://hazyresearch.stanford.edu/blog/2025-11-11-ipw).[↩︎](#fnref:1)[↩︎](#fnref1:1) -\n[Koomey’s law](https://en.wikipedia.org/wiki/Koomey%27s_law), Wikipedia.[↩︎](#fnref:2) -\nWatts measure power, the rate energy is drawn, & joules measure the energy itself; Koomey’s original metric was computations per joule, but the underlying trend is the same one intelligence per watt now tracks for AI models.\n\n[↩︎](#fnref:3) -\nAnson Ho, Ege Erdil & Tamay Besiroglu, “Limits to the Energy Efficiency of CMOS Microprocessors,” 2023.\n\n[arXiv:2312.08595](https://arxiv.org/abs/2312.08595).[↩︎](#fnref:4) -\n71.3% & 89% are two different measurements from the same 2025 data, not the same number at different times. 71.3% is the best single local model’s win/tie rate against a frontier model, the trend line this chart plots (23.2% in 2023, 48.7% in 2024, 71.3% in 2025). 89% is a separate, higher ceiling: routing each query to whichever of the 20+ local models tested handles it best beats any single model by 16.3 to 28.8 percentage points. That gain is a selection effect: the paper notes local routing draws from 20+ diverse models versus three frontier cloud models, so on some benchmarks the best-of-local ensemble even surpasses best-of-cloud. More candidates to choose from, not just smarter routing, is what raises accuracy. Both figures come from the same study & neither supersedes the other.\n\n[↩︎](#fnref:5)[↩︎](#fnref1:5) -\nJon Saad-Falcon, Avanika Narayan, et al., “Intelligence per Watt: Measuring Intelligence Efficiency of Local AI,” Stanford University & Together AI, November 2025. Cloud accelerators deliver at least 1.4x higher intelligence-per-watt than local chips running the same models, roughly a 40% efficiency premium.\n\n[arXiv:2511.07885](https://arxiv.org/abs/2511.07885).[↩︎](#fnref:6) -\n[Most AI Work Can Wait](https://tomtunguz.com/ai-execution-routing/), tomtunguz.com.[↩︎](#fnref:7)", "url": "https://wpnews.pro/news/mainframes-became-personal-so-will-your-data-center", "canonical_source": "https://www.tomtunguz.com/intelligence-per-watt/", "published_at": "2026-08-21 00:00:00+00:00", "updated_at": "2026-08-21 19:43:25.325191+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-infrastructure"], "entities": ["Stanford University", "Together AI", "Jon Saad-Falcon", "Avanika Narayan"], "alternates": {"html": "https://wpnews.pro/news/mainframes-became-personal-so-will-your-data-center", "markdown": "https://wpnews.pro/news/mainframes-became-personal-so-will-your-data-center.md", "text": "https://wpnews.pro/news/mainframes-became-personal-so-will-your-data-center.txt", "jsonld": "https://wpnews.pro/news/mainframes-became-personal-so-will-your-data-center.jsonld"}}