Cloud and Local, One System A Stanford study of local AI found that intelligence per watt across more than twenty local models and eight accelerators improved 5.3-fold from 2023 to 2025, while the share of single-turn chat and reasoning queries answered correctly by the best local model of each year rose from 23.2% to 71.3%. The study's authors argue the cloud will keep work that needs frontier models while capable agents on local machines take on tasks that live next to a user's files and tools, citing Epoch AI's finding that a single top gaming GPU under $2,500 runs open models matching the frontier of six to twelve months earlier. Counterpoint Research expects PCs with a neural processing unit to pass half of global shipments in 2026. Computing has done this before computing-has-done-this-before From 1946 to 2009, the energy efficiency of computing doubled about every eighteen months. That curve is what moved work from the mainframe to the personal computer. As a recent Stanford study of local AI puts it, the move happened when efficiency let personal devices meet people’s needs within their own power budget, not when PCs beat mainframes in raw performance. 1 user-content-fn-1 The same curve is now visible in AI. The Stanford team measured intelligence per watt, task accuracy per unit of power, across more than twenty local models and eight accelerators. From 2023 to 2025 it improved 5.3-fold, and the share of single-turn chat and reasoning queries answered correctly by the best local model of each year rose from 23.2% to 71.3%.