tokens too cheap to meter The cost of completing a given task with a machine learning model is falling sharply even though per-token prices for frontier models are not consistently declining, according to an analysis citing Epoch.AI data. GPU power efficiency is doubling roughly every two years, a logarithmic rate of 1.3 that the analysis says has not been seen since Moore's Law in the 1960s, and Artificial Analysis pareto-frontier charts show models getting smarter and cheaper per task across 2025. The author predicts LLMs will be integrated into every part of computing as infrastructure within one to two years and will run locally at current frontier quality on commodity hardware within three to six years, making quality and access rather than token volume the limiting factor. tokens too cheap to meter The price of using machine learning intelligence is decreasing by several orders of magnitude a year and shows no signs of slowing. We are likely to see LLMs integrated into every part of computing as infrastructure, not just as a product, in the next year or two. We are likely to see LLMs running locally at current frontier-quality on commodity hardware in the next 3-6 years. Starting very soon, we are likely to see quality and access become the limiting factor to AI