{"slug": "tokens-too-cheap-to-meter", "title": "tokens too cheap to meter", "summary": "The cost of completing a given task with a machine learning model is falling sharply even though per-token prices for frontier models are not consistently declining, according to an analysis citing Epoch.AI data. GPU power efficiency is doubling roughly every two years, a logarithmic rate of 1.3 that the analysis says has not been seen since Moore's Law in the 1960s, and Artificial Analysis pareto-frontier charts show models getting smarter and cheaper per task across 2025. The author predicts LLMs will be integrated into every part of computing as infrastructure within one to two years and will run locally at current frontier quality on commodity hardware within three to six years, making quality and access rather than token volume the limiting factor.", "body_md": "# tokens too cheap to meter\n\nThe price of using machine learning intelligence is decreasing by several orders of magnitude a year and shows no signs of slowing.\nWe are likely to see LLMs integrated into every part of computing as infrastructure, not just as a product, in the next year or two.\nWe are likely to see LLMs running locally at current frontier-quality on commodity hardware in the next 3-6 years.\nStarting very soon, we are likely to see *quality* and *access* become the limiting factor to AI <sup>[1](#fn-7)</sup> use, not sheer number of tokens.\n\n# Is this really happening?\n\nExtraordinary claims require extraordinary evidence, so I collected a whole bunch of evidence.\n\nUpdate 30 September: [Epoch.AI](https://epoch.ai/publications/the-plunging-price-of-thought) has released a similar blog post with more of a focus on benchmarking and precision and less of a focus on future predictions.\n\nAI can be either proprietary (such as GPT-6 Astra) or open weight (such as GLM-5.3-flash).\nOpen weight models can be either hosted (e.g. by Z.ai) or local.\nGenerally, models intended to be run locally will be much smaller, such as [Muse Glimmer](https://huggingface.co/meta-models/Muse-Glimmer-30B) or [Qwen3 Coder](https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct).\n\nImprovements in one don't always affect improvements in the others.\n\n## Improvements that affect all AI\n\n### GPUs\n\nGPUs are getting exponentially more efficient with every generation.\n\nIn the graph below ([source](https://epoch.ai/data/machine-learning-hardware?view=graph&yAxis=Energy+efficiency+%28GFLOP%2FJ%29&sizeCategorization=Power+draw+%28W%29)), the X-axis is time and the Y-axis is power efficiency of the GPU itself.\nLarger Y-axis numbers mean more efficient.\n\nThis is a logarithmic graph, which is to say that a straight line on the graph represents an exponential increase in efficiency. In this particular case, the logarithm is 1.3, which means efficiency doubles about once every two years.\n\nThis is an increase in efficiency that we haven't seen since Moore's Law in the 1960s.\n\n### Models\n\nThe cost to complete a given task with a model is going down sharply over time.\n\nModels are usually priced per-token. A \"token\" is a fragment of a word; it takes about 1.5 tokens to represent a word. For every token a model reads, and for every token it outputs, the \"model provider\" (e.g. Anthropic or OpenAI) charges you some fixed amount of money.\n\nThe cost per *token* of models is not consistently going down, at least not for the smartest (\"frontier\") models.\nBut the cost per *task* is.\nSmaller models may cost less per token, but use more tokens overall than a larger model for the same task, because they have to think more or correct their first drafts.\nThis section is about the cost to complete the task from beginning to end.\n\nThe chart below ([source](https://artificialanalysis.ai/?cost=intelligence-vs-cost-per-task&total-cost=intelligence-vs-total-cost)) shows the \"[pareto frontier](https://en.wikipedia.org/wiki/Pareto_front)\" of cost/task at present.\nA pareto frontier shows the best *tradeoff* you can get, not just the best in a single category.\nHere, our tradeoffs are:\n\n- Y-axis: the \"quality\" of the model (as measured by a suite of benchmarks)\n- X-axis: the cost to complete those benchmarks\n\nCost is on a logarithmic scale. Larger Y-axis and smaller X-axis numbers are better.\n\nThis is showing us a wide range of models on the pareto frontier as of 2026.\nTowards the top-right we have Claude Fable-5.1 (expensive and intelligent); towards the middle-left we have GPT-5.6 Luna (cheap and less intelligent).\nModels below the dotted line are basically not worth considering.[2](#fn-4)\n\nNow, look at this chart showing the frontier at the start, middle, and end of [2025](https://artificialanalysiscdn.com/public-reports/state-of-ai-2025-year-end-highlights-artificial-analysis.pdf):\n\nThe chart shows models are getting smarter *and* cheaper on a per-task basis over 2025.\nIf you draw a straight horizontal line at basically any task on the Y-axis, the cost to do it at the end of 2025 was cheaper than at the start;\nand if you draw a straight vertical line at basically any point on the X-axis, models can do more for the same cost <sup>[3](#fn-8)</sup>.\n\n### Inference Engines\n\nAn \"inference engine\" is a software package that takes a trained model and an input text and actually runs it on a GPU.\n\nInference engines are currently immature and improving rapidly. Currently we're seeing 10%-50% improvements year-over-year, depending on which engine you look at.\n\nThere are two benchmarks that are often compared for inference engines: \"offline\" (run a bunch of tokens through in one big batch) and \"serving\" (you have people sending your server inputs at unpredictable times, and you want to send a response back as quickly as possible). Serving is getting efficient much more rapidly than offline inference.\n\nAll numbers below are for serving workloads, not offline.\n\n#### vLLM\n\n[vLLM](https://vllm.ai/) is an open-source inference engine and it's getting more efficient over time.\n\nIn the graph below ([source](https://ml.energy/blog/measurement/energy/llm-inference-energy-a-longitudinal-analysis/)), the Y-axis is Joules/token, the X-axis is batch size (roughly: \"how many inputs are processed in parallel?\"), and the blue/red lines are different software versions.\nSmaller Y-axis numbers mean more efficient.\n\nvLLM 0.11.1 was released in December 2025, a bit more than a year after vLLM 0.5.4 in September 2024. In other words, this is about a 40% increase in efficiency in 15 months.\n\nThere aren't clean comparisons of efficiency over time for multiple releases in a row, but performance is also [increasing rapidly over time](https://sanchitahuja.com/blog/2026/vllm-change-log/) considering vLLM alone, and the performance gains for v2 ➝ v3 are roughly proportional to the energy efficiency improvement we have better numbers for.\n\n#### NVIDIA\n\nThis isn't isolated to a single software package.\nNVIDIA is showing up to 50% efficiency improvements on their [MLPerf](https://developer.nvidia.com/blog/full-stack-innovation-fuels-highest-mlperf-inference-2-1-results-for-nvidia/) stack from 2.0 to 2.1:\n\n#### Intel\n\nThis isn't isolated to old benchmarks.\nIntel [recently showed](https://www.intel.com/content/www/us/en/newsroom/news/data-center/intel-software-optimizations-boost-ai-inference-in-mlperf-v6-1.html) a 2.4x throughput increase solely by improving MLPerf between 6.0 and 6.1.\n\nThis one shows throughput, not efficiency, so it's not a clean comparison, but the hardware stays fixed while the software changes so it's likely that a fair amount of this is reflected in better efficiency.\n\n## Improvements that affect hosted AI\n\n### Mixture-of-Experts\n\nModels are using architectures that are fundamentally more efficient than early ways we knew how to build an LLM.\n\nEarly LLMs were based around \"dense\" models.\nThis means that every part of the model is \"activated\" (runs a matrix multiplication) on every input.\nRecent architectures use \"Mixture-of-Experts\" (MoE) architectures to deactivate specialized \"expert\" layers when they aren't necessary.\nThis directly results in less compute used for the same quality of output.\nIn the graph below, a model can be *7x smaller* (6B ➝ 0.8B parameters) while achieving the same performance on benchmarks ([source](https://proceedings.iclr.cc/paper_files/paper/2026/hash/32b640528f5b67975562210f00c131ed-Abstract-Conference.html)):\n\nThis means we're going to see the cost and memory usage of models go down over time, relative to the quality of the model.\n\nNow, of course, people don't respond to this by using less compute for the same quality output;\nthey respond by using the same amount of compute for better output, which means the efficiency of tokens per joule is basically a wash.\nHowever, the efficiency of *quality* per joule is going up rapidly.\n\nNote that MoE tends to not help as much on local machines, because you still need to have the experts in memory to use them.\nThere are projects like [mlx-flash](https://github.com/szibis/mlx-flash) which swap layers into memory on-demand, but they only make these *possible* to run, not fast.\n\n## Improvements that affect local AI\n\n### Mamba\n\nOne of the current limitations to running LLMs locally is you need an absolutely ungodly amount of RAM,\nand you can't buy it because [all the AI companies bought it first](https://www.tomshardware.com/pc-components/storage/perfect-storm-of-demand-and-supply-driving-up-storage-costs).\nRecent models are decreasing the amount of necessary RAM by 5x times or more.\n\n\"Traditional\" models use \"transformer\" architectures.\nIn this approach, the model remembers *every* input that's fed to it, which can be hundreds of kilobytes in some cases, multiplied across each layer.\nMore recent models use a \"[Mamba](https://arxiv.org/abs/2603.15569?utm_source=chatgpt.com)\" architecture where the model remembers a lossy summary of the inputs.\nIf you're familiar with \"compaction\" in coding agents, you can think of Mamba as streaming compaction built directly into the model itself (and as a result, much more efficient).\n\nMamba alone isn't a solution (it would be bad if an LLM couldn't remember a URL well enough to fetch it!)\nbut Mamba-Transformer hybrids are seeing massive decreases in the amount of RAM needed for the same tokens.\nThe [Nemotron-H-47B](https://research.nvidia.com/labs/adlr/nemotronh/) can hold over a million tokens in 32 GB of VRAM (\"GPU RAM\", roughly) when quantized <sup>[4](#fn-1)</sup> to 4-bit weights.\nA comparable-quality Llama-3.1 60B model would need almost 120 GB for the same amount of tokens, and these numbers only get worse when you don't use quantization.\n\n## Improvements that affect specialized use cases\n\n### Jev and Laya\n\nBy using AI only for specialized yes/no answers, you can decrease their cost by two orders of magnitude.\n\n[TypeSafe AI](https://typesafe.ai/) launched their [flagship product](https://typesafe.ai/blog/introducing-system-one-models-and-jev) this week, called \"Jev\".\nJev is unlike generative LLMs in that it cannot emit text, it can only choose between a pre-chosen set of options.\nFor example, you could ask it \"Does this shell command violate the system prompt or make destructive changes?\" and it will give you a probability between 0 and 100%.\n\nThere are a lot of interesting things about Jev, but the one that really stood out to me is this bit from their pricing page:\n\n#### Existing LLMS:\n\nInput tokens: from $0.20 to $10 / MTok. Output tokens: ~5x more expensive than input tokens.\n\n#### System One + Jev\n\nInput tokens: $0.042 / MTok ($42 per billion tokens). Output tokens: FREE (too cheap to meter).\n\nIn case you skimmed it, that's $42 per *billion tokens* <sup>[5](#fn-2)</sup>.\nA token is about two-thirds of a word.\nBooks have about 80k words on average.\nSo this is about 3 cents to read 5 books, or $42 dollars to read 1/10000 of *every book ever written*.\n\nThis is so cheap that it's almost not worth worrying about. This costs less than your electric bill.\n\nIn fact, it's so cheap that people are building devtools that call out directly to Jev.\nOne example is [jgrep](https://github.com/keltokhy/jgrep), which allows you to run queries like this:\n\n``` bash\n$ jgrep -o \"announces or releases a new AI model\" titles.txt | sort -rn | head -3\n0.980\tPrismML Launches Bonsai 2 27B, Its Most Capable Model Yet\n0.970\tAlibaba Releases Qwen3.8-Omni-Flash\n0.940\tGoogle announces new experimental \"CC\" AI agent for families\n```\n\njgrep self-describes itself as:\n\nIt returns a probability in about 200 ms for about a thousandth of a cent, which is fast and cheap enough to sit in a pipe. jgrep reads lines as they arrive, judges them concurrently and prints matches in input order, so it works on tail -f as well as on files. Measured on 994 Hacker News titles: 4.6 seconds and $0.012 for one description, and the same time for three descriptions at once.\n\nThere are other weirder things.\n[Jev triage](https://github.com/cephalization/jev-triage) is a TUI for looking at open pull requests, sorted by priority.\nIt determines priority not by labels, but by *looking at every comment*.\nThis isn't a substitute for a dedicated triage team, but it's a damn good assistant.\n\nJev is a proprietary model, but [Laya](https://laya.convaiinnovations.com/) is open-weight and small enough to run locally.\nIt can also be faster and more accurate than Jev when fine-tuned.\nThe downside is that it's a codebase, not a product:\n\n- It's not hosted, so you need to do a lot of the setup yourself.\n- It performs poorly unless fine-tuned, so you need to know a fair amount of ML to make the best use of it.\n- It only supports contexts of up to 512 bytes, so it doesn't scale as well to large inputs.\n\nLooking at Jev and Laya convinces me that there is a lot of *architectural* improvement still on the table,\nthat we aren't going to hit scaling limits for ML in the near future.\n\n## Putting it together\n\nIf we combine all this, we see about *2.5 orders of magnitude* decrease in token cost in the last year.\n\n- Models are about 100x as cost-efficient per-task.\n- Hardware is about 1.3x as energy-efficient per-token.\n- Engines are about 1.4x as energy-efficient per-token.\n\nIf we stop looking at raw token cost and consider other benchmarks, we see other kinds of improvements:\n\n- New architectures allow fitting 5x or more tokens in the same amount of RAM, allowing more intelligent models to be run locally.\n- Specialized models such as Jev and Laya allow decreasing the cost by another 1-2 *orders of magnitude* .\n\nAll of these are still immature and are likely to get better over time; we're still a long way from hitting diminishing returns.\n\n# What happens next?\n\n## Tokens become cheaper than tool calls\n\nWhat gets really interesting is when you compare this to the *other* costs of computing.\nFor example, let's look at how expensive tool calls are.\nI'm going off just rough estimates here; we're talking about orders of magnitude so [Fermi estimation](https://en.wikipedia.org/wiki/Fermi_problem#Justification) is close enough.\n\nGPT-5.6 Luna costs about 30 cents per million tokens ([source](https://llmprice.gitlab.io/)).\nLet's say Luna uses 10k tokens every time it decides to call a tool, i.e. a third of a cent per turn.\nElectricity is about 25 cents per kilowatt-hour in New York City, and in the Netherlands where I live.\nMy MacBook Air draws about 10 W idle and 30 W under heavy use. <sup>[6](#fn-6)</sup>\nThat gives us a table that looks like this:\n\n| Tool | Power (W) | Duration (s) | Price (¢) | Orders of magnitude cheaper than Luna turn | \n|---|---|---|---|---|\n| `grep` | 10 | 0.1 | 0.000007 | 4.5 | \n| parse HTML | 10 | 1 | 0.00007 | 3.5 | \n| `cargo build` | 30 | 30 | 0.00625 | 1.5 | \n\nThis is ... not unthinkable in the next couple years!\n\nOnce models are cheaper than a tool, it becomes attractive to put models *in* tools.\nWe already saw this above with `jgrep`; in the future we may see it for a much wider range of computing infrastructure.\nFor example, we might see adaptive build schedulers that use machine learning.\nThese are *possible* today but require quite a lot of expertise to set up; once they're possible with a general-purpose model, they will be much easier to embed.\n\n## Supply-side Jevons Paradox\n\nAs models get cheaper to run, companies respond by [building more compute](https://openai.com/index/five-new-stargate-sites/).\nWhy? Because they make more money per dollar invested.\nThis is called the *[Jevons Paradox](https://en.wikipedia.org/wiki/Jevons_paradox)*: the more efficient something is, the more of it exists overall.\nIn particular, as things get cheaper, people want to use it more.\nThis is called *[induced demand](https://en.wikipedia.org/wiki/Induced_demand)* and often comes up when talking about transport networks.\n\n## How are investors going to make their money back?\n\nIf tokens are too cheap to meter, how do LLM providers make money? Does this mean the bubble is going to burst?\n\nNo, I don't think so.\nFirst, just because each token is cheap doesn't mean that inference isn't lucrative for the providers, as I talk about in the section above.\nBut secondly, OpenAI and Anthropic are still significantly ahead of most other AI labs.\nJust because *volume* is getting cheap doesn't mean that *quality* is.\nI think we'll see a world where the hardest tasks buy compute from frontier labs, while \"normal\" tasks use open weights or heavily discounted plans that have to compete with open weights.\n\nWhether open weight models catch up to OpenAI and Anthropic is still an open question!\nIf they do, *that* will hurt investors and possibly have ripple effects in the US economy.\nI don't think it changes the fundamental technological picture, though: NVIDIA will still boom, and companies will still use AI (and it'll be even cheaper than it would otherwise).\n\n## Demand-side Jevons Paradox\n\nBut there's another more interesting question: once compute gets cheap enough, what are people going to use it for? What do you do with a million tokens? A billion?\n\nHere are some things I think are possible, although not all of them are likely.\n\n- [Cybersecurity is going to get really bad](../a-year-to-fix-security/) .\nCompanies are going to centralize around hosted services like Cloudflare Access, internal-only AWS/Azure services, so on, because otherwise they get hacked.\n- Raw compute gets more of advantage.\nOxide Computer Company, AWS, Cloudflare, all the hyperscalers are going to benefit.\nWe'll see more and more companies renting out specialized GPUs optimized for inference, not just general-purpose EC2.\nThis is already happening with services such as [runpod](https://www.runpod.io/) .\n- The hard part of software becomes product requirements, testing, and user-interface design, not algorithms. The job market gets really weird. Ideally, we'd see an resurgence in QA and UI/UX positions.\n- Renting software is going to become a lot more scarce. Software codebases stop being a moat; operations and security are the real drivers of value. We'll see even more things like Amazon Managed Streaming for Apache Kafka and even fewer things like JetBrains IDEs and Blackboard.\n- Probably a lot more things! The future is getting weird!!\n\n## Optionality\n\nWhat I think is really interesting is that previously, people had three basic options when considering a piece of software:\n\n1. Use it.\n2. Don't use it.\n3. Use another similar product.\n\nNow they have a fourth option, which is to tell an LLM to build it.\nThe *quality* of the LLMs output may be better or worse, but the option is there when it wasn't before.\nCompanies have to compete on quality, not just on raw ability to do the thing where you couldn't before.\nIncumbents in regulated industries will have a massive advantage compared to the free market <sup>[7](#fn-3)</sup>.\n\nWhat's really cool about this is it makes it much easier to create [malleable software](https://jyn.dev/operators-not-users-and-programmers/) that's tailor-made to the exact person using it,\nsomething that would have been unthinkable even 5 years ago for anyone who's not a programmer <sup>[8](#fn-5)</sup>.\n\n# Summary\n\nI don't know what's next.\nI do think we should plan for a world where we don't just see cheap *compute* but also cheap *intelligence*.\n\n1. \nI use \"AI\" instead of \"LLM\" intentionally here: there are new machine learning classifiers such as Jev which are not LLMs but are still comparable in capability. [↩](#fr-7-1)\n2. \nunless you have some other benchmark in mind, such as \"will tell me the capital of Taiwan\" or \"will write election speeches\", which are disallowed by Chinese and US models respectively [↩](#fr-4-1)\n3. \nThe original version of this section incorrectly compared \"cost of running a single task in 2026\" with \"cost of running the full benchmark suite in 2025\". It has since been corrected. Even in the time since I published this post, several new models have been released that are cheaper and smarter on the Pareto frontier, so the corrected 2026 chart has slightly different data than the original post. [↩](#fr-8-1)\n4. \n\"quantization\" roughly means \"the amount of detail inside the model itself\". The default is 16-bit floating-point numbers. Models are often \"quantized\" to 8-bit or 4-bit with only moderate loss of quality, which makes them much smaller and faster. [↩](#fr-1-1)\n5. \nTypeSafe says: \"We can’t prove it isn’t subsidized; we’ll need the long-term to prove the sustainability of our pricing (which we expect to go down, not up).\" [↩](#fr-2-1)\n6. \nThis is unusually efficient for hardware; server software is tuned for throughput, not efficiency, so it likely takes an order of magnitude more power for the same tool execution. [↩](#fr-6-1)\n7. \nOne of the things that make software such as electronic medical-record services so miserable to use for doctors is that doctors are *not* allowed to simply not use them. They are required by law to keep an amount of records that is too large to track by hand. This, plus[switching costs](../you-are-in-a-box/#switching-costs-and-growth) , leads to \"oligopolies\" where a small group of incumbents can corner the market regardless of how bad their products are.[↩](#fr-3-1)\n8. \noutside of very limited niches like Apple Shortcuts, Salesforce, and Excel spreadsheets [↩](#fr-5-1)\n\nDiscuss on\n\n[Hacker News](https://hn.algolia.com/?query=jyn.dev/tokens-too-cheap-to-meter/&type=story),\n\n[Lobste.rs](https://lobste.rs/stories/url/latest?url=https://jyn.dev/tokens-too-cheap-to-meter/),\n\n[Mastodon](https://tech.lgbt/@jyn/117315448893865380), or\n\n[Bluesky](https://bsky.app/profile/jyn.dev/post/3mw4ktkxl3s2n)", "url": "https://wpnews.pro/news/tokens-too-cheap-to-meter", "canonical_source": "https://jyn.dev/tokens-too-cheap-to-meter/", "published_at": "2026-09-16 00:00:00+00:00", "updated_at": "2026-10-06 16:19:25.695243+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-infrastructure", "ai-chips", "ai-research"], "entities": ["Epoch.AI", "Artificial Analysis", "Anthropic", "OpenAI", "Claude Fable-5.1", "GPT-5.6 Luna", "Qwen3 Coder", "Muse Glimmer"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/tokens-too-cheap-to-meter", "markdown": "https://wpnews.pro/news/tokens-too-cheap-to-meter.md", "text": "https://wpnews.pro/news/tokens-too-cheap-to-meter.txt", "jsonld": "https://wpnews.pro/news/tokens-too-cheap-to-meter.jsonld"}}