{"slug": "ai-is-getting-cheaper-but-ai-bills-could-still-go-up", "title": "AI Is Getting Cheaper, But AI Bills Could Still Go Up", "summary": "The cost of reaching a fixed level of AI benchmark performance has fallen roughly 47% per quarter since 2023, according to Epoch AI, with the price of hitting a 75% score on the GPQA Diamond benchmark dropping from about $0.30 per question for OpenAI's o3 in January 2025 to roughly $0.0004 per question for GPT-5.6 Luna. Epoch AI's tracking shows the decline varies by task type, falling about 50% to 52% per quarter for mathematics and pure logic, 39% to 43% for combinatorial games such as chess and puzzles, and about 27.5% per quarter on SWE-bench software engineering tasks. The publication notes that cheaper per-call inference does not necessarily lower overall AI bills, since systems increasingly chain multiple model calls to generate, check and refine answers.", "body_md": "Thirty cents.\n\nThat was the estimated cost of getting an AI model to a particular level of performance on a difficult benchmark in early 2025.\n\nLess than 18 months later, reaching that same performance level could cost around $0.0004.\n\nThat’s a huge drop. And it happened in a field where better models have generally meant spending more money to run them.\n\nThe number comes from [Epoch AI](https://epoch.ai/publications/the-plunging-price-of-thought), which has been tracking how much it costs to reach the same level of performance across different AI benchmarks.\n\nThe cost of reaching the same level of AI performance is falling faster than it has for most technologies we’ve seen before.\n\nSo what exactly is getting cheaper? And if AI is getting this cheap, why can the bills still go up?\n\n## **Table of Contents**\n\n## **The 47% Collapse in Token Costs**\n\nFor most of computing history, adding more intelligence meant adding more people.\n\nMore code reviews required more engineers. More scientific research required more researchers. More analysis required more analysts.\n\nAI changes that equation because the cost of running a capable model keeps falling.\n\nEpoch AI tracks this by looking at how much it costs to reach a specific level of performance on AI benchmarks. Instead of comparing one model’s price over time, it measures the cheapest available way to achieve the same result.\n\nSince 2023, that cost has been falling by roughly 47% every quarter.\n\nThe drop becomes easier to understand with one example.\n\nWhen OpenAI released o3 in January 2025, Epoch AI estimated that reaching a 75% score on the GPQA Diamond benchmark cost around $0.30 per question. The benchmark tests advanced knowledge across fields like chemistry, physics and biology.\n\nLater, GPT-5.6 Luna reached the same performance level at around $0.0004 per question.\n\nThe cost of the same level of AI capability had fallen hundreds of times in a relatively short period.\n\nBut the important part is what happens after intelligence becomes cheap.\n\nBecause cheaper AI does not automatically mean every AI task becomes cheap.\n\n## **Cheap Intelligence Meets Expensive Problems**\n\nA model can solve a difficult math problem in a test environment, but a software engineer using AI has to deal with an existing codebase, unclear requirements, changing priorities, and decisions that are difficult to measure.\n\nThe gap becomes clearer when looking across different types of tasks.\n\nMathematics and pure logic have seen some of the fastest cost reductions, falling around 50% to 52% per quarter. Structured problems are easier to optimize because the path to a correct answer is clearer.\n\nHard science tasks have followed a similar downward trend.\n\nBut areas like combinatorial games and software engineering have moved more slowly. Chess and puzzle-style tasks fell around 39% to 43% per quarter, while SWE-bench data showed a slower decline of around 27.5% per quarter.\n\nSoftware engineering is a good example of the difference between passing a test and doing useful work. Fixing a real bug often requires understanding why a system was built a certain way, not just producing a correct piece of code.\n\n## **How Cheaper AI Changes Software Architecture**\n\n**How Cheaper AI Changes Software Architecture**\n\nA coding agent can now write a change, run tests, inspect failures and try again without every step being treated as a major expense. A separate model can review the output. Another can compare approaches before the final result is returned.\n\nThe software is no longer built around a single AI response. It is built around a process that can generate, check and refine answers.\n\nThe focus shifts from How do we reduce the number of AI calls? to How do we design systems that use those calls effectively?\n\n##### **Also Read:** [OpenAI’s Agents Leaked 53 Images. What Else Did Its Agents Do?](https://firethering.com/openai-agents-53-images/) \n\n## **The inference paradox: cheap calls, bigger bills**\n\nLower AI prices create a strange situation, using AI more can still make companies spend more.\n\nA simple chatbot interaction might use a few hundred tokens. An agentic workflow can use many times more because it has to read context, call tools, check results and repeat steps until it reaches an acceptable answer.\n\nA coding agent working through a large repository is a good example. It may inspect files, make changes, run tests, review errors and try again multiple times. Each individual AI call may be cheaper, but the total number of calls grows quickly.\n\nThis is why cheaper models do not always translate into smaller AI budgets.\n\nGartner describes this as the Inference Paradox. As AI systems become capable of handling more complex tasks, companies often build workflows that require more inference, not less.\n\nGartner also expects the total inference cost of agentic workflows to increase more than fivefold through 2028 as these systems become more common.\n\nThe cost of each AI action is falling but the amount of AI work being assigned is growing even faster.\n\n## **The New Challenge Is Orchestration**\n\nThe cost of AI intelligence is falling quickly, but using that intelligence effectively is becoming a different challenge.\n\nAs software starts relying on more agents, more verification steps, and longer reasoning loops, the total cost of a workflow can still increase even when each individual AI call becomes cheaper.\n\nThat makes the design of the system just as important as the model behind it.\n\nA well-built AI system knows when to use a cheaper model, when a more capable one is worth the cost, how much context to provide, and where additional reasoning actually improves the result.\n\nThe advantage will come from building systems that can get more value from every AI call without wasting the savings that cheaper intelligence provides.", "url": "https://wpnews.pro/news/ai-is-getting-cheaper-but-ai-bills-could-still-go-up", "canonical_source": "https://firethering.com/ai-costs-are-falling/", "published_at": "2026-10-01 21:03:52+00:00", "updated_at": "2026-10-01 21:49:13.383942+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-agents", "ai-infrastructure"], "entities": ["Epoch AI", "OpenAI", "o3", "GPT-5.6 Luna", "GPQA Diamond", "SWE-bench"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-is-getting-cheaper-but-ai-bills-could-still-go-up", "markdown": "https://wpnews.pro/news/ai-is-getting-cheaper-but-ai-bills-could-still-go-up.md", "text": "https://wpnews.pro/news/ai-is-getting-cheaper-but-ai-bills-could-still-go-up.txt", "jsonld": "https://wpnews.pro/news/ai-is-getting-cheaper-but-ai-bills-could-still-go-up.jsonld"}}