Before GPT-6 Astra made headlines for finding zero-days, it did something quieter and, to me, more economically startling: it produced proofs for ten mathematics and theoretical computer science problems that had sat unsolved for decades, and it did it for a reported few thousand dollars of compute. Everyone focused on the math. I keep thinking about the invoice. Because the story that matters for anyone who budgets infrastructure is not "AI is smart now," it is what happens to planning when a unit of genuinely novel intellectual work drops to the price of a rounding error.
Decades-unsolved problems are, by definition, things that resisted a lot of expert human effort. The traditional cost of solving one is measured in careers, not dollars. Astra reportedly cleared ten for an amount of compute that would not survive a single line-item review on most cloud bills.
Set aside whether every proof holds up (that verification matters and is its own story). The direction is the point: the marginal cost of attempting hard, novel, high-value intellectual work is collapsing toward the cost of the compute to run the attempt. That is a different economic regime, and it breaks assumptions that FinOps and capacity planning are built on.
For years, expensive cognitive work was a fixed, scarce, human input you planned around. You could not "scale up" a research breakthrough by renting more of it. Now, increasingly, you can attempt to, by spending compute. That changes three things about how you budget: None of this is exotic if you already do cloud cost work. It is the same muscles, pointed at a new kind of spend:
Astra solving ten decades-old problems cheaply is being read as a capability milestone, and it is one. But the more durable lesson for anyone who plans infrastructure spend is economic: the cost of attempting hard, novel work is collapsing toward compute, which means categories of work are migrating from headcount onto your cloud bill, where they behave like every other usage-metered cost, cheap per unit, dangerous in aggregate, and worthless unless you connect the spend to verified value. The teams that treat this like real FinOps, attribute, cap, verify, alert, will get the leverage. The teams that treat it as "it's only a few thousand dollars" will find out how fast a few thousand dollars, many times over, adds up.
Is your organization starting to spend real compute on open-ended, novel work yet? And if so, is that spend tracked as a real cost line with attribution, or is it still "it's just some API calls"? That gap is where I expect the next round of budget surprises.