AI Infra Summit in Santa Clara is bigger than ever. 8,000 attendees (up from 3,500 last year) crowd the courtyards, parking lots, hallways, and expo halls — you don’t have to look far for a conversation.
The organizers framed the show with a single line: “the age of inference is here, and it’s time to connect infrastructure investment to enterprise ROI”.
After a busy first day at the show, these are our top three takeaways for you, whether you’re running a platform, standing up an AI factory, or trying to get more value out of the GPUs you already bought.
Tokens per watt is the new benchmark #
At last year’s event, the conversation was still mostly about FLOPS. This year, every major session and vendor talked about tokens per watt, in some form.
Dave Patterson opened the mainstage with a talk titled “The Bitter Lesson, Revisited: How Far Does More Compute Get Us?” NVIDIA’s Ian Buck followed with a keynote on tokens per watt for AI factories. NVIDIA’s headline claim was Vera Rubin combined with Groq 3 LPX delivering up to 35x higher token throughput per megawatt than GB200 NVL72, on 2-trillion-plus-parameter models at long context.
When power is the constraint, the most important thing is to maximize the inference you can extract from your available capacity. That’s why tokens per watt is becoming a C-level metric, and every architecture choice (silicon, cooling, serving stack, orchestration) either helps that ratio or hurts it.
Agentic workloads are driving a token bill hockeystick #
One figure that stood out came from NVIDIA: agentic AI sessions can consume roughly 15 times as many tokens as a simple chat request. Buck’s team used SemiAnalysis’s AgentX benchmark, which measures inference on recorded real-world agentic coding sessions rather than single-request tests, with context growth, tool calls, and sub-agent spawning all preserved.
Apply that 15x ratio to your own engineering teams. As agentic coding becomes mainstream, your inference bill will skyrocket.
Our co-founder and CTO Saad Malik will take the Data & Models stage at noon today to discuss how to manage that bill with a platform approach. He’ll share the session with AMD’s CVP of Enterprise AI, Kumaran Siva, covering intelligent routing between local and frontier models, metering, quotas, policy control, and a real-time view of token economics.
To keep spending under control as you deploy more agents, you’ll need to treat inference as an infrastructure discipline, building the routing, governance and metering to scale sustainably. Several vendors on the show floor, ourselves included, were making that case.
The enterprise inference control plane is becoming a category #
Across the agenda, sessions on multi-model routing, distributed AI operations, and regulated industries point to a shared concern: how to operate inference across different models and infrastructure. Several introduced terms like ‘the inference control plane’ or ‘service plane’, arguing that GPUs and models are commodities, and the value enterprises are looking for lives in the layer between the two.
The control plane is where you enforce policy and isolate tenants. It’s also where you apply jurisdictional requirements and decide how teams share GPU capacity, with usage data to show what each team spends. Those controls matter for sovereign AI, too. Multiple sessions used the same working definition: infrastructure choice, local control, software portability, and the ability to evolve deployments as national, industry, and workload requirements change. If your platform layer cannot do that, you do not have a sovereign AI capability. You have a rented one.
Our positioning has always been that the platform manages the inference environment, not the models. The first day at AI Infra Summit was a room full of practitioners and vendors arriving at the same conclusion from different starting points.
Plan for inference efficiency and agent costs #
If you’re planning your AI infrastructure roadmap for the next year, I’d start here. Measure tokens per watt alongside cost per successful task and utilization. Test with your own workloads so you can judge what the extra capacity is worth.
Treat agentic workloads as a distinct scaling problem. Measure that 15x token consumption across complete agent sessions, including tool calls, retries, and growing context. Design routing and governance from the start.
Invest in the platform layer between GPUs and models. If your current stack lets you choose silicon, choose deployment location, meter usage per tenant, enforce policy per request, and swap models without rewriting your applications, you’re in a good position. If it doesn’t, those gaps belong on your roadmap.
Talk inference with us at booth 646 #
We’re at booth 646 through Thursday. Come find us to talk through what an enterprise-grade inference control plane could look like in your environment.
If you can’t make it to Santa Clara, take a look at what we launched with AMD and Supermicro at Ai4, and the PaletteAI Inference Launchpad product page. More to come from Day 2.