cd /news/ai-infrastructure/building-infrastructure-for-an-ai-wo… · home topics ai-infrastructure article
[ARTICLE · art-132835] src=newsroom.amd.com ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

Building Infrastructure for an AI World in Motion

At the AI Infra Summit, AMD argued that AI infrastructure should be designed around adaptability rather than any single chip or fabric, citing three auditable properties — adaptable, economical and open — as the way to preserve the freedom to move to whatever workload comes next. AMD said its portfolio spans AMD EPYC server CPUs, AMD Radeon and AMD Ryzen AI platforms for local AI, AMD Instinct PCIe cards for large language model inference and fine-tuning, eight-way AMD Instinct GPU systems, and the AMD Helios rackscale solution for mega-scale deployments. The company said the metric to watch for agentic AI is cost per finished task rather than cost per token, since one agentic request becomes a chain of inferences.

by read4 min views1 publishedSep 17, 2026
Building Infrastructure for an AI World in Motion
Image: Newsroom (auto-discovered)

No one in this industry can tell you what AI will look like five years from now.

No one knows the next dominant model architecture. No one knows where the mix between training and inference will settle. And no one knows how much AI will run in your data center, in someone else’s data center or on your laptop (or on something smaller). But infrastructure leaders still have to make decisions and investments today that will be the IT foundations for their businesses for five to seven years.

That is the challenge. The simple truth is that the greatest risk is not that you pick the wrong technology today. It’s that you build infrastructure that can’t adapt to changing requirements as workloads evolve tomorrow.

At the AI Infra Summit this week, I made the case that we need to rethink how we design AI infrastructure. It shouldn’t be around a chip or a fabric, but around one fundamental question: How easily can the system adapt when requirements change?

Adaptability Becomes the Design Principle #

A few years ago, “AI infrastructure” mostly meant large training runs creating predictable, batch-shaped demand on one big cluster. And while that workload still matters, it now has to coexist with inference, agentic AI and an increasingly distributed set of workloads with very different requirements.

Inference is always on and cares about latency and cost per request. Agentic AI compounds that demand because one request can become a series of steps: It retrieves and shapes the inputs, calls tools, executes code, coordinates sub-agents and holds state that outlives any single request.

That changes more than the amount of compute required; it changes the shape of the system. Infrastructure optimized for yesterday’s training clusters isn’t just short on capacity, it’s the wrong shape: The balance of compute, memory and network it was built around no longer matches where the work actually goes. And you can’t buy your way out of a shape problem by adding more of what you already have.

The key is to optimize the system, not only the individual components. Compute, memory, networking, storage, software, power and reliability all have to work together. Scaling something familiar is a budgeting exercise, while absorbing something new is an architectural and operations problem.

Three Properties that Keep You Free #

If the requirements keep changing, then the design goal isn’t to predict the winner – it’s to preserve the freedom to move to what’s next. In practice, that comes down to three auditable properties: Adaptable: You need a portfolio deep enough that when the workload shifts, the answer is a different configuration, not a different vendor. Match the right compute to the right job, with a common software foundation that makes those choices practical.

Economical: This doesn't mean cheap – it means paying for the right tool for the job. Training was a capital decision you made once; inference is an operating decision you make a billion times a day. And an agentic request isn't one inference but a chain of them, so the number to watch is cost per finished task, not cost per token.

Open: Open ecosystems provide choice through broadly adopted standards that allow customers to avoid dependence on a single vendor’s product roadmap. That’s real choice across suppliers, and it’s why open interconnects and open networking matter.

Where the AMD Portfolio Comes In #

This is why breadth of technology is important. It’s why AMD has invested across the entire stack rather than in a single hero product. Today’s AI requirements span:

  • Inference, tokenization, orchestration and other general-purpose compute on AMD EPYC™ Server CPUs extending to small models and local AI on AMD Radeon™ and AMD Ryzen™ AI platforms.
  • Large language model inference and fine-tuning on AMD Instinct™ PCIe cards.
  • Distributed inference and training on eight-way AMD Instinct™ GPU systems, scaling to the AMD Helios™ rackscale solution for mega-scale deployments.

Individual products aren’t really the point. Breadth without a common software story is just a catalog. Breadth with that software is flexibility. Unified by AMD ROCm™ software, a workload’s hardware can move without starting over.

We build across the stack because we think each piece has to win on its own merits not because the stack only works if you buy all of it. Open standards are what make that true: Choose the AMD pieces that are best for your workload and keep the freedom to choose differently elsewhere.

Organizations that do well over the next five years won’t be the ones that predicted every turn. They’ll be the ones that have the freedom and flexibility to make the turn when they arrive at it.

Build adaptable, so the future – whatever shape it takes – still fits. Build economically, so you're paying for the right tool, not the biggest one. Build open, so you preserve choice – because when you don’t, you pay for it.

What we can guarantee is that AI will keep changing. We should build the infrastructure with that assumption from the start.

[Artificial Intelligence](https://newsroom.amd.com/category/ai/)

[Data Center](https://newsroom.amd.com/category/data-center/)

[Thought Leadership Blogs](https://newsroom.amd.com/category/thought-leadership-blogs/)

[#epyc](https://newsroom.amd.com/tag/epyc/)

[#radeon-pro](https://newsroom.amd.com/tag/radeon-pro/)

[#ryzen](https://newsroom.amd.com/tag/ryzen/)

[#instinct](https://newsroom.amd.com/tag/instinct/)

[#rocm](https://newsroom.amd.com/tag/rocm/)

[#helios](https://newsroom.amd.com/tag/helios/)

 Press inquiries:  [corporate.pressinquiry@amd.com](mailto:corporate.pressinquiry@amd.com)
── more in #ai-infrastructure 4 stories · sorted by recency
── more on @amd 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/building-infrastruct…] indexed:0 read:4min 2026-09-17 ·