cd /news/artificial-intelligence/how-the-arm-agi-cpu-supports-key-ai-… · home topics artificial-intelligence article
[ARTICLE · art-69090] src=newsroom.arm.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

How the Arm AGI CPU supports key AI workloads at scale

Arm Holdings plc introduced the Arm AGI CPU, designed to balance compute, memory, I/O, and power for AI agent workloads at scale. The chip sustains high utilization in video processing and streaming tasks, outperforming legacy x86 architectures by maintaining stable performance at 90% allocation versus x86's 50% limit. Arm claims the AGI CPU enables greater performance per rack, improved power efficiency, and lower total cost of ownership as agentic AI deployments scale.

read6 min views2 publishedJul 22, 2026
How the Arm AGI CPU supports key AI workloads at scale
Image: Newsroom (auto-discovered)

As AI agents drive a major shift in computing, infrastructure must coordinate thousands of concurrent tasks while continuously moving data between CPUs, memory, storage and AI accelerators. This means AI workloads are becoming increasingly distributed and dependent on sustained performance.

Unlike traditional AI inference, agentic AI rarely stresses compute alone. AI agents continuously retrieve and validate information, access external tools, reason through problems and execute multi-step workflows. This creates an entirely different workload profile with different system bottlenecks, which changes the role of the CPU. Rather than simply supporting accelerators, the CPU becomes the orchestration engine for AI infrastructure. CPUs orchestrate workflows, exchange data with accelerators, manage inference pipelines and execute many parallel processes simultaneously.

As agentic workloads scale, legacy x86 architectures can become constrained by memory and I/O resources, reducing utilization, limiting throughput and requiring additional servers to maintain performance. The Arm AGI CPU is designed to keep compute, memory, I/O and power in balance. With significantly higher memory and I/O bandwidth per core, lower thermal design power (TDP) and greater rack density, it helps sustain high utilization across CPU driven tasks and AI host-node workloads while keeping accelerators productive. This balanced architecture enables greater performance per rack, improved power efficiency and more deterministic performance as agentic AI deployments scale across the most demanding use cases.

Scaling up performance per rack for video processing #

Video processing is a strong example of how AI agents are changing data center workloads. As video is processed, compressed and converted for end users, each frame is analyzed for object detection, metadata tagging, translation and summarization. These steps place sustained pressure on infrastructure. Video requires predictable performance at 30 fps, and the CPU plays a central role in orchestrating inference, memory and data flows across the pipelines.

In the demo scenario above, an x86 rack remains stable up to only 50 percent allocation before performance slows and frame rates drop. As utilization increases, x86 performance break, servers become underutilized and additional racks are required to sustain performance. With the Arm AGI CPU, the same workload stays stable at 90 percent allocation, with consistent frame rates, higher core utilisation and more performance per rack. This helps data centers improve efficiency and total cost of ownership as AI agent workloads scale.

Higher quality, better utilization, smarter video streaming #

Alongside video processing, streaming is one of the largest and most demanding workloads on the internet. Whether people are watching a live event, joining a video call or viewing a product review, they expect the video to load quickly and play without buffering, lag or quality drops.

The scale behind these experiences are significant. Video currently accounts for around 82 percent of all internet traffic, which means data centers need to process, compress and deliver huge numbers of streams while maintaining consistent playback quality. Then, as agentic AI expands, video streams are becoming an important source of real-time data, enabling AI agents to analyze scenes, detect events and automate decisions while delivering high-quality video experiences to users.

The Arm AGI CPU helps data centers deliver smooth, high-quality video experiences while using less power and supporting more streams per rack. In the demo above, a high-end x86 server can support video streams up to 256 in 4K at 30 frames per second, but performance begins to break down when the number of streams exceeds 256.

An Arm AGI CPU server delivers 256 simultaneous 4K streams at 30 frames per second with no dropped frames or interruptions, while delivering 55 percent more streams per rack within the same power envelope. This helps platforms deliver smarter video services at scale, with fewer interruptions, lower power use and more streams per rack.

Powering multi-agent AI workloads for finance at scale #

Each day companies process large volumes of financial documents, operational events and customer interactions. At enterprise scale, that can mean thousands of workloads running across finance systems, databases, and approval processes. This creates a need for orchestrated decision-making across many AI agents. In finance workflows, agents may scan SAP databases and turn financial data into invoice documents. At the same time, other agents can validate data integrity, compare historical trends, assess governance risk and detect inconsistencies.

The Arm AGI CPU powers sustained multi-agent AI workloads at scale, helping these agents run in parallel while keeping data moving through the workflow.** **Invoices can be categorized by client, work order, invoice risk distribution, work order risk distribution and validation ratio. For SAP users, this can help reduce the time spent on manual financial checks and support faster decisions, greater operational efficiency and more consistent real-world outcomes.

Accelerating EDA workloads #

Advanced chip design depends on electronic design automation (EDA). These workloads push compute infrastructure hard, using massive datasets, thousands of simultaneous processes and run times that can last days or weeks. Traditionally, engineering teams have relied heavily on cloud infrastructure to handle these workloads. But balancing performance, cost and workload placement across cloud and on-premises environments remains a challenge.

Siemens Questa Visualizer and Cadence Innovus run smoothly on Arm AGI CPU, supporting massive EDA workloads across cloud and on-premises infrastructure. Working with design datasets that can exceed hundreds of terabytes, the Arm AGI CPU enables engineering teams to maintain consistent performance while scaling workloads across cloud and on-premises environments. This gives engineering teams the flexibility to run jobs locally or burst into the cloud for larger simulations, while maintaining consistent performance and optimizing infrastructure utilization and costs.

Running customer agents at scale with 2x more requests per rack #

Customer service is becoming a practical entry point for agentic AI in the enterprise. These systems interpret customer intent, route tasks to specialized agents and execute multi-step workflows across knowledge bases, Customer Relationship Management (CRM) platforms and back-office systems.

The Arm AGI CPU helps enterprises run autonomous customer service applications at scale by coordinating AI agents, managing workflow state and keeping data moving efficiently across enterprise systems. Working alongside AI accelerators, it enables responsive, multi-step customer interactions while maintaining high system utilization.

In a rack power budget under 40kW, an Arm AGI CPU platform paired with Rebellions’ RebelCard accelerators can support approximately 2x more automated customer requests per rack than a comparable legacy x86-based deployment, while maintaining similar token throughput per accelerator and using only about the half of the system power.*

Built for the next phase of AI infrastructure #

Agentic AI is redefining what matters in CPU architecture. Modern AI infrastructure depends on high memory bandwidth per core to keep AI agents supplied with data, high I/O bandwidth to keep AI accelerators productive, and sustained compute utilization to efficiently orchestrate thousands of concurrent workflows. At the same time, all of this must happen within the power and thermal limits of a modern data center. The Arm AGI CPU is built to meet these demands, delivering greater performance per rack, higher efficiency and scalable performance across key use cases that are redefining modern data center workloads.

*Based on Arm estimates

Any re-use permitted for informational and non-commercial or personal use only.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arm holdings plc 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-the-arm-agi-cpu-…] indexed:0 read:6min 2026-07-22 ·