{"slug": "axelera-ai-data-center-inference-performance-in-the-power-envelope-of-embedded", "title": "Axelera AI: Data Center Inference Performance in the Power Envelope of Embedded Systems", "summary": "Axelera AI launched Europa, a purpose-built inference accelerator delivering 629 TOPS at INT8 across eight second-generation AI cores at 35 watts typical during LLM inference, announced at the AI Infra Summit 2026 in Santa Clara, California. Co-founder and CEO Fabrizio Del Maffeo said the chip ships as the Axelera Edge 232p half-height, half-length PCIe card and the Axelera Server 250p full-height, full-length card with four AI processing units, paired with the Voyager SDK that takes PyTorch models with no intermediate conversion and the new agentic developer assistant Voyager Wingman. Del Maffeo framed power, economics and fragmentation as the three barriers to scaling physical AI, arguing inference must run inside power envelopes measured in tens of watts where data is created.", "body_md": "Physical AI is hitting three walls: power, economics, and fragmentation. Fabrizio Del Maffeo explains how to break them.\n\nRecorded at the AI Infra Summit 2026 in Santa Clara, California, Axelera AI Co-Founder and CEO Fabrizio Del Maffeo sits down with Sally on why intelligence was never the bottleneck to AI adoption. Infrastructure is.\n\n### The three walls to scaling physical AI\n\nDel Maffeo’s keynote framed three barriers standing between today’s pilots and real deployment.\n\nPower and physics. The data center was built for training: persistent, high-wattage, liquid-cooled, and tolerant of moving data long distances. Physical AI is the inverse. Inference happens where data is created—in factories, vehicles, hospitals and city infrastructure—inside power envelopes measured in tens of watts. Shrinking a training architecture does not fix that, because the dominant cost is not computation, it is data movement. Axelera’s Digital In-Memory Compute performs matrix-vector multiplication, roughly 80% of the inference workload, inside the memory cell itself. Less movement, less energy, more work per watt.\n\n[View All](https://www.eetimes.com/category/sponsored-content/)\n\n[EE Times](https://www.eetimes.com/author/ee-times/)10.05.2026\n\nEconomics and control. Cloud inference made sense for experimentation. For production workloads running continuously, the bill never stops growing, and in regulated industries, the data cannot leave the building in the first place. Del Maffeo argues the decision has shifted from a cost question to a control question: your models, your data, your cost structure.\n\nFragmentation. Every new accelerator brings its own toolchain, conversion step and lock-in. Developers pay that tax, and it is why projects stall between demo and deployment.\n\n### Start at the edge, then scale up\n\nMost semiconductor companies design for the cloud and strip the architecture down for smaller devices. Axelera did the opposite: solve the hardest, most power-constrained environments first, then carry that efficiency up. Europa, launched at the show, is what that looks like at enterprise scale.\n\nEuropa is a purpose-built inference accelerator, delivering 629 TOPS at INT8 with eight second-generation AI cores, 16 on-chip RISC-V vector processors for pre- and post-processing, LPDDR5 at 200 GB/s, hardware HEVC/H.265 decode, and a Kudelski KSE3 secure enclave, at 35 watts typical during LLM inference. It ships as the Axelera Edge 232p, a half-height, half-length PCIe card, and the Axelera Server 250p, a full-height, full-length card with four AI processing units. Standard slot, existing server, no new rack, no new cooling.\n\n### Software is the missing link\n\nSilicon is the easy half. Del Maffeo makes the case that hardware and software have to be co-designed, and that manual optimization is where developer velocity dies. The Voyager SDK takes PyTorch models directly with no intermediate conversion, covering CNNs, transformer vision models, VLMs and LLMs, with quantization and graph optimization handled by the compiler. Voyager Wingman, the new agentic developer assistant, turns plain language into a working pipeline, runs it on real hardware and iterates until it performs. Built on RISC-V, open by design, no proprietary lock-in.\n\n### Ecosystem validation\n\nA chip alone deploys nothing. Del Maffeo closes on the partner ecosystem, from defense and agricultural robotics to retail analytics and rail, and how system integrators, model builders and hardware vendors working against one toolchain are collapsing the traditional 12- to 18-month deployment cycle.\n\n##### Resources:\n\n[https://axelera.ai/ai-accelerators/aipu/europa](https://axelera.ai/ai-accelerators/aipu/europa)\n\n[https://axelera.ai/ai-software/voyager-sdk#wingman](https://axelera.ai/ai-software/voyager-sdk#wingman)", "url": "https://wpnews.pro/news/axelera-ai-data-center-inference-performance-in-the-power-envelope-of-embedded", "canonical_source": "https://www.eetimes.com/axelera-ai-data-center-inference-performance-in-the-power-envelope-of-embedded-systems/", "published_at": "2026-10-06 18:03:14+00:00", "updated_at": "2026-10-06 18:19:33.131720+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-chips", "ai-tools", "developer-tools", "artificial-intelligence"], "entities": ["Axelera AI", "Fabrizio Del Maffeo", "Europa", "Axelera Edge 232p", "Axelera Server 250p", "Voyager SDK", "Voyager Wingman", "AI Infra Summit 2026"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/axelera-ai-data-center-inference-performance-in-the-power-envelope-of-embedded", "markdown": "https://wpnews.pro/news/axelera-ai-data-center-inference-performance-in-the-power-envelope-of-embedded.md", "text": "https://wpnews.pro/news/axelera-ai-data-center-inference-performance-in-the-power-envelope-of-embedded.txt", "jsonld": "https://wpnews.pro/news/axelera-ai-data-center-inference-performance-in-the-power-envelope-of-embedded.jsonld"}}