# Axelera AI: Data Center Inference Performance in the Power Envelope of Embedded Systems

> Source: <https://www.eetimes.com/axelera-ai-data-center-inference-performance-in-the-power-envelope-of-embedded-systems/>
> Published: 2026-10-06 18:03:14+00:00

Physical AI is hitting three walls: power, economics, and fragmentation. Fabrizio Del Maffeo explains how to break them.

Recorded at the AI Infra Summit 2026 in Santa Clara, California, Axelera AI Co-Founder and CEO Fabrizio Del Maffeo sits down with Sally on why intelligence was never the bottleneck to AI adoption. Infrastructure is.

### The three walls to scaling physical AI

Del Maffeo’s keynote framed three barriers standing between today’s pilots and real deployment.

Power and physics. The data center was built for training: persistent, high-wattage, liquid-cooled, and tolerant of moving data long distances. Physical AI is the inverse. Inference happens where data is created—in factories, vehicles, hospitals and city infrastructure—inside power envelopes measured in tens of watts. Shrinking a training architecture does not fix that, because the dominant cost is not computation, it is data movement. Axelera’s Digital In-Memory Compute performs matrix-vector multiplication, roughly 80% of the inference workload, inside the memory cell itself. Less movement, less energy, more work per watt.

[View All](https://www.eetimes.com/category/sponsored-content/)

[EE Times](https://www.eetimes.com/author/ee-times/)10.05.2026

Economics and control. Cloud inference made sense for experimentation. For production workloads running continuously, the bill never stops growing, and in regulated industries, the data cannot leave the building in the first place. Del Maffeo argues the decision has shifted from a cost question to a control question: your models, your data, your cost structure.

Fragmentation. Every new accelerator brings its own toolchain, conversion step and lock-in. Developers pay that tax, and it is why projects stall between demo and deployment.

### Start at the edge, then scale up

Most semiconductor companies design for the cloud and strip the architecture down for smaller devices. Axelera did the opposite: solve the hardest, most power-constrained environments first, then carry that efficiency up. Europa, launched at the show, is what that looks like at enterprise scale.

Europa is a purpose-built inference accelerator, delivering 629 TOPS at INT8 with eight second-generation AI cores, 16 on-chip RISC-V vector processors for pre- and post-processing, LPDDR5 at 200 GB/s, hardware HEVC/H.265 decode, and a Kudelski KSE3 secure enclave, at 35 watts typical during LLM inference. It ships as the Axelera Edge 232p, a half-height, half-length PCIe card, and the Axelera Server 250p, a full-height, full-length card with four AI processing units. Standard slot, existing server, no new rack, no new cooling.

### Software is the missing link

Silicon is the easy half. Del Maffeo makes the case that hardware and software have to be co-designed, and that manual optimization is where developer velocity dies. The Voyager SDK takes PyTorch models directly with no intermediate conversion, covering CNNs, transformer vision models, VLMs and LLMs, with quantization and graph optimization handled by the compiler. Voyager Wingman, the new agentic developer assistant, turns plain language into a working pipeline, runs it on real hardware and iterates until it performs. Built on RISC-V, open by design, no proprietary lock-in.

### Ecosystem validation

A chip alone deploys nothing. Del Maffeo closes on the partner ecosystem, from defense and agricultural robotics to retail analytics and rail, and how system integrators, model builders and hardware vendors working against one toolchain are collapsing the traditional 12- to 18-month deployment cycle.

##### Resources:

[https://axelera.ai/ai-accelerators/aipu/europa](https://axelera.ai/ai-accelerators/aipu/europa)

[https://axelera.ai/ai-software/voyager-sdk#wingman](https://axelera.ai/ai-software/voyager-sdk#wingman)
