Netherlands-based AI accelerator start-up Axelera is targeting a major challenge facing enterprise AI deployments with its new Europa platform. Europa is designed to deliver more AI inference performance than competing solutions within the constrained power and cooling budgets of enterprise data centers.
The power demands of AI infrastructure have become one of the primary limiting factors of current data center buildouts. Hyperscalers can secure dedicated generation capacity, negotiate directly with utilities and, in some cases, pursue entirely new power sources to support massive AI clusters. Many corporate data centers, however, don’t have these luxuries.
As businesses move AI applications from experimentation into production, performance-per-watt becomes critical. AI compute often has to fit within the infrastructure and power envelope enterprises already have, making useful inference capacity per watt a critical deployment metric.
Axelera’s new Europa AI Processing Unit (AIPU) takes the company’s power-efficient approach to edge and physical AI and scales it to enterprise server-class workloads. Europa will initially appear in two PCIe add-in card-based products: the half-height, half-length Edge 232p and full-height, full-length Server 250p, allowing organizations to add AI inference capacity to existing server environments with a standard PCIe card upgrade.
Axelera
Axelera’s product roadmap initially targeted intelligent edge applications, where power is a hard design constraint and inference requires consistent performance under tight latency, thermal and power limits. Its premise with Europa is that an architecture designed around those constraints can scale effectively into servers as more power, cooling and physical space becomes available.
The Europa AIPU delivers a claimed 629 TOPS within a 45-watt power envelope, with eight AIPU cores and 200 GB/s of memory bandwidth. The architecture also integrates RISC-V vector cores for pre- and post-processing along with an H.265 decoder, reducing data movement between the accelerator and host processor for machine vision-oriented workloads like surveillance systems.
Those characteristics become particularly relevant as enterprises deploy AI across hundreds of servers, where accelerator power consumption affects server configuration, rack density, cooling and ultimately how much inference capacity can fit within a facility’s electrical envelope.
Axelera claims its Edge 232p can deliver up to six times more tokens per second per watt than a competing GPU-based solution across several Llama and Qwen models. Those numbers combine Axelera’s internal testing with publicly available competitor benchmarks, so independent validation across a wider range of workloads will be important.
However, I’ve previously seen evidence that the company’s architectural approach can deliver strong efficiency. In a HotTech study of AI accelerators for machine vision, Axelera’s first-generation Metis accelerator delivered the highest performance in multi-stream inference testing, while its PCIe and M.2 implementations consistently delivered the best energy efficiency.
Europa takes that same power-conscious design philosophy to substantially larger enterprise workloads.
The Server 250p is perhaps the more interesting product from an enterprise data center perspective. It’s a full-height, full-length dual-slot PCIe accelerator available with either 128GB or 256GB of LPDDR5 memory.
Axelera says a single Server 250p can deliver 4,205 tokens per second running Qwen3 8B at INT4 when serving a large batch of simultaneous requests, with an efficiency of 32.1 tokens per second per watt. Up to eight cards can be installed in a server, providing a path to scale inference capacity within conventional server infrastructure.
That deployment model could be useful for enterprises serving internal AI applications to many simultaneous users, allowing organizations to deploy inference locally and add accelerators as utilization grows, rather than sending every request to the public cloud or building dedicated high-power GPU infrastructure.
Security and governance are considerations as well. Financial services, healthcare, legal, defense and government organizations face restrictions on where sensitive information is processed and stored. Europa provides an option for running those workloads on-prem within local infrastructure that an organization controls itself.
Of course, new server solutions also need to be available in full system solutions that enterprise IT can buy, deploy and support. Axelera’s Edge 232p card is available in fully validated Dell and Supermicro systems, while Axelera’s broader roster of validated OEM partners also includes HPE, Lenovo, Advantech and Axiomtek.
In terms of software enablement, Axelera’s Voyager SDK spans its existing Metis products and the new Europa architecture, providing a common environment across embedded, edge and server deployments, with support for a multitude of computer vision models, LLMs, VLMs, diffusion models, speech and other AI workloads.
To automate setup, Axelera’s Voyager Wingman uses natural-language prompts to help developers create or port inference pipelines, while AxeleraScript, or AxScript, provides a Python-enabled domain-specific language with lower-level AIPU control for custom operators and transformer models.
This could prove every bit as important as Europa’s performance and efficiency. Enterprises already have models, development environments and application stacks. Extensive rewriting or specialized expertise adds development and operational costs that can quickly undermine savings on hardware and power.
Axelera isn’t positioning Europa as hardware for training the next frontier model. Its opportunity lies in production enterprise inference, delivering performance while staying within existing power, cooling, security and budget constraints. The company also appears to be gaining commercial traction, reporting deployments with more than 600 customers globally and a sales pipeline exceeding $1.5 billion.
For CIOs and infrastructure engineers, Europa may represent another option between consuming AI entirely from the public cloud and deploying power-hungry dedicated GPU infrastructure on premises. If Axelera’s performance-per-watt claims hold up across a broader range of enterprise workloads, the economics could be compelling. Hyperscalers have the capital and scale to chase new power sources and rack architectures for AI. Conversely, most enterprises have to work within the power limits, cooling capacity and data center footprint they already own. This makes efficiency a fundamental infrastructure constraint and opens an opportunity for power-efficient architectures like Axelera’s Europa in the burgeoning enterprise AI market.