Delos Data Targets Heterogeneous AI With Data Interface Delos Data announced at the AI Infra Summit in Santa Clara, California, that it is developing Apollo, a data interface chiplet offered as a 30+ Tbps I/O chiplet next to an XPU, a 10+ Tbps near-packaged optical interface, or a 400+ Gbps card for CPUs and memory endpoints, and that it has raised more than $100 million in venture capital. Delos Data CTO Dan Daly said the chiplet bridges GPUs, accelerators, CPUs, memory and storage into a single flat, low-latency domain by handling load balancing, topology and failure handling at the endpoint rather than replacing existing switches. The company previously announced its Mosaic data center orchestration software and Asterion server architecture for scale-up domains of up to 1,000 GPUs. SANTA CLARA, Calif. – Delos Data is adding silicon to its portfolio, the startup announced at the AI Infra Summit. Delos previously announced data center orchestration software, Mosaic, and a new server architecture, Asterion https://www.eetimes.com/startup-boosts-scale-up-to-1000-gpus-in-a-single-domain/ , for huge scale-up domains, but is now working on Apollo, a data interface chiplet. Apollo comes either as a 30+ Tbps I/O chiplet that sits next to an XPU, as a 10+ Tbps near-packaged optical interface, or on a 400+ Gbps card for CPUs and memory endpoints. It provides endpoints with guaranteed bandwidth and latency, handling load balancing, topology, and failure handling so the endpoint doesn’t have to do it. Delos is targeting order-of-magnitude improvements in speed, resiliency, and scale versus what endpoints can do today. The company also announced that it has raised more than $100 million in venture capital. New AI infrastructure will cater to either regular inference, agentic AI, or both, Delos Data CTO Dan Daly told EE Times. View All https://www.eetimes.com/category/sponsored-content/ “The theme we see in these emerging new types of infrastructure buildouts is a mixture of hardware,” Daly said. “The CPU folks love to talk about how the CPU is back, but scale up and scale out only connects GPUs. There is another domain needed that includes GPUs, accelerators, CPUs, memory and storage. We’re offering a data interface that can bridge all these different types of devices, each with their own semantics, and put them in a single, flat, low-latency domain.” For example, in prefill-decode disaggregated architectures, GPUs and dataflow accelerators have completely different approaches to memory, Daly said. GPUs place data in HBM and all compute units can access the data via the HBM’s memory map; it’s about data objects and addresses. In a dataflow architecture, a stream of data goes through a sequence of operations—the location determines what happens to the data. The semantics of these architectures are very different and need to be bridged. Communication semantics can also be difficult, since a CPU and a memory controller for example may not speak the same language; today this is handled by a switch, but it incurs latency. It would be more efficient if interconnects could be aware of the semantics of both endpoints, Daly argues, and that is what Delos’ Apollo chiplet is designed to achieve. “There’s an opportunity here to be able to greatly simplify the switching to really maximize the value of those switches—radix, latency, low power, flexibility of media—all the things you want from a switch and less of the things you don’t…legacy things from cloud or scale-out,” Daly said. “There’s an opportunity to open up the aperture here, to enable flexibility in terms of what switches go there, and to potentially simplify the network that goes in between these devices, where it could even be optical.” Inference platforms may use copper and/or optical links, different types of switches, and different topologies, which require detailed co-design and customization. Delos does not want to replace switches; rather, it wants to add Apollo at each endpoint to reuse existing switches of any kind in a more efficient way. Resiliency across multiple hops that could mean bridging multiple types of semantics is also critical. “A lot of these devices were not born in this new environment, they were born on the board,” Daly said. “On the board, when you access memory or Flash, it always works. So the semantics of data memory expects that reads and writes always work. So we need to build in that assumption that even if a switch fails, it should still work.” Delos’ data interface has to be in silicon because the lowest possible latency is required, Daly said. Apollo would replace today’s I/O chiplet in the package next to a multi-die GPU or XPU, or as a packaged chip on a board in near-packaged or co-packaged optics systems. There is also demand for a form factor resembling a traditional NIC, which can temporarily help bridge endpoints—particularly CPUs, which may not need the bandwidth but would benefit from latency improvements—into a cluster without requiring the Apollo chiplet, Delos CEO Ed Doe told EE Times. “The common attribute across these form factors is latency, resiliency, and ultimately, the ability to bring everything in scale into a single domain,” Doe said. Delos currently has an FPGA version of the card to offer, with NPO chips to appear next year, Doe said. The chiplet version is the same tapeout as the NPO chip version, but because of XPU design cycles, the chiptlet will not appear in XPU products for another couple of years, Doe said. Delos has also combined its hardware and software into a reference architecture and offers Morpheus, a development platform that bridges customers’ pre-silicon environments and lab environments to help customers co-design system topologies based on workloads and understand the impact of changes in connectivity. Target customers for the Apollo chiplet include AI chip companies, such as hyperscalers and frontier labs working on their own silicon, Doe said. According to the company, Delos’ Mosaic cluster management software https://www.eetimes.com/startup-boosts-scale-up-to-1000-gpus-in-a-single-domain/ is already deployed in production environments using customers’ existing infrastructure. The server, Asterion, will be sampling at the end of Q4. The development platform, Morpheus, is available now. Also read: Engineering Heterogeneity at Scale https://www.eetimes.com/engineering-heterogeneity-at-scale/ Cost-Efficient AI at Scale is a Software Problem https://www.eetimes.com/cost-efficient-ai-at-scale-is-a-software-problem/ Chiplets Are The New Baseline for AI Inference Chips https://www.eetimes.com/chiplets-are-the-new-baseline-for-ai-inference-chips/