{"slug": "metas-custom-transport-protocol-embraces-packet-chaos-to-boost-ai-throughput", "title": "Meta’s custom transport protocol embraces packet chaos to boost AI throughput", "summary": "Meta has developed MetaRoCE, a custom RDMA transport protocol that sprays packets out of order to boost AI throughput on Ethernet-based scale-out networks, maintaining around 86% throughput at 1% packet loss in tests on AMD hardware with Pensando NICs across a 64-node cluster. Meta is opening the full spec to the Open Compute Project (OCP) at the OCP Global Summit in mid-October, aiming to improve network utilization and reduce latency for AI workloads.", "body_md": "Meta developed its own transport protocol for higher-throughput Ethernet-based scale-out networks running AI workloads.\n\nThe aptly named [MetaRoCE](https://engineering.fb.com/2026/08/24/networking-traffic/metaroce-rdma-transport-ai-ethernet/) is a remote direct memory access (RDMA) transport designed entirely from scratch. Unlike conventional high-performance protocols, MetaRoCE deliberately sprays packets out of order to achieve mass cross-sectional bandwidth and network utilization. Each write carries its destination in every packet, meaning a send will land correctly even if messages arrive out of order.\n\nMeta engineers argued that standard RDMA’s sequential sending of packets stalls network speeds, with receiving network interface card (NIC) hardware being forced to wait until the missing packet arrives to patch the gap. With MetaRoCE, data is written straight to its final memory location as it lands, or as the hyperscaler put it, “the fabric sees packets, but the NIC sees intent.”\n\n“Traditional architectures centralize intelligence in the fabric, relying on switches to enforce losslessness and maintain order,” a company blog post explains. “By moving intelligence to the endpoint, MetaRoCE decomposes the network into many fine-grained logical paths, each with its own real-time telemetry.”\n\nThe Facebook parent tested MetaRoCE on AMD hardware, leveraging the Pensando programmable NICs across a 64-node cluster. Results showed that in packet loss conditions that would typically degrade a protocol like RDMA over converged Ethernet version 2 (RoCEv2), MetaRoCE maintained around 86% throughput at just 1% packet loss.\n\nMeta’s home-brewed protocol was found to have consistently maintained higher throughput and lower flow completion times than rival alternatives, even when scaled across four- and eight-plane topologies with up to 4,000 concurrent connections.\n\n“By designing for loss from day one and pushing intelligence to the edge, you get a transport that performs better in ideal conditions and degrades gracefully when things go wrong,” Meta engineers wrote.\n\nMetaRoCE was teased a few weeks prior at [Advancing AI](https://www.sdxcentral.com/control-plane/5-things-we-learned-from-amds-advancing-ai-2026/), where senior director for data center and AI networking Omar Baldonado and AMD’s Soni Jiandani told [ The Cube](https://www.youtube.com/watch?v=EUpd1AFZdoE&t=1201s) that the Pensando NIC running the transport would make vast pools of graphic processing unit (GPU) resources more accessible and efficient.\n\nAlthough designed for scale-out networking, Meta said the concept could be applied to scale-up inside the rack to help remove sources of latency, including reorder buffers and priority-based flow control (PFC), as packets are sent out unsequentially by design. While on the distributed computing or scale-across side, Meta engineers claim it could be used to reduce the time between disparate links.\n\nMeta confirmed it was opening the full spec to the Open Compute Project (OCP), a move that is set to coincide with the OCP Global Summit in mid-October, and that while it was tested on AMD NICs, additional implementations were planned with “other vendors.”\n\n“We’re building this in the open because the challenges ahead benefit from broad industry collaboration. If you’re building NICs, switches, or AI infrastructure, we invite you to join us,” Meta’s blog concludes.", "url": "https://wpnews.pro/news/metas-custom-transport-protocol-embraces-packet-chaos-to-boost-ai-throughput", "canonical_source": "https://www.sdxcentral.com/news/metas-custom-transport-protocol-embraces-packet-chaos-to-boost-ai-throughput/", "published_at": "2026-08-25 12:37:14+00:00", "updated_at": "2026-08-25 12:43:33.845293+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-infrastructure", "ai-research"], "entities": ["Meta", "MetaRoCE", "AMD", "Pensando", "Open Compute Project", "Omar Baldonado", "Soni Jiandani", "The Cube"], "alternates": {"html": "https://wpnews.pro/news/metas-custom-transport-protocol-embraces-packet-chaos-to-boost-ai-throughput", "markdown": "https://wpnews.pro/news/metas-custom-transport-protocol-embraces-packet-chaos-to-boost-ai-throughput.md", "text": "https://wpnews.pro/news/metas-custom-transport-protocol-embraces-packet-chaos-to-boost-ai-throughput.txt", "jsonld": "https://wpnews.pro/news/metas-custom-transport-protocol-embraces-packet-chaos-to-boost-ai-throughput.jsonld"}}