How DOCA GPUNetIO Unifies GPU-Initiated Networking Across the NVIDIA Software Stack NVIDIA's DOCA GPUNetIO has evolved from a standalone GPU-centric packet-processing framework into a unified GDA-KI foundation now integrated across NVSHMEM, NCCL, Aerial 5G SDK, UCX/NIXL, NVQLink with the Holoscan Sensor Bridge operator, Holoscan Advanced Network Operator, and DeepEP/HybridEP. The SDK lets CUDA kernels directly drive Ethernet, RDMA, Verbs, and DMA operations via GPUDirect RDMA, GPUDirect Async Kernel-Initiated (GDA-KI), and GDRCopy, keeping the CPU off the application critical path, and NVIDIA is also releasing an open-source version alongside the full DOCA SDK superset. Consolidating previously separate per-library GDA-KI RDMA implementations into one shared codebase reduces duplicated engineering and speeds feature, optimization, and bug-fix propagation across the stack. GPU applications increasingly need networking and data movement to behave like first-class GPU-controlled operations rather than host-driven services. When the CPU sits in the middle of every network transaction, it becomes a bottleneck on the critical path, adding latency and limiting how efficiently distributed applications can respond in real time. NVIDIA DOCA GPUNetIO, a GPU-centric networking SDK layer for real-time packet processing and data movement, addresses this directly. It brings together technologies such as GPUDirect RDMA, GPUDirect Async Kernel-Initiated GDA-KI and GDRCopy so CUDA kernels can directly drive Ethernet, RDMA, Verbs, and DMA operations while keeping the CPU out of the application critical path. Specifically, DOCA GPUNetIO provides both: - CPU functions on the control path to export on GPU memory transport objects created with DOCA Ethernet, DOCA Verbs, DOCA DMA, DOCA CommChannel - GPU CUDA functions on the data path to allow the creation of CUDA kernels that can manipulate the transport objects exported Since NVIDIA first introduced GPU-centric packet processing with DOCA GPUNetIO https://developer.nvidia.com/blog/inline-gpu-packet-processing-with-nvidia-doca-gpunetio/ , the framework has matured significantly, evolving from a powerful way to remove the CPU from the critical path into a unified GDA-KI foundation integrated across NVSHMEM https://github.com/NVIDIA/nvshmem , NCCL https://github.com/NVIDIA/nccl , Aerial 5G SDK https://developer.nvidia.com/aerial-sdk , UCX https://github.com/openucx/ucx / NIXL https://github.com/ai-dynamo/nixl/tree/main , NVQLink https://developer.nvidia.com/blog/nvidia-nvqlink-architecture-integrates-accelerated-computing-with-quantum-processors/ with Holoscan Sensor Bridge operator https://github.com/nvidia-holoscan/holoscan-sensor-bridge/tree/main/src/hololink/operators/gpu roce transceiver , Holoscan Advanced Network Operator https://github.com/nvidia-holoscan/holohub/tree/main/operators/advanced network , DeepEP/HybridEP https://github.com/deepseek-ai/DeepEP/tree/hybrid-ep , and others. This post covers what’s new: the open-source library, the unified architecture, and how these integrations work in practice. Why a unified GPUNetIO umbrella matters Before GPUNetIO became the shared foundation, every communication library had built its own separate implementation of GDA-KI-style, GPU-initiated RDMA. Each one worked independently, with its own assumptions, its own code, and its own maintenance burden. None of them shared anything with the others. Bringing the different GDA-KI-style implementations under the GPUNetIO umbrella provides a great advantage: instead of each library building and maintaining its own GDA-KI-based RDMA communication path, they can converge on a single GPUNetIO GDA-KI implementation that everyone can use, improve, and extend in one place. That creates a shared investment point for features, optimizations, and bug fixes, reduces duplicated engineering across SDKs, and allows innovations developed for one framework to become available to others much more quickly. In other words, GPUNetIO becomes the common GDA-KI foundation that multiple communication libraries can build on, rather than a collection of parallel, partially overlapping implementations. This unified approach is especially valuable for higher-level libraries such as NIXL/UCX, NCCL, and NVSHMEM. They are collaborating to enrich GPUNetIO, all taking and benefit from improvements made across the stack. From an ecosystem perspective, that means less duplicated code, less fragmentation in behavior, and a much stronger path for long-term evolution. GPUNetIO: Open Source vs SDK NVIDIA ships GPUNetIO in two closely related forms because the ecosystem needs both breadth and openness. The full DOCA SDK version is the superset: it is the GPUNetIO implementation documented in the DOCA Programming Guide https://networking-docs.nvidia.com/doca/sdk/doca-gpunetio and it spans the broader DOCA stack, including but not limited to Verbs, Ethernet, DMA, and Comm Channel integration, while preserving the core GPUNetIO model of moving the control path closer to the GPU and removing the CPU from the application critical path. In parallel, NVIDIA also publishes an open-source GPUNetIO project https://github.com/NVIDIA-DOCA/gpunetio as a lighter-weight, RDMA-Verbs-focused implementation designed for frameworks that want to remain fully open in how they integrate networking transports. Today, the open version is the smaller Verbs-oriented subset, while the DOCA SDK version is the superset that adds broader capabilities such as richer RDMA support, Ethernet, and DMA. That split is not about creating two divergent software stacks. It is about providing a common open-source foundation for GPU-initiated RDMA communications while still leaving a path to more advanced functionality when the DOCA SDK is present. Indeed, the open-source implementation can detect the presence of DOCA SDK https://github.com/NVIDIA-DOCA/gpunetio enable-sdk-mode at runtime and, when available, call selected closed-source DOCA SDK functions through dlopen ; if the SDK is absent, it continues to operate using the open-source implementation. This is the key architectural advantage: instead of every communication library building and maintaining its own private version of GDAKI-style RDMA plumbing, GPUNetIO becomes the shared implementation point that multiple frameworks and libraries can adopt, harden, and improve together. On the CUDA device side, the gap is intentionally small for the Verbs path: both the open and SDK variants expose a device-facing API, which keeps the GPU programming model largely aligned even as the host-side implementation scales from a lightweight open-source path to the broader DOCA SDK feature set. Programming model This section covers the key concepts and programming model elements that apply to any GPUNetIO application. CPU control path In applications that use DOCA GPUNetIO, an initial host-side CPU configuration phase is required to execute the control path, followed by a second phase in which the data path runs on the GPU through CUDA kernels. The control-path workflow generally follows this sequence: - The GPU and network devices are initialized and configured, and the required memory is allocated. - Network transport objects are created. For example, DOCA Verbs or DOCA Ethernet creates a network queue object via mlx5dv on the CPU. - A GPUNetIO CPU function exports the relevant elements of the network queue object into a descriptor stored in GPU memory and provides a GPU address for it. - The application launches a CUDA kernel, passing the GPU descriptor for the network queue as one of the input parameters. To simplify some of these operations for DOCA Verbs, a set of GPUNetIO high-level functions for example, doca gpu verbs create qp hl has been introduced to condense the steps required to create and connect RDMA QPs. These functions are present in both SDK and open-source versions. Once the control path is complete, the data path can start on the GPU, where one or more CUDA kernels use GPUNetIO CUDA functions to operate on the transport objects exported into GPU memory to send or receive traffic. GPU data path The generic structure of a CUDA kernel for GPU communication typically consists of the following steps: - Post WQEs: one or more CUDA threads post Work Queue Entries WQEs , such as RDMA Write, RDMA Read, or Ethernet Send/Recv, to the network queue object. - Ring the doorbell: one or more CUDA threads notify the network card that new WQEs are ready by writing to its registers that is, “ringing the doorbell” . - Poll CQEs optional : one or more CUDA threads wait for Completion Queue Entries CQEs to confirm that the WQEs have completed successfully. API: High-level vs low-level For Ethernet and RDMA Verbs transports, DOCA GPUNetIO offers two different levels of API: high-level and low-level. The high-level API offers implementation of complex and composite operations. As an example, on the Verbs side, the doca gpu dev verbs put signal offers a pre-implemented thread-safe combination of combining an RDMA Write Work Queue Entry WQE with RDMA Atomic Fetch & Add WQE executed per-thread scope or per-warp scope all threads in the warp cooperate to the operation posting WARP SIZE RDMA Write and just a single RDMA Atomic at the end . The function takes care about concurrent submissions of WQEs from other CUDA threads on the same network queue and rings the network card doorbell to avoid race conditions. Similarly, on the Ethernet side, another example is the doca gpu dev eth txq send , offering a thread-safe combination of posting an Ethernet Send WQE on the network queue and then ringing the doorbell at thread, warp or block scope. The same applies to the receiving side. The low-level API, by contrast, offers basic building blocks an application can use to create its own customized composite operations like posting WQEs with different opcodes, ringing the network card doorbell, poll the Completion Queue waiting for Completion Queue Entries CQE . These APIs are not thread safe, so it’s the application’s responsibility to properly synchronize the concurrent access if present to the same network queue object. Ring the doorbell Once WQEs have been posted to the network queue, the network card must be notified so it can execute them. This step is known as “ringing the doorbell.” In GPUNetIO, it refers to the CUDA kernel notifying the network card that new work is ready for execution. DOCA GPUNetIO offers several ways to ring the doorbell: - Regular doorbell: This is the canonical mode, in which the NIC registers are MMIO-mapped into the CUDA memory space, allowing CUDA threads to write to them directly. - BlueFlame: This follows the same general model as the regular doorbell, but in this case the entire WQE, rather than just a notification, is written into the NIC registers. BlueFlame is typically used in latency-sensitive applications with a small number of network queues. - CPU-assisted doorbell: In this mode, the NIC register is mapped into CPU memory rather than GPU memory. The GPU writes a notification to a shared memory area that is polled by a CPU thread. Once the CPU thread detects the GPU update, it rings the NIC doorbell. This mode is mainly useful to enable GDA-KI in systems without a direct GPU-to-NIC connection, such as DGX Spark. Examples To facilitate both API exploration and performance benchmarking, GPUNetIO offers various reference implementations. Developers can access the open-source examples https://github.com/NVIDIA-DOCA/gpunetio/tree/main/examples directly within the project repository, while the more comprehensive SDK samples https://networking-docs.nvidia.com/doca/sdk/gpunetio-sample-guide and a dedicated Ethernet-based application https://networking-docs.nvidia.com/doca/sdk/doca-gpu-packet-processing-application-guide are distributed as part of the full DOCA SDK. Furthermore, these samples have been integrated into GitHub https://github.com/NVIDIA-DOCA/doca-samples for easier accessibility in recent releases. The following sections highlight specific code snippets to illustrate how high-level and low-level primitives can be leveraged to implement similar logic using different programming semantics. Ethernet example A GPU-initiated Ethernet packet generator can be implemented using high-level API like in the code snippet below. Specifically, the high-level function doca gpu dev eth txq send facilitates the concurrent submission of WQEs across multiple CUDA blocks to a shared Send Queue, handling necessary synchronization internally to ensure thread-safe operation without requiring custom application-level locks. Also, it abstracts away the granular management of queue structures, such as tracking the specific indices of WQEs and CQEs. global void send packets struct doca gpu eth txq txq, uint8 t addr, const uint32 t mkey, const size t size, uint32 t exit cond { enum doca gpu eth send flags flags = DOCA GPUNETIO ETH SEND FLAG NONE; doca gpu dev eth ticket t out ticket; uint32 t num completed = 0; / Only the last thread in the block requests a CQE. / if threadIdx.x == blockDim.x - 1 flags = DOCA GPUNETIO ETH SEND FLAG NOTIFY; while DOCA GPUNETIO VOLATILE exit cond == 0 { doca gpu dev eth txq send