Control How Your GPU Shares Work with Green Contexts NVIDIA green contexts, available in the Driver API since CUDA 12.4, became accessible through the CUDA Runtime API starting with CUDA 13.1, letting applications explicitly partition GPU execution resources such as SMs and workqueues within a single process. The Runtime API represents green contexts with the cudaExecutionContext_t type, and cudaGreenCtxCreate() returns a handle that can be passed directly to cudaExecutionCtxStreamCreate() instead of relying on implicit thread-local device or context state. The change targets interference between concurrent workloads — for example a latency-sensitive operator alongside a throughput-oriented background kernel — by reducing false serialization and hardware context-switch overhead. GPU applications increasingly consist of multiple independent components running at the same time within a single process: a latency-sensitive operator alongside a throughput-oriented background kernel; a data preprocessing stage alongside model inference; or multiple stages of a processing workflow sharing a single GPU. Controlling how GPU resources are shared between them remains difficult. Components can interfere unpredictably, and existing tools offer limited ability to partition resources. Green contexts address this by letting applications explicitly select a subset of GPU execution resources and target work to those resources directly. Traditional CUDA contexts weren’t designed for this usage model. They are heavyweight, incur hardware context-switch overhead, and reflect assumptions from an earlier era, when GPUs were smaller and applications typically ran as a single dominant workload. Green contexts have been available in the Driver API https://docs.nvidia.com/cuda/cuda-driver-api/group CUDA GREEN CONTEXTS.html group CUDA GREEN CONTEXTS since NVIDIA CUDA 12.4. Starting with CUDA 13.1, they are also accessible through the Runtime API https://docs.nvidia.com/cuda/cuda-runtime-api/group CUDART EXECUTION CONTEXT.html group CUDART EXECUTION CONTEXT , allowing an application, within its process, to explicitly define where work runs and how execution resources are divided. How green contexts work One primary use of green contexts is SM partitioning . By assigning a specific subset of SMs to a green context, applications can target work submitted through that green context to those SMs. This can enable multiple workloads to run concurrently on the GPU without competing for the same compute units. In addition to SM partitioning, green contexts can provision workqueue resources . In the traditional model, independent stream-ordered workloads may map to the same underlying workqueues, introducing unintended serialization even when sufficient execution resources are available. By provisioning workqueues explicitly, green contexts allow applications to express expected concurrency and reduce false dependencies. Green contexts are lightweight to create and destroy, and creating or destroying one doesn’t implicitly synchronize unrelated GPU work. They also provide a more explicit programming model where applications target work to a chosen green context rather than relying only on implicit/current device state. Explicit programming model with green contexts Historically, CUDA Runtime applications typically selected a device with cudaSetDevice , created streams, and submitted work to those streams. The stream’s execution target was inferred from the current device or context for the calling thread at the time the stream was created. Green contexts make that targeting more explicit. An application creates a green context for a chosen set of resources, then creates streams from that green context. Work submitted to those streams is associated with the green context’s resources. In the CUDA Runtime API, green contexts are represented with the cudaExecutionContext t type, a Runtime abstraction for CUDA contexts. For green-context use, the important point is that cudaGreenCtxCreate returns a handle that can be passed directly to APIs such as cudaExecutionCtxStreamCreate , instead of relying on implicit thread-local device or context state. For streams, this shift is reflected directly in how applications create them. Note: For brevity, the following code snippets omit full error checking. Production code should check all CUDA Runtime API return values, use cudaGetLastError or cudaPeekAtLastError after kernel launches, and check synchronization calls such as cudaStreamSynchronize for asynchronous execution errors. Traditional Runtime flow: stream target comes from current device state cudaSetDevice device ; cudaStream t s; cudaStreamCreate &s ; kernel<<