{"slug": "control-how-your-gpu-shares-work-with-green-contexts", "title": "Control How Your GPU Shares Work with Green Contexts", "summary": "NVIDIA green contexts, available in the Driver API since CUDA 12.4, became accessible through the CUDA Runtime API starting with CUDA 13.1, letting applications explicitly partition GPU execution resources such as SMs and workqueues within a single process. The Runtime API represents green contexts with the cudaExecutionContext_t type, and cudaGreenCtxCreate() returns a handle that can be passed directly to cudaExecutionCtxStreamCreate() instead of relying on implicit thread-local device or context state. The change targets interference between concurrent workloads — for example a latency-sensitive operator alongside a throughput-oriented background kernel — by reducing false serialization and hardware context-switch overhead.", "body_md": "GPU applications increasingly consist of multiple independent components running at the same time within a single process: a latency-sensitive operator alongside a throughput-oriented background kernel; a data preprocessing stage alongside model inference; or multiple stages of a processing workflow sharing a single GPU.\n\nControlling how GPU resources are shared between them remains difficult. Components can interfere unpredictably, and existing tools offer limited ability to partition resources.\n\nGreen contexts address this by letting applications explicitly select a subset of GPU execution resources and target work to those resources directly. Traditional CUDA contexts weren’t designed for this usage model. They are heavyweight, incur hardware context-switch overhead, and reflect assumptions from an earlier era, when GPUs were smaller and applications typically ran as a single dominant workload.\n\nGreen contexts have been available in the [Driver API](https://docs.nvidia.com/cuda/cuda-driver-api/group__CUDA__GREEN__CONTEXTS.html#group__CUDA__GREEN__CONTEXTS) since NVIDIA CUDA 12.4. Starting with CUDA 13.1, they are also accessible through the [Runtime API](https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__EXECUTION__CONTEXT.html#group__CUDART__EXECUTION__CONTEXT), allowing an application, within its process, to explicitly define where work runs and how execution resources are divided.\n\n## How green contexts work\n\nOne primary use of green contexts is **SM partitioning**. By assigning a specific subset of SMs to a green context, applications can target work submitted through that green context to those SMs. This can enable multiple workloads to run concurrently on the GPU without competing for the same compute units.\n\nIn addition to SM partitioning, **green contexts can provision workqueue resources**. In the traditional model, independent stream-ordered workloads may map to the same underlying workqueues, introducing unintended serialization even when sufficient execution resources are available. By provisioning workqueues explicitly, green contexts allow applications to express expected concurrency and reduce false dependencies.\n\nGreen contexts are lightweight to create and destroy, and creating or destroying one doesn’t implicitly synchronize unrelated GPU work. They also provide a more explicit programming model where applications target work to a chosen green context rather than relying only on implicit/current device state.\n\n## Explicit programming model with green contexts\n\nHistorically, CUDA Runtime applications typically selected a device with `cudaSetDevice()`, created streams, and submitted work to those streams. The stream’s execution target was inferred from the current device or context for the calling thread at the time the stream was created.\n\nGreen contexts make that targeting more explicit. An application creates a green context for a chosen set of resources, then creates streams from that green context. Work submitted to those streams is associated with the green context’s resources.\n\nIn the CUDA Runtime API, green contexts are represented with the `cudaExecutionContext_t` type, a Runtime abstraction for CUDA contexts. For green-context use, the important point is that `cudaGreenCtxCreate()` returns a handle that can be passed directly to APIs such as `cudaExecutionCtxStreamCreate()`, instead of relying on implicit thread-local device or context state.\n\nFor streams, this shift is reflected directly in how applications create them.\n\nNote: For brevity, the following code snippets omit full error checking. Production code should check all CUDA Runtime API return values, use `cudaGetLastError()` or `cudaPeekAtLastError()` after kernel launches, and check synchronization calls such as `cudaStreamSynchronize()` for asynchronous execution errors.\n\n**Traditional Runtime flow: stream target comes from current device state**\n\n```\ncudaSetDevice(device);\ncudaStream_t s;\ncudaStreamCreate(&s);\nkernel<<<grid, block, 0, s>>>();\n```\n\n**Green-context flow: stream target is selected explicitly**\n\n```\ncudaExecutionContext_t greenCtx;\ncudaGreenCtxCreate(&greenCtx, desc, device, 0);\n\ncudaStream_t s;\ncudaExecutionCtxStreamCreate(&s, greenCtx, 0, 0);\nkernel<<<grid, block, 0, s>>>();\n```\n\nThe code changes are minimal, but the mental model is clearer. Instead of relying on hidden thread-local device state, applications explicitly choose a green context and create streams for it. The rest of the application code can continue using the same stream-based CUDA programming model.\n\nApplications that want to target the full device can continue using the traditional Runtime model. For APIs that take an explicit context handle, they can obtain the device-wide context with `cudaDeviceGetExecutionCtx()`.\n\n# Example\n\nGPU workloads may often require a small, latency-sensitive kernel to share a device with a larger throughput-oriented worker. A canonical example is communication/GEMM overlap in distributed training and inference, or latency-sensitive operators in AI sensor processing platforms like [NVIDIA Holoscan](https://docs.nvidia.com/holoscan/index.htm) that require starting as soon as possible.\n\nOne of the standard tools for prioritizing such critical work is [CUDA stream priority](https://docs.nvidia.com/cuda/cuda-programming-guide/03-advanced/advanced-host-programming.html#stream-priorities). Unfortunately, setting priority isn’t enough to guarantee immediate execution of a higher priority kernel, when the bulk kernel fully occupies all of the GPU’s SMs. If bulk kernels are queued up and the critical kernel arrives, the scheduler hands critical kernel threads blocks to the next SM(s) that free(s) up. Stream priorities can’t preempt a block that’s already executing on an SM, so it needs to wait for some of the blocks to drain.\n\nLet’s create a sample and look at some results. The setup code is the new piece we need to write:\n\n```\n// 1. Query all SMs on the device.\ncudaDevResource all_sm {};\ncudaDeviceGetDevResource(dev, &all_sm, cudaDevResourceTypeSm);\n\n// 2. Carve out one critical group (arch's coscheduled alignment).\ncudaDevSmResourceGroupParams crit_params {};\ncrit_params.smCount = all_sm.sm.smCoscheduledAlignment;\n\ncudaDevResource critical_res{}, remaining_res{};\ncudaDevSmResourceSplit(&critical_res, 1, &all_sm, &remaining_res, 0, &crit_params);\n\n// 3. Pack each partition with its own workqueue-config resource so queue\n//    pressure is isolated too, then generate the descriptor.\ncudaDevResource wq {};\nwq.type = cudaDevResourceTypeWorkqueueConfig;\nwq.wqConfig.device = dev;\nwq.wqConfig.sharingScope = cudaDevWorkqueueConfigScopeGreenCtxBalanced;\nwq.wqConfig.wqConcurrencyLimit = 2;\n\ncudaDevResource crit_pack[2] = { critical_res, wq };\ncudaDevResourceDesc_t crit_desc{};\ncudaDevResourceGenerateDesc(&crit_desc, crit_pack, 2);\n\ncudaDevResource bulk_pack[2] = { remaining_res, wq };\ncudaDevResourceDesc_t bulk_desc{};\ncudaDevResourceGenerateDesc(&bulk_desc, bulk_pack, 2);\n\n// 4. Create each green context, then a stream on it.\ncudaExecutionContext_t crit_ctx{};\ncudaGreenCtxCreate(&crit_ctx, crit_desc, dev, 0);\n\ncudaStream_t crit_stream{};\ncudaExecutionCtxStreamCreate(&crit_stream, crit_ctx, cudaStreamNonBlocking, prio_high);\n\ncudaExecutionContext_t bulk_ctx{};\ncudaGreenCtxCreate(&bulk_ctx, bulk_desc, dev, 0);\n\ncudaStream_t bulk_stream{};\ncudaExecutionCtxStreamCreate(&bulk_stream, bulk_ctx, cudaStreamNonBlocking, prio_low); // define prio_low\n```\n\nWhile the rest of the code should look familiar:\n\n```\n// Saturate the bulk partition.\nfor (int i = 0; i < BULK_LAUNCHES; ++i) {\n    bulk_kernel<<<bulk_grid, bulk_block, 0, bulk_stream>>>(\n        d_bulk, N_BULK, BULK_ITERS);\n}\n// Launch the critical kernel on its own partition.\ncudaEventRecord(t_start, crit_stream);\ncritical_kernel<<<crit_grid, crit_block, 0, crit_stream>>>(d_crit, N_CRIT);\ncudaEventRecord(t_stop, crit_stream);\ncudaStreamSynchronize(crit_stream);\nfloat crit_ms = 0.0f;\ncudaEventElapsedTime(&crit_ms, t_start, t_stop);\n```\n\nThe stream carries the partition and the kernel sees only the SMs it’s allowed to run on.\n\nWe test three modes, same critical + bulk workload in each:\n\n- **Mode A** : Green-ctx partition, high-priority critical stream\n- **Mode B** : Default context, high-priority critical stream (no partition)\n- **Mode C** : Default context, normal priority on both streams (no stream priorities; no partition)\n\nThe critical kernel is a tiny integer workload. The bulk kernel is 20 back-to-back launches of a 4M-thread `sqrtf` loop designed to saturate the device. We measure the wall-clock latency of the critical kernel while bulk is running.\n\nOn an NVIDIA Blackwell GPU with 148 SMs:\n\n| Mode | Critical Kernel Latency | \n|---|---|\n| Green-ctx (8 SMs for critical / 140 SMs for bulk) | 0.007 ms | \n| Default ctx, high-priority critical | 0.140 ms | \n| Default ctx, equal priority | 3.727 ms | \n\n*Table 1. Critical kernel latency by execution mode*\n\nStream priority alone is a huge win over not using it. Roughly 27x faster than equal-priority streams competing for the same SMs. But priority has a ceiling. An additional 20x is wasted waiting for bulk blocks to drain, even with the GPU scheduler prioritizing the critical kernel. Green contexts skip that wait entirely because the critical SMs are dedicated to that workload. The tradeoff is that less SMs are available for the bulk kernel’s execution.\n\nThe same skeleton can be used to achieve desired concurrency rather than latency improvement, such as overlapping a communication kernel with a GEMM kernel so both make progress. Advanced use-cases often use both green context partitions and stream priorities to achieve desired performance targets.\n\n## When to use green contexts\n\nThey are most useful when your application:\n\n- Runs multiple independent workloads concurrently on the same GPU.\n- Needs more predictable performance rather than best-effort stream scheduling\n- Wants explicit control over how GPU execution resources are divided\n- Sees degraded performance due to limited control over workqueues\n- Is evolving toward pipeline-based or multi-component execution within a single process\n\n## Getting started\n\n- Build your application with CUDA 13.1 or newer.\n- Query device resources and create a green context for the resource partition you want to target.\n- Create streams for that green context and submit work using the same stream-based CUDA programming model.\n- For APIs that take an explicit context handle and should target the full device, use `cudaDeviceGetExecutionCtx()` .\n\nGreen contexts are opt-in and additive. Existing applications can continue to target the full device unchanged and adopt green contexts incrementally where finer control over execution is needed.\n\nFor additional details and examples, see the [CUDA Programming Guide](https://docs.nvidia.com/cuda/cuda-programming-guide/04-special-topics/green-contexts.html) and [CUDA Runtime API](https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__EXECUTION__CONTEXT.html#group__CUDART__EXECUTION__CONTEXT) documentation, and try green contexts today.", "url": "https://wpnews.pro/news/control-how-your-gpu-shares-work-with-green-contexts", "canonical_source": "https://developer.nvidia.com/blog/control-how-your-gpu-shares-work-with-green-contexts/", "published_at": "2026-10-06 15:00:00+00:00", "updated_at": "2026-10-06 15:17:09.645418+00:00", "lang": "en", "topics": ["ai-infrastructure", "developer-tools"], "entities": ["NVIDIA", "CUDA", "CUDA 12.4", "CUDA 13.1", "cudaExecutionContext_t", "cudaGreenCtxCreate()", "cudaExecutionCtxStreamCreate()", "cudaSetDevice()"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/control-how-your-gpu-shares-work-with-green-contexts", "markdown": "https://wpnews.pro/news/control-how-your-gpu-shares-work-with-green-contexts.md", "text": "https://wpnews.pro/news/control-how-your-gpu-shares-work-with-green-contexts.txt", "jsonld": "https://wpnews.pro/news/control-how-your-gpu-shares-work-with-green-contexts.jsonld"}}