AOTInductor Triton SASS Inspection
AOTInductor, PyTorch's ahead-of-time inference backend compiler, emits Triton kernels as cubin files stored in the pt2 zip archive, which developers can unzip and disassemble with nvdisasm or cuobjdum…
AOTInductor, PyTorch's ahead-of-time inference backend compiler, emits Triton kernels as cubin files stored in the pt2 zip archive, which developers can unzip and disassemble with nvdisasm or cuobjdum…
NVIDIA's CUDA_DEVICE_MAX_CONNECTIONS environment variable defaults to 8 concurrent hardware connections to the GPU, capping real concurrency regardless of how many CUDA streams a developer creates, ac…
A technical blog post analyzes how CUDA thread block swizzle algorithms change L2 cache residency and GEMM kernel performance by remapping program IDs to tile locations. The post compares four swizzle…
A blog post explains the Residual-Quantized Variational Autoencoder (RQ-VAE), an architecture that extends the Vector Quantized Variational Autoencoder (VQ-VAE) introduced in arXiv paper 1711.00937 to…
NVIDIA's CUDA Multi-Process Service (MPS) enables concurrent kernel execution across multiple processes, improving GPU utilization compared to default time-slicing, according to a technical blog post …
AOTInductor, a PyTorch compiler backend, supports in-place input mutations in optimized inference engines, contrary to the assumption that it relies solely on functionalization. The key is to explicit…
A new GitHub example demonstrates how to store AOTInductor model weights outside the shared library and update them at inference runtime in a thread-safe manner, addressing the lack of documentation f…
NVIDIA CUDA Graph, a feature designed to reduce CPU overhead and GPU bubbles in single-stream kernel launches, may not improve latency or GPU utilization in multi-stream execution when bubbles stem fr…
PyTorch's torch._assert_async API enables asynchronous assertion on GPU device streams, avoiding graph breaks in torch.compile. The API accepts a boolean tensor and reports failures only when the GPU …
NVIDIA's CUDA programming guide introduces a swizzling technique to eliminate shared memory bank conflicts without wasting memory, using an XOR-based formula that rearranges shared memory indices. The…
Vector quantization and product quantization offer an alternative to scaling-based neural network quantization, achieving high compression ratios by mapping high-dimensional vectors to a finite codebo…
PyTorch's multiprocessing module provides a CUDA Inter-Process Communication (IPC) API that enables sharing model weights across multiple processes for inference, avoiding duplication in GPU VRAM. The…
Predicated execution and conditional execution are two approaches to handling control flow in programming, with predicated execution running all instructions but committing only those meeting a condit…
TorchRec's KeyedJaggedTensor, designed to efficiently combine sparse features in recommendation systems without padding, introduces GPU-CPU synchronization that degrades system performance. The data t…
PyTorch users can now implement custom operations in C++ and CUDA for use in both Python and C++ inference programs, with automatic device dispatch between CPU and CUDA implementations. The approach s…
PyTorch AOTInductor now compiles exported programs with hybrid CPU-GPU execution plans into a single executable package, eliminating the need to manually split models into separate device sub-models. …
PyTorch has introduced transparent tracing and compilation for Triton kernels, allowing custom operations to be visible to the compiler for optimization. The framework now supports compiling Triton ke…
PyTorch introduced a "fake export" method that allows developers to verify the exportability of large deep learning models using `torch.export` APIs without requiring actual GPU memory. The approach u…