cd/sources/leimao-auto-discovered· home› sources› Leimao (auto-discovered)
cat /sources/leimao-auto-discovered.feed | wc -l → 18

Leimao (auto-discovered)

articles 18 domain leimao.github.io → feed RSS
06:16
2026-10-02
leimao.github.io
artificial-intelligence

AOTInductor Triton SASS Inspection

AOTInductor, PyTorch's ahead-of-time inference backend compiler, emits Triton kernels as cubin files stored in the pt2 zip archive, which developers can unzip and disassemble with nvdisasm or cuobjdum…

16:30
2026-09-25
leimao.github.io
ai-infrastructure

CUDA Device Max Connections

NVIDIA's CUDA_DEVICE_MAX_CONNECTIONS environment variable defaults to 8 concurrent hardware connections to the GPU, capping real concurrency regardless of how many CUDA streams a developer creates, ac…

07:00
2026-09-20
leimao.github.io
ai-infrastructure

CUDA Thread Block Swizzle

A technical blog post analyzes how CUDA thread block swizzle algorithms change L2 cache residency and GEMM kernel performance by remapping program IDs to tile locations. The post compares four swizzle…

04:28
2026-09-14
leimao.github.io
machine-learning

Residual-Quantized Variational Autoencoder

A blog post explains the Residual-Quantized Variational Autoencoder (RQ-VAE), an architecture that extends the Vector Quantized Variational Autoencoder (VQ-VAE) introduced in arXiv paper 1711.00937 to…

05:31
2026-09-08
leimao.github.io
ai-infrastructure

CUDA Multi-Process Service

NVIDIA's CUDA Multi-Process Service (MPS) enables concurrent kernel execution across multiple processes, improving GPU utilization compared to default time-slicing, according to a technical blog post …

15:55
2026-08-31
leimao.github.io
ai-infrastructure

AOTInductor Input Mutation

AOTInductor, a PyTorch compiler backend, supports in-place input mutations in optimized inference engines, contrary to the assumption that it relies solely on functionalization. The key is to explicit…

04:48
2026-08-27
leimao.github.io
ai-infrastructure

AOTInductor External Weight Storage and Weight Streaming Update

A new GitHub example demonstrates how to store AOTInductor model weights outside the shared library and update them at inference runtime in a thread-safe manner, addressing the lack of documentation f…

15:12
2026-08-19
leimao.github.io
ai-infrastructure

CUDA Graph In The Context of Multi-Stream Execution

NVIDIA CUDA Graph, a feature designed to reduce CPU overhead and GPU bubbles in single-stream kernel launches, may not improve latency or GPU utilization in multi-stream execution when bubbles stem fr…

01:50
2026-08-14
leimao.github.io
developer-tools

PyTorch Asynchronous Assert

PyTorch's torch._assert_async API enables asynchronous assertion on GPU device streams, avoiding graph breaks in torch.compile. The API accepts a boolean tensor and reports failures only when the GPU …

14:13
2026-08-13
leimao.github.io
developer-tools

CUDA Shared Memory Swizzling

NVIDIA's CUDA programming guide introduces a swizzling technique to eliminate shared memory bank conflicts without wasting memory, using an XOR-based formula that rearranges shared memory indices. The…

07:00
2026-08-07
leimao.github.io
machine-learning

Vector Quantization and Product Quantization

Vector quantization and product quantization offer an alternative to scaling-based neural network quantization, achieving high compression ratios by mapping high-dimensional vectors to a finite codebo…

07:00
2026-07-01
leimao.github.io
machine-learning

Predicated Execution VS Conditional Execution

Predicated execution and conditional execution are two approaches to handling control flow in programming, with predicated execution running all instructions but committing only those meeting a condit…

07:00
2026-06-05
leimao.github.io
machine-learning

Synchronizations With TorchRec KeyedJaggedTensor

TorchRec's KeyedJaggedTensor, designed to efficiently combine sparse features in recommendation systems without padding, introduces GPU-CPU synchronization that degrades system performance. The data t…

15:39
2026-06-01
leimao.github.io
machine-learning

PyTorch Custom Operation

PyTorch users can now implement custom operations in C++ and CUDA for use in both Python and C++ inference programs, with automatic device dispatch between CPU and CUDA implementations. The approach s…

07:00
2026-05-28
leimao.github.io
machine-learning

PyTorch AOTInductor Hybrid Lowering

PyTorch AOTInductor now compiles exported programs with hybrid CPU-GPU execution plans into a single executable package, eliminating the need to manually split models into separate device sub-models. …

07:00
2026-05-22
leimao.github.io
machine-learning

PyTorch Triton Kernel Transparent Tracing and Compilation

PyTorch has introduced transparent tracing and compilation for Triton kernels, allowing custom operations to be visible to the compiler for optimization. The framework now supports compiling Triton ke…

07:00
2026-05-17
leimao.github.io
machine-learning

PyTorch Fake Export

PyTorch introduced a "fake export" method that allows developers to verify the exportability of large deep learning models using `torch.export` APIs without requiring actual GPU memory. The approach u…