07:00
2026-09-20
leimao.github.io
ai-infrastructure
CUDA Thread Block Swizzle
A technical blog post analyzes how CUDA thread block swizzle algorithms change L2 cache residency and GEMM kernel performance by remapping program IDs to tile locations. The post compares four swizzleβ¦