A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation
Researchers proposed FIBER, a thread-register decoupled GPU execution model that extends the SIMT architecture to improve tensor computation efficiency, achieving a 2.25x end-to-end speedup on Ampere, 1.8x on Hopper, and…