Why INT4 Weight-Only Quantization Doesn't Speed Up Prefill
A developer's analysis shows that INT4 weight-only quantization speeds up decode but not prefill, because prefill is compute-bound while decode is memory-bound. The crossover point where a GEMM become…