00:00
2026-09-01
sailresearch.com
artificial-intelligence
Improving Decode Throughput on Intel Gaudi 3
Intel Gaudi 3's vLLM-Gaudi PagedAttention implementation underutilizes hardware for sliding window attention, but optimizations for serving Gemma 4 31B expanded usable KV cache by 3.5×, raised decode …