Accelerating LLM Inference via Vector Index Based Output Embeddings
Researchers propose replacing dense output projections in LLMs with an HNSW-based vector index to speed up token selection, achieving up to 82% faster end-to-end decoding throughput for Gemma 3 270M o…