{"slug": "hllm-single-pass-decoding-for-generative-reranking", "title": "hLLM: Single Pass Decoding for Generative Reranking", "summary": "Researchers introduced hLLM (Hungarian LLM), a decoding strategy that generates rankings from large language models in a single forward pass, achieving 28 ms end-to-end inference, a 64x speed-up over autoregressive decoding while maintaining ranking quality on par with the teacher model. The method reads an N x K item-position score matrix from the LLM's prefill hidden states and decodes ordinals via the Hungarian algorithm, ensuring a valid permutation by construction.", "body_md": "arXiv:2609.01807v1 Announce Type: new\nAbstract: Large language models (LLMs) achieve state-of-the-art generative ranking quality, but the ranking they produce must be decoded, and autoregressive decoding spends one sequential forward pass per emitted token. We observe that the only tokens a ranker must emit are the $N$ ordinal values naming the items in ranked order, and that this narrow, permutation-structured output format admits decoding strategies which are much more efficient than left-to-right generation. We introduce hLLM (Hungarian LLM), a format-specialized decoding strategy that decodes all $N$ ordinals in $O(1)$ forward passes. hLLM reads an $N \\times K$ item-position score matrix off the LLM's prefill hidden states with a lightweight self-attention head, then decodes the ordinals as the optimal bipartite assignment of that matrix via the Hungarian algorithm, yielding a valid permutation by construction rather than by repair. Through a systematic study of training signals and backbone adaptation, we show that LoRA-based fine-tuning combined with teacher ranking distillation reaches 28 ms end-to-end inference, a speed-up of $64\\times$ while maintaining ranking quality on par with the teacher. We provide a complete ablation decomposing the contributions of architecture, training signal, and backbone adaptation. Our framework connects generative ranking to combinatorial optimization, opening a path toward other $O(1)$-decode mechanisms for real-time ranking.", "url": "https://wpnews.pro/news/hllm-single-pass-decoding-for-generative-reranking", "canonical_source": "https://arxiv.org/abs/2609.01807", "published_at": "2026-09-03 04:00:00+00:00", "updated_at": "2026-09-03 04:23:34.893059+00:00", "lang": "en", "topics": ["large-language-models", "artificial-intelligence", "ai-research"], "entities": ["hLLM", "Hungarian LLM", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/hllm-single-pass-decoding-for-generative-reranking", "markdown": "https://wpnews.pro/news/hllm-single-pass-decoding-for-generative-reranking.md", "text": "https://wpnews.pro/news/hllm-single-pass-decoding-for-generative-reranking.txt", "jsonld": "https://wpnews.pro/news/hllm-single-pass-decoding-for-generative-reranking.jsonld"}}