{"slug": "retrieval-scoring-and-decoding-shape-performance-and-stability-in-llm-based", "title": "Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation", "summary": "A new arXiv study (2609.00086v1) evaluating LLM-based rerankers for conversational recommendation on the ReDial benchmark found that the best proprietary reranker achieves NDCG@10 of 0.1497 with a shared semantic top-250 candidate pool and strict candidate-aware scoring, versus 0.0939 for the strongest non-LLM baseline, but no open-weight LLM outperforms a tuned shallow autoencoder under the same protocol. The study also reports that switching from semantic to collaborative-filtering candidates raises NDCG@10 by more than 50% for the strongest rerankers, and that raising temperature from 0 to 1.0 increases top-10 Jaccard distance from 0.0900 to 0.1240 for the best proprietary reranker with negligible NDCG change, leading the authors to urge reporting candidate generation, pool size, scoring policy, and decoding configuration as standard fields.", "body_md": "arXiv:2609.00086v1 Announce Type: new\nAbstract: Large language models (LLMs) are increasingly used as rerankers in conversational recommender systems, yet measured gains depend strongly on the retrieval and inference protocol. On the ReDial conversational movie recommendation benchmark, we compare proprietary, open-weight, and fine-tuned LLM rerankers with collaborative-filtering and sequential baselines in a shared retrieve-then-rerank pipeline. We vary candidate-pool size, first-stage retriever, and decoding temperature. With a shared semantic top-250 candidate pool and strict candidate-aware scoring, the best proprietary reranker reaches NDCG@10 of 0.1497, compared with 0.0939 for the strongest non-LLM baseline. The same reranker reaches 0.2925 in zero-shot generation, showing that unconstrained scoring can yield a much larger apparent advantage than matched-pool evaluation. No evaluated open-weight LLM outperforms the tuned shallow autoencoder baseline under this protocol. For the strongest proprietary and open-weight rerankers, switching from semantic to collaborative-filtering candidates raises NDCG@10 by more than 50%, showing that measured reranker performance is highly sensitive to candidate generation. For the best proprietary reranker, raising temperature from 0 to 1.0 increases top-10 Jaccard distance from 0.0900 to 0.1240 while mean NDCG@10 changes negligibly, whereas weaker LLMs show larger degradation. These ReDial results support treating candidate generation, candidate-pool size, scoring policy, and decoding configuration as required reporting fields rather than implementation details.", "url": "https://wpnews.pro/news/retrieval-scoring-and-decoding-shape-performance-and-stability-in-llm-based", "canonical_source": "https://arxiv.org/abs/2609.00086", "published_at": "2026-09-02 04:00:00+00:00", "updated_at": "2026-09-02 04:26:10.776112+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "natural-language-processing"], "entities": ["arXiv", "ReDial"], "alternates": {"html": "https://wpnews.pro/news/retrieval-scoring-and-decoding-shape-performance-and-stability-in-llm-based", "markdown": "https://wpnews.pro/news/retrieval-scoring-and-decoding-shape-performance-and-stability-in-llm-based.md", "text": "https://wpnews.pro/news/retrieval-scoring-and-decoding-shape-performance-and-stability-in-llm-based.txt", "jsonld": "https://wpnews.pro/news/retrieval-scoring-and-decoding-shape-performance-and-stability-in-llm-based.jsonld"}}