{"slug": "polaris-learning-to-generate-table-descriptions-from-retrieval-feedback", "title": "Polaris: Learning to Generate Table Descriptions from Retrieval Feedback", "summary": "Researchers introduced Polaris, a system that trains a large language model to generate table descriptions optimized for retrieval effectiveness using Direct Preference Optimization (DPO) on existing table retrieval benchmarks. Polaris outperforms the state-of-the-art AutoDDG solution, often by a significant margin, by expanding abbreviated table and column names before generation to reduce vocabulary mismatch.", "body_md": "arXiv:2608.17171v1 Announce Type: new\nAbstract: Many table-centric NLP tasks such as NL2SQL first retrieve relevant tables from large collections using keyword search. Recent work uses LLMs to generate natural-language table descriptions to improve retrieval, but they are typically optimized for fluency rather than retrieval effectiveness. We present Polaris, a system that trains an LLM to generate table descriptions directly from retrieval feedback. Our key insight is that existing table retrieval benchmarks already contain the supervision needed for this task: given query-table relevance judgments, we generate multiple candidate descriptions for each table, rank them by their BM25 retrieval effectiveness, and use the resulting preference pairs to fine-tune the LLM with Direct Preference Optimization (DPO). Polaris further expands abbreviated table and column names before generation to reduce vocabulary mismatch. Extensive experiments show that Polaris outperforms the state-of-the-art AutoDDG solution, often by a significant margin. More broadly, our results demonstrate that retrieval benchmarks can be repurposed as supervision for training LLMs to generate retrieval-oriented metadata.", "url": "https://wpnews.pro/news/polaris-learning-to-generate-table-descriptions-from-retrieval-feedback", "canonical_source": "https://www.machinebrief.com/news/polaris-learning-to-generate-table-descriptions-from-retriev-ejm3", "published_at": "2026-08-19 04:00:00+00:00", "updated_at": "2026-08-19 04:11:20.928767+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "natural-language-processing", "ai-research"], "entities": ["Polaris", "AutoDDG", "Direct Preference Optimization (DPO)", "BM25", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/polaris-learning-to-generate-table-descriptions-from-retrieval-feedback", "markdown": "https://wpnews.pro/news/polaris-learning-to-generate-table-descriptions-from-retrieval-feedback.md", "text": "https://wpnews.pro/news/polaris-learning-to-generate-table-descriptions-from-retrieval-feedback.txt", "jsonld": "https://wpnews.pro/news/polaris-learning-to-generate-table-descriptions-from-retrieval-feedback.jsonld"}}