{"slug": "multi-label-topic-assignment-via-llm-distillation-a-comparative-analysis-of-vs", "title": "Multi-Label Topic Assignment via LLM Distillation: A Comparative Analysis of Generative vs. Discriminative Student Models", "summary": "A comparative study of multi-label topic assignment for user-generated content found that while discriminative models (DeBERTa-V3 and ModernBERT) outperform ultra-lightweight generative models on structured product reviews, even the smallest 1B-parameter generative model surpasses discriminative baselines on complex multi-turn conversational data. The analysis, spanning Small Language Model parameter scales of 1B, 4B, and 8B, also found that discriminative baselines suffer a 35% drop in Macro-F1 under massive label-set expansion up to 112 topics and severe long-tail distributions, while generative models maintain robust performance. The authors report successful production deployment of the optimized models across product review and conversational domains at a global marketplace scale with strict latency compliance.", "body_md": "arXiv:2610.09063v1 Announce Type: new \nAbstract: Multi-label topic assignment for user-generated content (UGC) -- including product reviews and buyer-seller conversations -- poses unique scalability challenges in large-scale e-commerce due to informal language, extreme label sparsity, and rapidly evolving taxonomies. While utilizing Large Language Models (LLMs) as labeling oracles to distill ground-truth data has emerged as an industry standard to bypass prohibitive manual annotation costs, determining the optimal, low-latency architecture for the resulting student models remains an open challenge. To address this, we conduct a comprehensive evaluation across Small Language Model (SLM) parameter scales (1B, 4B, and 8B) and architectural paradigms (causal generative versus bidirectional discriminative). Comparing generative text-to-label classifiers against discriminative baselines (DeBERTa-V3 and ModernBERT), our analysis reveals a crucial data-dependent trade-off: while discriminative models outperform ultra-lightweight generative models on structured product reviews, even the smallest 1B generative model surpasses discriminative baselines on complex, multi-turn conversational data. Furthermore, generative models maintain robust performance under massive label-set expansion (up to 112 topics) and severe long-tail distributions, whereas discriminative baselines suffer a 35% drop in Macro-F1 at scale. Finally, we detail the successful production deployment of these optimized models across both product review and conversational domains, demonstrating strict latency compliance and tangible business impact at a global marketplace scale.", "url": "https://wpnews.pro/news/multi-label-topic-assignment-via-llm-distillation-a-comparative-analysis-of-vs", "canonical_source": "https://www.machinebrief.com/news/multi-label-topic-assignment-via-llm-distillation-a-comparat-dtp6", "published_at": "2026-10-08 04:00:00+00:00", "updated_at": "2026-10-08 05:16:42.262739+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "natural-language-processing", "ai-research"], "entities": ["DeBERTa-V3", "ModernBERT", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/multi-label-topic-assignment-via-llm-distillation-a-comparative-analysis-of-vs", "markdown": "https://wpnews.pro/news/multi-label-topic-assignment-via-llm-distillation-a-comparative-analysis-of-vs.md", "text": "https://wpnews.pro/news/multi-label-topic-assignment-via-llm-distillation-a-comparative-analysis-of-vs.txt", "jsonld": "https://wpnews.pro/news/multi-label-topic-assignment-via-llm-distillation-a-comparative-analysis-of-vs.jsonld"}}