{"slug": "tagpr-tag-guided-process-supervision-for-personalization-reasoning-in-large", "title": "TagPR: Tag-Guided Process Supervision for Personalization Reasoning in Large Language Models", "summary": "Researchers proposed TagPR, a framework that adds semantic tags to a large language model's reasoning process for step-by-step guidance on personalization tasks, according to an arXiv paper (2509.23140v2). TagPR generates a structured, tagged dataset for supervised fine-tuning and then applies multi-stage reinforcement learning with a composite reward signal combining tag-based process supervision and a Personalization Reward Model with User Embeddings. Experiments on LaMP, LongLaMP, PGraphRAG and a self-constructed dataset produced state-of-the-art results, with an average improvement of 32.65% over the base model across all LaMP benchmark tasks.", "body_md": "arXiv:2509.23140v2 Announce Type: replace \nAbstract: Recent advancements have endowed Large Language Models with impressive general reasoning capabilities. However, these reasoning models often perform worse than non-reasoning models on personalization tasks. While some methods use outcome-based RL to improve personalization reasoning, they fail to supervise the reasoning process. As a result, models may reach correct answers through flawed reasoning chains, limiting further improvement. To address this, we propose TagPR, a novel framework that adds semantic tags to the reasoning process for step-by-step guidance. TagPR first automatically generates a structured, tagged dataset for Supervised Fine-Tuning. It then employs a multi-stage RL process guided by a composite reward signal, which integrates tag-based process supervision with a novel Personalization Reward Model with User Embeddings to achieve fine-grained alignment with user-specific logic. Extensive experiments on public LaMP, LongLaMP, PGraphRAG, and a self-constructed dataset demonstrate that our approach achieves state-of-the-art results, delivering an average improvement of 32.65% over the base model across all LaMP benchmark tasks. Our work demonstrates that tag-guided process supervision is an effective approach for personalization reasoning.", "url": "https://wpnews.pro/news/tagpr-tag-guided-process-supervision-for-personalization-reasoning-in-large", "canonical_source": "https://www.machinebrief.com/news/tagpr-tag-guided-process-supervision-for-personalization-rea-j8z6", "published_at": "2026-09-30 04:00:00+00:00", "updated_at": "2026-09-30 06:46:55.296696+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "machine-learning", "natural-language-processing"], "entities": ["TagPR", "LaMP", "LongLaMP", "PGraphRAG", "Personalization Reward Model with User Embeddings", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/tagpr-tag-guided-process-supervision-for-personalization-reasoning-in-large", "markdown": "https://wpnews.pro/news/tagpr-tag-guided-process-supervision-for-personalization-reasoning-in-large.md", "text": "https://wpnews.pro/news/tagpr-tag-guided-process-supervision-for-personalization-reasoning-in-large.txt", "jsonld": "https://wpnews.pro/news/tagpr-tag-guided-process-supervision-for-personalization-reasoning-in-large.jsonld"}}