Adaptive Emotion Management in Human-robot Dialogue using Online Group Relative Policy Optimization Researchers Anna Manaseryan and Casey Kennington presented an online reinforcement learning framework at the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue that adapts robot emotional behavior during live dialogue using binary human feedback. The system integrates a DeBERTa-v3-base emotion classifier and applies Group Relative Policy Optimization (GRPO) to select emotion candidates, with experiments showing the robot was perceived as responsive and emotionally consistent. Results indicate that online GRPO with human feedback enables effective real-time emotion adaptation in embodied interaction. Abstract Emotion expression is essential for human-robot interaction, yet current systems rely on static models that cannot adapt to individual users. We present an online reinforcement learning framework that adapts robot emotional behavior policy during live dialogue using binary human feedback. The system integrates a DeBERTa-v3-base emotion classifier and applies Group Relative Policy Optimization GRPO in a human-robot dialogue system. At each dialogue turn, the classifier samples a group of emotion candidates and the selected emotion is passed to a generative model that synthesizes a novel robot emotional behavior. We evaluate the system in three experiments: 1 offline supervised fine-tuning followed by GRPO on synthetic dialogue data, 2 a live GRPO training with a human teacher and 3 a final experiment with human participants. Results indicate that the robot was perceived as responsive and emotionally consistent, with high ratings for personality coherence and contextual appropriateness of emotional behaviors. Results further show that online GRPO with human feedback enables effective real-time emotion adaptation in embodied interaction.- Anthology ID: - 2026.sigdial-1.41 - Volume: Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue /volumes/2026.sigdial-1/ - Month: - August - Year: - 2026 - Address: - Atlanta, Georgia, USA - Editors: Jinho D. Choi /people/jinho-d-choi/ , Yun-Nung Chen /people/yun-nung-chen/ , Kotaro Funakoshi /people/kotaro-funakoshi/ , Ali Emami /people/ali-emami/ - Venue: SIGDIAL /venues/sigdial/ - SIG: SIGDIAL /sigs/sigdial/ - Publisher: - Association for Computational Linguistics - Note: - Pages: - 585–596 - Language: - URL: https://aclanthology.org/2026.sigdial-1.41/ https://aclanthology.org/2026.sigdial-1.41/ - DOI: - Cite ACL : - Anna Manaseryan and Casey Kennington. 2026. Adaptive Emotion Management in Human-robot Dialogue using Online Group Relative Policy Optimization https://aclanthology.org/2026.sigdial-1.41/ . In Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue , pages 585–596, Atlanta, Georgia, USA. Association for Computational Linguistics. - Cite Informal : Adaptive Emotion Management in Human-robot Dialogue using Online Group Relative Policy Optimization https://aclanthology.org/2026.sigdial-1.41/ Manaseryan & Kennington, SIGDIAL 2026 - PDF: https://aclanthology.org/2026.sigdial-1.41.pdf https://aclanthology.org/2026.sigdial-1.41.pdf