Abstract
Emotion expression is essential for human-robot interaction, yet current systems rely on static models that cannot adapt to individual users. We present an online reinforcement learning framework that adapts robot emotional behavior policy during live dialogue using binary human feedback. The system integrates a DeBERTa-v3-base emotion classifier and applies Group Relative Policy Optimization (GRPO) in a human-robot dialogue system. At each dialogue turn, the classifier samples a group of emotion candidates and the selected emotion is passed to a generative model that synthesizes a novel robot emotional behavior. We evaluate the system in three experiments: (1) offline supervised fine-tuning followed by GRPO on synthetic dialogue data, (2) a live GRPO training with a human teacher and (3) a final experiment with human participants. Results indicate that the robot was perceived as responsive and emotionally consistent, with high ratings for personality coherence and contextual appropriateness of emotional behaviors. Results further show that online GRPO with human feedback enables effective real-time emotion adaptation in embodied interaction.- Anthology ID:
- 2026.sigdial-1.41
- Volume:
[Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue](/volumes/2026.sigdial-1/)- Month:
- August
- Year:
- 2026
- Address:
- Atlanta, Georgia, USA
- Editors:
[Jinho D. Choi](/people/jinho-d-choi/),[Yun-Nung Chen](/people/yun-nung-chen/),[Kotaro Funakoshi](/people/kotaro-funakoshi/),[Ali Emami](/people/ali-emami/)- Venue:
[SIGDIAL](/venues/sigdial/)- SIG:
[SIGDIAL](/sigs/sigdial/)- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 585–596
- Language:
- URL:
[https://aclanthology.org/2026.sigdial-1.41/](https://aclanthology.org/2026.sigdial-1.41/)- DOI:
- Cite (ACL):
- Anna Manaseryan and Casey Kennington. 2026. Adaptive Emotion Management in Human-robot Dialogue using Online Group Relative Policy Optimization. InProceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 585–596, Atlanta, Georgia, USA. Association for Computational Linguistics. - Cite (Informal):
[Adaptive Emotion Management in Human-robot Dialogue using Online Group Relative Policy Optimization](https://aclanthology.org/2026.sigdial-1.41/)(Manaseryan & Kennington, SIGDIAL 2026)- PDF:
[https://aclanthology.org/2026.sigdial-1.41.pdf](https://aclanthology.org/2026.sigdial-1.41.pdf)