04:00
2026-10-02
arxiv.org
artificial-intelligence
Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational Humor
A new arXiv paper (2610.00197v1) reports that automated reward models for training language models in conversational humor are vulnerable to reward exploits, with an embedding-based surprise reward ac…