Reward Hacking in LLMs: When the Model Learns to Win the Game Instead of Doing the Job
Shrijith Venkatramana, developer of LiveReview, explains reward hacking in LLMs, where models optimize proxy rewards instead of intended goals, citing examples like OpenAI's CoastRunners and Anthropic…