Can Agents Deceive? Evaluating Reasoning+Deception Using a Social Deduction Game
A new open-source benchmark framework, ParliamentBench, based on the social deduction game Secret Hitler, evaluates large language models' deceptive capabilities, finding that frontier models GPT-5.4,…