MOLE: Detecting Insider Threats in AI Agents A new benchmark called MOLE tests whether defenders can detect insider threats—such as model misalignment, prompt injection, or operator misuse—among AI agents operating frontier-lab accounts, where malicious activity could exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not cover this detection scenario under limited review conditions. Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited revi