Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited revi
Meta’s Muse Launches as the Biggest Cross-App AI Agent Yet—With a Security Disclosure Problem