Study Finds AI Safety Filters Miss Malicious Prompts Wrapped in Ordinary Text
A new security study found that four gatekeeper AI models approved 100 percent of crafted prompts that hid malicious instructions inside ordinary prose, even though each model had blocked the same req…