17:46
2026-09-01
promptcube3.com
ai-safety
Visual inputs are bypassing LLM safety filters in ways that
Researchers analyzing 10 vision-language models (VLMs) found that visual inputs bypass safety filters because text-based refusal relies on a tiny cluster of about 88 neurons (less than 0.01% of total)…