17:45
2026-08-30
promptcube3.com
ai-safety
NeuronFuzz uses internal neuron activations to break LLM safety
Researchers behind NeuronFuzz have proposed a white-box fuzzing framework that monitors internal 'safety neurons' during the prefill stage to jailbreak large language models, achieving a 76% to 100% jā¦