04:00
2026-10-01
arxiv.org
ai-safety
The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models
A study of 17 instruction-tuned models spanning 8 architecture families and 1.5B to 72B parameters found that safety system prompts barely alter internal transformer representations, with changes statβ¦