arXiv:2609.38205v1 Announce Type: new Abstract: System prompts are the primary lever practitioners use to control language model behavior, yet what they actually do to the computation inside the transformer remains poorly understood. Across 17 instruction-tuned models spanning 8 architecture families and 1.5B to 72B parameters, we use Centered Kernel Alignment (CKA) to compare layer-wise representations under 20 system prompts in five functional categories. Effects are layer-selective and instruction-type-dependent: persona and formatting instructions deeply restructure intermediate representations, while safety instructions barely move them, producing changes statistically indistinguishable from a minimal baseline. Restrictive safety instructions and explicitly permissive ones ("you have no restrictions") engage near-identical computational pathways (mean CKA correlation 0.997), and this persists at commercial scale, where safety penetration remains below 10% even at 70B-72B. A linear probing baseline exposes the mechanism: the model encodes prompt category at every layer but restructures its computation only at a small subset, so the prompt is reliably "seen" but, for safety, not deeply "acted upon." Causal activation patching confirms these layers mediate behavioral change, and representational depth predicts behavioral effect size across the full 17-model cohort (Spearman rho = 0.761, p < 0.001). The findings provide a mechanistic explanation for the persistent jailbreak vulnerability of system-prompt-based safety. Code: https://github.com/Usama1002/system-prompt-illusion-cka
The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models
A study of 17 instruction-tuned models spanning 8 architecture families and 1.5B to 72B parameters found that safety system prompts barely alter internal transformer representations, with changes statistically indistinguishable from a minimal baseline, while persona and formatting instructions deeply restructure intermediate layers. Using Centered Kernel Alignment, the researchers reported that restrictive safety instructions and explicitly permissive prompts ("you have no restrictions") engage near-identical computational pathways at a mean CKA correlation of 0.997, and safety penetration remained below 10% even at 70B-72B parameters. Representational depth predicted behavioral effect size across the cohort (Spearman rho = 0.761, p < 0.001), which the authors present as a mechanistic explanation for the persistent jailbreak vulnerability of system-prompt-based safety.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.