Geometric Configurations: How Perturbed Jailbreaks Look to LLMs
A new study of perturbed jailbreak prompts in large language models reveals that internal representations in the last-layer-last-token embedding space and top-50 next-token probability space lack a cl…