How LLMs Defend Against Jailbreak and Prompt Injection
Prompt injection manipulates an LLM's instructions to ignore original constraints, while jailbreaking bypasses safety filters to produce restricted content. Reinforcement Learning from Human Feedback …