Why Are Some LLMs Harder to Jailbreak Than Others?
A model's resistance to jailbreaking depends primarily on the rigor of its alignment training, including the volume of safety preference pairs and techniques like Constitutional AI, according to an an…