LLM Refusal Behavior on Open-Weight Model
Researchers from ZENTARA Labs detail how refusal behavior in open-weight large language models can be removed cheaply and surgically, exposing latent capabilities inherited from pretraining. The analysis outlines three r…