Revealing Backdoors in LLMs: New Detection Framework Emerges
Researchers have developed a new framework for detecting backdoor attacks in large language models, addressing the challenge of discrete input spaces. The framework introduces Class Subspace Orthogonalization (CSO) to en…