arXiv:2610.00663v1 Announce Type: new Abstract: Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning (by filtering poisoned data or interrupting its acquisition during optimization) or learn, then purify (by repairing model weights or gating inputs after a fully backdoored model has formed). We propose a third strategy, learn, but channel: allow backdoor formation during training but route it into a designated, quarantined component that can be disabled at deployment. To this end, we propose Quarantined Expert Shutdown QES, a computationally efficient containment strategy built in a regularization-steered MoE-like setting. Specifically, given a poisoned dataset, QES augments a Transformer-based language model with routed expert-specific LoRA branches and lightweight routers, and uses auxiliary routing objectives to attract trigger-conditioned behavior into a designated expert while preserving benign capability elsewhere. At deployment, mitigation reduces to a single constant-time operation: zeroing the quarantined expert's routing weight, without trigger screening or further updating model weights. Empirically, our methods reduce the attack success rate ASR from 100% to 0-10% on most settings across two tasks, three attacks, and four model families, while downstream utility is often preserved or only modestly affected. These results establish learn, but channel as a previously unexplored regime for backdoor containment in generative LLMs.
Backdoor Containment via Expert Quarantine and Shutdown in LLMs
A new arXiv paper (2610.00663v1) proposes Quarantined Expert Shutdown (QES), a backdoor-containment method that lets a backdoor form during training but routes trigger-conditioned behavior into a designated, quarantined expert that can be disabled at deployment by zeroing that expert's routing weight. Built in a regularization-steered MoE-like setting with routed expert-specific LoRA branches and lightweight routers, QES reduced attack success rate from 100% to 0-10% on most settings across two tasks, three attacks, and four model families, while downstream utility was often preserved or only modestly affected. The authors describe the approach as a third strategy, "learn, but channel," distinct from suppressing backdoor learning or learning then purifying.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.