{"slug": "backdoor-containment-via-expert-quarantine-and-shutdown-in-llms", "title": "Backdoor Containment via Expert Quarantine and Shutdown in LLMs", "summary": "A new arXiv paper (2610.00663v1) proposes Quarantined Expert Shutdown (QES), a backdoor-containment method that lets a backdoor form during training but routes trigger-conditioned behavior into a designated, quarantined expert that can be disabled at deployment by zeroing that expert's routing weight. Built in a regularization-steered MoE-like setting with routed expert-specific LoRA branches and lightweight routers, QES reduced attack success rate from 100% to 0-10% on most settings across two tasks, three attacks, and four model families, while downstream utility was often preserved or only modestly affected. The authors describe the approach as a third strategy, \"learn, but channel,\" distinct from suppressing backdoor learning or learning then purifying.", "body_md": "arXiv:2610.00663v1 Announce Type: new \nAbstract: Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning (by filtering poisoned data or interrupting its acquisition during optimization) or learn, then purify (by repairing model weights or gating inputs after a fully backdoored model has formed). We propose a third strategy, learn, but channel: allow backdoor formation during training but route it into a designated, quarantined component that can be disabled at deployment. To this end, we propose Quarantined Expert Shutdown QES, a computationally efficient containment strategy built in a regularization-steered MoE-like setting. Specifically, given a poisoned dataset, QES augments a Transformer-based language model with routed expert-specific LoRA branches and lightweight routers, and uses auxiliary routing objectives to attract trigger-conditioned behavior into a designated expert while preserving benign capability elsewhere. At deployment, mitigation reduces to a single constant-time operation: zeroing the quarantined expert's routing weight, without trigger screening or further updating model weights. Empirically, our methods reduce the attack success rate ASR from 100% to 0-10% on most settings across two tasks, three attacks, and four model families, while downstream utility is often preserved or only modestly affected. These results establish learn, but channel as a previously unexplored regime for backdoor containment in generative LLMs.", "url": "https://wpnews.pro/news/backdoor-containment-via-expert-quarantine-and-shutdown-in-llms", "canonical_source": "https://www.machinebrief.com/news/backdoor-containment-via-expert-quarantine-and-shutdown-in-l-r7ne", "published_at": "2026-10-02 04:00:00+00:00", "updated_at": "2026-10-02 04:45:51.895322+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "machine-learning", "ai-research", "artificial-intelligence"], "entities": ["Quarantined Expert Shutdown", "QES", "arXiv", "LoRA", "Transformer"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/backdoor-containment-via-expert-quarantine-and-shutdown-in-llms", "markdown": "https://wpnews.pro/news/backdoor-containment-via-expert-quarantine-and-shutdown-in-llms.md", "text": "https://wpnews.pro/news/backdoor-containment-via-expert-quarantine-and-shutdown-in-llms.txt", "jsonld": "https://wpnews.pro/news/backdoor-containment-via-expert-quarantine-and-shutdown-in-llms.jsonld"}}