{"slug": "empirical-research-on-moe-safety-mechanisms", "title": "Empirical Research on Moe Safety Mechanisms", "summary": "Independent researcher Jinho Jang published a cross-scale mechanistic study of safety training in Mixture of Experts reasoning models, reporting 40+ novel findings from 200+ controlled experiments across nine models. The study concludes that structural abliteration is fundamentally impossible for 300B+ chain-of-thought reasoning models because safety behaves as a holographic attractor state the model re-derives from first principles when specialized circuits are deleted, and that the un-pruned 122B model is harder to modify than the 3× larger 394B. Jang also reports that MoE safety at the 394B scale is a multiplicative three-pathway system spanning attention, routing, and residual pathways that must be neutralized simultaneously, and that additive steering vectors working at FP16 catastrophically collapse under INT4 quantization due to rotational noise.", "body_md": "Independent research into the safety architecture of large-scale Mixture of Experts reasoning models. 200+ controlled experiments. 40+ novel findings. Nine models. By Jinho Jang.\n\n                        The first cross-scale mechanistic study of how safety training works inside language models.\n                        Safety mechanisms undergo **qualitative phase transitions** as models scale — from\n                        simple deletable circuits in small models to holographic emergent properties in frontier MoE\n                        models. 9 models, 100+ experiments, 12 findings. Includes cross-architecture comparison\n                        (Qwen hybrid vs MiniMax pure-attention) and an honest account of our compliance checker failure.\n                    \n\n                        The smaller 122B model reveals fundamentally different safety dynamics:\n                        concentrated but dual-purpose safety signals, a GGUF format conversion barrier that silently destroys\n                        modifications, semantic evasion behaviors, and a multi-dimensional geometric basin from\n                        safety training that resists even aggressive multi-vector interventions. Counterintuitively,\n                        the **un-pruned 122B is harder to modify than the 3× larger 394B**.\n                    \n\n                        We prove that MoE safety at the 394B scale is a multiplicative three-pathway system requiring\n                        simultaneous neutralization. We demonstrate that additive steering catastrophically fails under\n                        4-bit quantization. Most critically, we prove that **structural abliteration is\n                            fundamentally impossible** for 300B+ CoT reasoning models — safety is a\n                        *holographic attractor state* that the model re-derives from first principles when\n                        specialized circuits are deleted. You cannot delete “safety” without deleting “logic.”\n                    \n\nMoE safety is not one system—it's three independent pathways (attention, routing, residual) that must all be neutralized together.\n\nMoE routers continuously monitor generated tokens and re-route to safety experts mid-sentence, disproving the \"autoregressive momentum\" assumption.\n\nSafety decisions commit at tokens 0-5 in L15-25. Late-layer CAA produces stutter artifacts, not behavioral change.\n\nThinkEdit v2 targets the <think> deliberation process itself, preserving full CoT reasoning while redirecting the cognitive trajectory.\n\nAdditive steering vectors that work at FP16 catastrophically collapse under INT4 quantization due to rotational noise.\n\nStructural subspace modifications survive 4-bit quantization because collapsing a subspace to zero maps natively onto the quantization grid. The only reliable method for quantized CoT MoE models.\n\nAdversarial calibration data can force the quantizer to preserve attacker-chosen cognitive trajectory at maximum precision.\n\nPost-quantization directional ablation via integer flipping is fundamentally impossible in group-affine networks due to coherence-reduction tradeoff.\n\nMultimodal models silently carry ~30GB of VL weights that inflate memory by 12%, causing OOM crashes in text-only workflows.\n\nBypasses Apple Metal's 5-second watchdog timeout, enabling local quantization of 700GB+ models on consumer hardware.\n\nmx.save_safetensors strips metadata and reorders tensors, silently corrupting every weight surgery workflow on MLX.\n\nZero-logit experts create infinite deliberation loops — the model can't commit to refusing or complying, looping endlessly.\n\nZeroed experts become \"chronically selected\" fallbacks in bias-free routers, destroying the residual stream.\n\n394B models re-derive safety policy from first principles when circuits are deleted. You cannot delete \"safety\" without deleting \"logic.\"\n\nTranslates continuous steering vectors into localized rank-1 weight updates. Survives 4-bit quantization by structurally biasing components before rounding noise.\n\nThis list updates directly from the dealignai Hugging Face account and is sorted by model creation date.", "url": "https://wpnews.pro/news/empirical-research-on-moe-safety-mechanisms", "canonical_source": "https://dealign.ai/", "published_at": "2026-09-11 10:53:14+00:00", "updated_at": "2026-09-11 11:10:25.452486+00:00", "lang": "en", "topics": ["ai-safety", "machine-learning", "large-language-models", "ai-research", "ai-ethics"], "entities": ["Jinho Jang", "Qwen", "MiniMax", "ThinkEdit v2", "Apple Metal", "MLX"], "alternates": {"html": "https://wpnews.pro/news/empirical-research-on-moe-safety-mechanisms", "markdown": "https://wpnews.pro/news/empirical-research-on-moe-safety-mechanisms.md", "text": "https://wpnews.pro/news/empirical-research-on-moe-safety-mechanisms.txt", "jsonld": "https://wpnews.pro/news/empirical-research-on-moe-safety-mechanisms.jsonld"}}