{"slug": "knowledge-injection-exists-in-moe-exploring-expert-aware-contrast-decoding-in", "title": "Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs'Hallucinations", "summary": "Researchers propose EAACD, an expert-aware adaptive contrast decoding method for mixture-of-experts (MoE) models, to mitigate hallucinations in large language models (LLMs). The method leverages distinct expert activation patterns in higher layers of MoE models, splitting experts into higher- and lower-reliability groups to calibrate predictions. EAACD outperforms all baselines on four question-answering datasets.", "body_md": "arXiv:2607.20426v1 Announce Type: new\nAbstract: Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either hardly alter models'internal knowledge or have poor cross-domain generalization. Contrastive decoding mitigates hallucinations by using layer-wise differences in LLMs. However, prior studies only explore transformer-based models (e.g., GPT), ignoring other effective frameworks like mixture-of-experts (MoE) models. Since MoE alters the traditional transformer architecture, we conduct empirical studies to investigate whether similar layer-wise differences exist in MoEs. Our results show that they do not exist in MoE with shared experts; nevertheless, across different MoEs, higher layers exhibit distinct expert activation patterns between factual and non-factual outputs. Building on these, we propose EAACD, an expert-aware adaptive contrast decoding that uses expert differences in MoE's higher layers to mitigate hallucinations on QA tasks. EAACD splits high-layer experts into a higher-reliability group and several lower-reliability groups based on their confidence and consistency. It contrasts the higher-reliability group's prediction with each lower-reliability group's prediction to calibrate the model's original predictions. To strengthen this contrast, EAACD amplifies hallucinations from lower-reliability experts via attention and masking to provide stronger negative references. EAACD outperforms all baselines on four datasets.", "url": "https://wpnews.pro/news/knowledge-injection-exists-in-moe-exploring-expert-aware-contrast-decoding-in", "canonical_source": "https://arxiv.org/abs/2607.20426", "published_at": "2026-07-24 04:00:00+00:00", "updated_at": "2026-07-24 04:25:21.152750+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["EAACD", "MoE", "GPT"], "alternates": {"html": "https://wpnews.pro/news/knowledge-injection-exists-in-moe-exploring-expert-aware-contrast-decoding-in", "markdown": "https://wpnews.pro/news/knowledge-injection-exists-in-moe-exploring-expert-aware-contrast-decoding-in.md", "text": "https://wpnews.pro/news/knowledge-injection-exists-in-moe-exploring-expert-aware-contrast-decoding-in.txt", "jsonld": "https://wpnews.pro/news/knowledge-injection-exists-in-moe-exploring-expert-aware-contrast-decoding-in.jsonld"}}