{"slug": "e-moe-enhanced-mixture-of-experts-for-non-factorized-diffusion-language-models", "title": "E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models", "summary": "E-MoE, a new method from researchers behind arXiv paper 2609.37533v1, builds the reverse process of masked diffusion models as a mixture of factorized distributions over a discrete shared latent set by the expert-routing decisions of a Mixture-of-Experts backbone, without increasing active parameters over the factorized baseline. Across synthetic multi-modal benchmarks, binarized MNIST, and LM1B, E-MoE improves few-step generation over factorized baselines, addressing the posterior collapse that plagues prior continuous Gaussian latent approaches.", "body_md": "arXiv:2609.37533v1 Announce Type: new \nAbstract: Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most. A recent line of work introduces a continuous Gaussian latent, trained as a variational autoencoder, to capture correlations across positions, but such approaches are prone to posterior collapse, where the latent is silently ignored. We propose Enhanced Mixture-of-Experts (E-MoE), which builds the reverse process as a mixture of factorized distributions over a discrete shared latent given by the expert-routing decisions of a Mixture-of-Experts (MoE) backbone, without increasing active parameters over the factorized baseline. Across synthetic multi-modal benchmarks, binarized MNIST, and LM1B, E-MoE improves few-step generation over factorized baselines.", "url": "https://wpnews.pro/news/e-moe-enhanced-mixture-of-experts-for-non-factorized-diffusion-language-models", "canonical_source": "https://www.machinebrief.com/news/e-moe-enhanced-mixture-of-experts-for-non-factorized-diffusi-j1un", "published_at": "2026-09-30 04:00:00+00:00", "updated_at": "2026-09-30 04:47:26.719248+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "generative-ai", "ai-research", "large-language-models"], "entities": ["E-MoE", "Mixture-of-Experts", "masked diffusion models", "arXiv", "LM1B", "MNIST"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/e-moe-enhanced-mixture-of-experts-for-non-factorized-diffusion-language-models", "markdown": "https://wpnews.pro/news/e-moe-enhanced-mixture-of-experts-for-non-factorized-diffusion-language-models.md", "text": "https://wpnews.pro/news/e-moe-enhanced-mixture-of-experts-for-non-factorized-diffusion-language-models.txt", "jsonld": "https://wpnews.pro/news/e-moe-enhanced-mixture-of-experts-for-non-factorized-diffusion-language-models.jsonld"}}