{"slug": "taming-outlier-tokens-in-diffusion-transformers", "title": "Taming Outlier Tokens in Diffusion Transformers", "summary": "Researchers from Apple and academic institutions introduced Dual-Stage Registers (DSR), a register-based intervention that reduces outlier tokens in Diffusion Transformers (DiTs) for image generation, improving generation quality on ImageNet and large-scale text-to-image tasks. The study, led by Xiaoyu Wu and Yifei Wang, found that outlier tokens appear in both the encoder and denoiser of Representation Autoencoder (RAE)-DiT pipelines, and that masking high-norm tokens alone does not help, indicating corrupted local patch semantics. DSR uses trained registers, recursive test-time registers, and diffusion registers to consistently reduce artifacts and enhance performance.", "body_md": "[content type paper](/research/)published August 2026\n\nTaming Outlier Tokens in Diffusion Transformers\n\nAuthorsXiaoyu Wu†*, Yifei Wang†*, Tsu-Jui Fu, Liang-Chieh Chen, Zhe Gan, Chen Wei†\n\nTaming Outlier Tokens in Diffusion Transformers\n\nAuthorsXiaoyu Wu†*, Yifei Wang†*, Tsu-Jui Fu, Liang-Chieh Chen, Zhe Gan, Chen Wei†\n\nWe study outlier tokens in Diffusion Transformers (DiTs) for image generation. Prior work has shown that Vision Transformers (ViTs) can produce a small number of high-norm tokens that attract disproportionate attention while carrying limited local information, but their role in generative models remains underexplored. We show that this phenomenon appears in both the encoder and denoiser of modern Representation Autoencoder (RAE)-DiT pipelines: pretrained ViT encoders can produce outlier representations, and DiTs themselves can develop internal outlier tokens, especially in intermediate layers. Moreover, simply masking high-norm tokens does not improve performance, indicating that the problem is not only caused by a few extreme values, but is more closely related to corrupted local patch semantics. To address this issue, we introduce Dual-Stage Registers (DSR), a register-based intervention for both components: trained registers when available, recursive test-time registers otherwise, and diffusion registers for the denoiser. Across ImageNet and large-scale text-to-image generation, these interventions consistently reduce outlier artifacts and improve generation quality. Our results highlight outlier-token control as an important ingredient in building stronger DiTs.\n\nDiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation\n\nDecember 11, 2025[research area Computer Vision](/research/?domain=Computer%20Vision)\n\nIn this work, we empirically study Diffusion Transformers (DiTs) for text-to-image generation, focusing on architectural choices, text-conditioning strategies, and training protocols. We evaluate a range of DiT-based architectures—including PixArt-style and MMDiT variants—and compare them with a standard DiT variant which directly processes concatenated text and noise inputs. Surprisingly, our findings reveal that the performance of standard…\n\nOn Inductive Biases That Enable Generalization of Diffusion Transformers\n\nSeptember 22, 2025[research area Computer Vision](/research/?domain=Computer%20Vision)[conference NeurIPS](/research/?event=NeurIPS)\n\nRecent work studying the generalization of diffusion models with UNet-based denoisers reveals inductive biases that can be expressed via geometry-adaptive harmonic bases. However, in practice, more recent denoising networks are often based on transformers, e.g., the diffusion transformer (DiT). This raises the question: do transformer-based denoising networks exhibit inductive biases that can also be expressed via geometry-adaptive harmonic…", "url": "https://wpnews.pro/news/taming-outlier-tokens-in-diffusion-transformers", "canonical_source": "https://machinelearning.apple.com/research/taming-outlier-tokens", "published_at": "2026-08-05 00:00:00+00:00", "updated_at": "2026-08-09 13:17:20.616478+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "generative-ai"], "entities": ["Apple", "Xiaoyu Wu", "Yifei Wang", "Tsu-Jui Fu", "Liang-Chieh Chen", "Zhe Gan", "Chen Wei", "Diffusion Transformers (DiTs)"], "alternates": {"html": "https://wpnews.pro/news/taming-outlier-tokens-in-diffusion-transformers", "markdown": "https://wpnews.pro/news/taming-outlier-tokens-in-diffusion-transformers.md", "text": "https://wpnews.pro/news/taming-outlier-tokens-in-diffusion-transformers.txt", "jsonld": "https://wpnews.pro/news/taming-outlier-tokens-in-diffusion-transformers.jsonld"}}