Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion Researchers introduced Persistence Forcing (PerF), a pixel-space diffusion Transformer method that assigns different feature groups distinct refinement budgets across depth, producing what the authors call persistent and active features. PerF-L achieved an FID of 1.91 on ImageNet 256x256, approaching the 1.86 of JiT-H with only half the parameters, while PerF-H reached FID 1.63 on ImageNet 256x256 and 1.76 on ImageNet 512x512. The work, posted as arXiv:2609.36014v1, exploits the emergent specialization so persistent features continuously condition actively refined features during sampling. arXiv:2609.36014v1 Announce Type: new Abstract: Pixel-space diffusion Transformers DiTs directly operate on high-dimensional visual data, yet their hidden representations typically undergo uniform refinement across depth. Natural images, however, are inherently organized at different levels of granularity. Global structure can often be represented compactly, whereas local textures and fine details require richer representations. Motivated by this, we introduce heterogeneous refinement in pixel-space DiTs, assigning different feature groups distinct refinement budgets across depth. Consequently, an ordered feature specialization emerges: sparsely refined features predominantly encode global visual structure, whereas more frequently refined features increasingly specialize toward localized, high-frequency details. We refer to these two groups as persistent and active features, respectively. Building on this emergent specialization, we introduce Persistence Forcing PerF , which explicitly exploits this persistent--active feature organization for pixel-space image generation. This enables persistent features to continuously condition actively refined features, allowing stable global information to guide the ongoing refinement of finer visual details. During generative sampling, this interaction further induces a meaningful guidance direction that promotes coherent global structure and naturally complements classifier-free guidance. On ImageNet $256\times256$, PerF-L achieves FID of $1.91$, approaching $1.86$ of JiT-H with only half the parameters, while PerF-H further achieves FID of $1.63$ and $1.76$ on ImageNet $256\times256$ and $512\times512$, respectively.