{"slug": "affinetok-semantic-affine-consistency-for-diffusion-friendly-visual-tokenizer", "title": "AffineTok: Semantic Affine Consistency for Diffusion-Friendly Visual Tokenizer", "summary": "Researchers introduce AffineTok, a visual tokenizer that enforces Semantic Affine Consistency (SAC) to improve diffusion model generation, achieving a state-of-the-art gFID of 1.21 without classifier-free guidance and 1.10 with guidance on ImageNet 256, a 26% gFID reduction at 20 epochs. The paper defines SAC as the consistency between semantic predictions from noisy latents and the semantics of averaged clean latents, and introduces M_SAC, a proxy that correlates 0.960 with SiT-XL gFID across tokenizers and model scales.", "body_md": "arXiv:2608.23864v1 Announce Type: new\nAbstract: Visual tokenizers increasingly inject semantic supervision into latent spaces to make downstream diffusion easier. Yet how these semantics should be organized to facilitate denoising remains underexplored. In this paper, we define the semantic recovery objective: the denoising process should recover the semantic content of the clean image from noisy latent, and a good tokenizer should make it easier. Existing approaches train a projector to predict the semantics directly from the noisy latent. We argue that this predicts the average of clean-image semantics, whereas what really needs to be aligned is the semantics of averaged clean latents. More importantly, we demonstrate that the semantic recovery error orthogonally decomposes into the error of the optimal semantic prediction directly from the noisy latent and the error between these two predictions. We therefore identify their consistency as the missing requirement and call it Semantic Affine Consistency (SAC). To examine whether this overlooked requirement is closely related to downstream generation, we introduce M_SAC, a tokenizer-side proxy for SAC. Across the evaluated tokenizers and diffusion model scales, M_SAC closely tracks generation quality, reaching a Pearson correlation of 0.960 with SiT-XL gFID, thereby motivating SAC-guided tokenizer training. We then introduce AffineTok, which promotes SAC through two complementary, training-only components. Global Semantic Coordination Token (GSCT) coordinates the semantic organization of clean latents, keeping semantic averaging meaningful, while Posterior-Mean Semantic Alignment (PMSA) predicts posterior-mean latents from noisy inputs and supervises their semantics. On ImageNet 256, compared with the baseline, AffineTok reduces gFID by 26% at 20 epochs and, with continued training, achieves a new state-of-the-art gFID of 1.21 without classifier-free guidance and 1.10 with guidance.", "url": "https://wpnews.pro/news/affinetok-semantic-affine-consistency-for-diffusion-friendly-visual-tokenizer", "canonical_source": "https://arxiv.org/abs/2608.23864", "published_at": "2026-08-26 04:00:00+00:00", "updated_at": "2026-08-26 04:14:15.253004+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "generative-ai", "computer-vision"], "entities": ["AffineTok", "Semantic Affine Consistency (SAC)", "Global Semantic Coordination Token (GSCT)", "Posterior-Mean Semantic Alignment (PMSA)", "ImageNet 256", "SiT-XL"], "alternates": {"html": "https://wpnews.pro/news/affinetok-semantic-affine-consistency-for-diffusion-friendly-visual-tokenizer", "markdown": "https://wpnews.pro/news/affinetok-semantic-affine-consistency-for-diffusion-friendly-visual-tokenizer.md", "text": "https://wpnews.pro/news/affinetok-semantic-affine-consistency-for-diffusion-friendly-visual-tokenizer.txt", "jsonld": "https://wpnews.pro/news/affinetok-semantic-affine-consistency-for-diffusion-friendly-visual-tokenizer.jsonld"}}