{"slug": "conditional-predictive-sufficient-statistics-for-visual-representation-learning", "title": "Conditional Predictive Sufficient Statistics for Visual Representation Learning", "summary": "A new arXiv paper (2609.30647v1) formalizes visual representation learning as a conditional predictive sufficient statistic (CPSS), showing that predicting the next patch embedding with a cosine loss is maximum likelihood for a von Mises-Fisher model of that embedding's direction. Using small causal Transformers on MNIST and CIFAR-10 as diagnostics, the authors report that on MNIST the future shift and stop-gradient move probe accuracy by tens of points and the CPSS readout peaks before the output, while on CIFAR-10 every objective lands near a linear classifier on pixels. The paper also finds that removing the stop-gradient collapses the effective rank of the embedding even when the pretext loss looks perfect.", "body_md": "arXiv:2609.30647v1 Announce Type: new \nAbstract: A useful visual representation is a statistic of the observed past that retains the latent factors shared with the future and discards patch-private noise. We formalize this requirement as a conditional predictive sufficient statistic (CPSS). Under a shared-factor model of image patches, the mutual information between the past and the next patch equals the information the past carries about the shared factor, up to a remainder that the next patch itself fails to reveal. Predicting the next patch embedding with a cosine loss is maximum likelihood for a von Mises-Fisher model of that embedding's direction, and is therefore a tractable surrogate for the predictive information. The same population loss is also minimized by a constant embedding, so stop-gradient does not by itself select the sufficient statistic; it only blocks the symmetric gradient that implements the constant solution in one step. The regression target is a shallow embedding, which forces the network output back into that shallow range and leaves the sufficient statistic in intermediate blocks. Small causal Transformers on MNIST and CIFAR-10 are used as diagnostics, not as a leaderboard. On MNIST the future shift and the stop-gradient move probe accuracy by tens of points, and the CPSS readout peaks before the output. On CIFAR-10, with the same short budget and no augmentation, every objective lands near a linear classifier on pixels. What still matches the derivation is the geometry: the CPSS output is a worse readout than its best intermediate block, next-pixel regression does not pay that penalty, and removing the stop-gradient collapses the effective rank of the embedding even when the pretext loss looks perfect.", "url": "https://wpnews.pro/news/conditional-predictive-sufficient-statistics-for-visual-representation-learning", "canonical_source": "https://arxiv.org/abs/2609.30647", "published_at": "2026-09-28 04:00:00+00:00", "updated_at": "2026-09-28 04:20:42.449215+00:00", "lang": "en", "topics": ["machine-learning", "computer-vision", "ai-research", "neural-networks"], "entities": ["arXiv", "MNIST", "CIFAR-10", "von Mises-Fisher"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/conditional-predictive-sufficient-statistics-for-visual-representation-learning", "markdown": "https://wpnews.pro/news/conditional-predictive-sufficient-statistics-for-visual-representation-learning.md", "text": "https://wpnews.pro/news/conditional-predictive-sufficient-statistics-for-visual-representation-learning.txt", "jsonld": "https://wpnews.pro/news/conditional-predictive-sufficient-statistics-for-visual-representation-learning.jsonld"}}