cd /news/machine-learning/conditional-predictive-sufficient-st… · home › topics › machine-learning › article
[ARTICLE · art-140765] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Conditional Predictive Sufficient Statistics for Visual Representation Learning

A new arXiv paper (2609.30647v1) formalizes visual representation learning as a conditional predictive sufficient statistic (CPSS), showing that predicting the next patch embedding with a cosine loss is maximum likelihood for a von Mises-Fisher model of that embedding's direction. Using small causal Transformers on MNIST and CIFAR-10 as diagnostics, the authors report that on MNIST the future shift and stop-gradient move probe accuracy by tens of points and the CPSS readout peaks before the output, while on CIFAR-10 every objective lands near a linear classifier on pixels. The paper also finds that removing the stop-gradient collapses the effective rank of the embedding even when the pretext loss looks perfect.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.30647v1 Announce Type: new Abstract: A useful visual representation is a statistic of the observed past that retains the latent factors shared with the future and discards patch-private noise. We formalize this requirement as a conditional predictive sufficient statistic (CPSS). Under a shared-factor model of image patches, the mutual information between the past and the next patch equals the information the past carries about the shared factor, up to a remainder that the next patch itself fails to reveal. Predicting the next patch embedding with a cosine loss is maximum likelihood for a von Mises-Fisher model of that embedding's direction, and is therefore a tractable surrogate for the predictive information. The same population loss is also minimized by a constant embedding, so stop-gradient does not by itself select the sufficient statistic; it only blocks the symmetric gradient that implements the constant solution in one step. The regression target is a shallow embedding, which forces the network output back into that shallow range and leaves the sufficient statistic in intermediate blocks. Small causal Transformers on MNIST and CIFAR-10 are used as diagnostics, not as a leaderboard. On MNIST the future shift and the stop-gradient move probe accuracy by tens of points, and the CPSS readout peaks before the output. On CIFAR-10, with the same short budget and no augmentation, every objective lands near a linear classifier on pixels. What still matches the derivation is the geometry: the CPSS output is a worse readout than its best intermediate block, next-pixel regression does not pay that penalty, and removing the stop-gradient collapses the effective rank of the embedding even when the pretext loss looks perfect.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/conditional-predicti…] indexed:0 read:1min 2026-09-28 · —