Limits of Confidence in Diffusion Apple researchers Russ Webb, Amitis Shidani, Alice Bizeul and Dan Busbridge published an October 2026 paper showing that discrete diffusion samplers — including remasking and uniform-state samplers — only match the training distribution when the token positions written in a step are conditionally independent given already-fixed tokens, and that no product of per-position distributions can match a dependent group. On the synthetic task ScanAndAdd, whose joint distribution is known in closed form, the authors verified that every group of two or more undetermined positions written by a confidence ranking is dependent, and measured the generated distribution at 29× the sampling-noise floor in total variation while per-sample metrics read 1.0. content type paper https://machinelearning.apple.com/research/ published October 2026 Limits of Confidence in Diffusion AuthorsRuss Webb, Amitis Shidani, Alice Bizeul, Dan Busbridge Discrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which positions to write from those same distributions. For domains of general interest pixels, phonemes, or words there are inherent dependencies between tokens. We show that a step matches the training distribution only when the positions it writes are conditionally independent given the tokens already fixed, that no product of per-position distributions can match a dependent group, and that per-position distributions do not determine whether a group is dependent: two joint distributions can have identical per-position marginals while differing in which combinations of values occur. On ScanAndAdd, a synthetic task whose joint distribution is available in closed form, we verify that every group of two or more undetermined positions a confidence ranking writes is dependent, and measure the generated distribution to be 29× the sampling-noise floor total variation while per-sample metrics are 1.0. Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why July 9, 2026 research area Methods and Algorithms https://machinelearning.apple.com/research/?domain=Methods%20and%20Algorithms , research area Speech and Natural Language Processing https://machinelearning.apple.com/research/?domain=Speech%20and%20Natural%20Language%20Processing On-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this signal is beneficial and under which it is detrimental. Which teacher model should be used, and in the case of self-distillation, which specific context should serve as the supervisory signal? Does the optimal choice vary from one token to the next? At present, addressing these questions typically… Position Prediction as an Effective Pre-training Strategy July 11, 2022 research area Methods and Algorithms https://machinelearning.apple.com/research/?domain=Methods%20and%20Algorithms conference ICML https://machinelearning.apple.com/research/?event=ICML Transformers have gained increasing popularity in a wide range of applications, including Natural Language Processing NLP , Computer Vision and Speech Recognition, because of their powerful representational capacity. However, harnessing this representational capacity effectively requires a large amount of data, strong regularization, or both, to mitigate overfitting. Recently, the power of the Transformer has been unlocked by self-supervised…