18:21
2026-10-02
arxiv.org
machine-learning
SFT matches RL if you MCMC the training data first
A paper submitted to arXiv on 1 October 2026 introduces a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces into more on-policy data for finetuning, alβ¦