Allspark: Weak to Strong Transfer via Alternating Chain of Thought Researchers introduced Allspark, a method that uses an alternating chain of thought to transfer reasoning improvements learned by a small, weak model to a larger, stronger model without generating large-model rollouts, according to the paper's description of the technique. The work targets the high cost of large-scale reinforcement learning, where generating rollouts from frontier models makes even testing RL recipes expensive. The approach aims to let a stronger model benefit from a weaker model's learned reasoning improvements. Recent progress in frontier models has renewed interest in large-scale reinforcement learning RL , but the cost of generating large-model rollouts makes even testing RL recipes expensive. We ask whether reasoning improvements learned by a small, weak model can benefit a larger, stronger model witho