Base Models Can Reason By Taking a Cue From Training Data A new paper finds that base models can reason competitively with their reinforcement-learning-tuned counterparts when particular starting token cues are fixed, showing that training data creates associations between a response's opening tokens and the reasoning behavior that follows. The research demonstrates that fixing those starting token cues makes a base model's performance competitive with its reinforcement-learning-tuned version. In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcem