Thought without systematicity? Evaluating reasoning models on rule induction tasks A new evaluation tests whether reasoning models exhibit systematicity by measuring their performance on structurally equivalent variants of rule induction tasks, according to the research. The study asks whether reasoning models robustly generalize understanding of one concept to close variations of that concept, a central tenet of human cognition. The work evaluates reasoning models specifically on rule induction tasks to assess that consistency. A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance on structurally equivalent va