Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions
Large language models (LLMs) show reduced accuracy on math problems when simple details like names or numbers are changed, and a new study finds that using code execution methods does not improve this robustness. Researc…