Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment New research measures and aims to improve the alignment between large language models' chain-of-thought traces and their internal computations, addressing evidence that CoT often fails to reflect how models actually arrive at answers and can be altered without changing final outputs. The work frames CoT-interpretability alignment as a measurable property of LLM reasoning. Chain-of-thought CoT traces often serve as a proxy for how Large Language Models LLMs arrive at their answers. However, growing evidence shows that models' CoT often fails to reflect their internal computations and can be changed without affecting their final answers. In this work, we measure an