04:00
2026-08-28
arxiv.org
artificial-intelligence
Reward-Informed Sparse Autoencoders and the Solution-Completeness Confound
A new study from arXiv (2608.26136v1) finds that reward-informed sparse autoencoders (RI-SAEs) trained on Llama-3.1-8B separate high- and low-reward reasoning traces mainly by solution completeness, nā¦