Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without modifying these layers, overlooking the spectral properties of linea
A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning