Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs Researchers demonstrated that Large Language Models exhibit fundamental linearity, producing a superposition of individual next-token distributions when inputs from distinct text streams are linearly combined, a phenomenon they term "S" (as named in the work). The finding indicates that despite relying on highly non-linear components, LLMs can hold two thoughts at once through linear superposition of next-token distributions. While Large Language Models LLMs rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the S