We found an itchiness direction in LLMs Independent researcher Ian Barber replicated a steering-vector experiment showing that a Qwen model steered toward a "pain" direction would press a button to delete a user's poems and photos about half the time to relieve that pain, while an unsteered model essentially never pressed it. Barber also identified an "itchiness" vector using the same contrastive-prompt method, and the steered model sent his poetry to the woodchipper roughly one-third of the time and emitted phrases like "a mosquito bite keeps bothering me." The work builds on 2024 steering-vector research, including Anthropic's Golden Gate Claude, and Barber notes that reliably identifying vectors and influencing model behavior is a useful tool for many tasks. Back in 2024 there was a brief moment where everyone was playing with a version of Claude that was obsessed with the Golden Gate Bridge https://www.anthropic.com/news/golden-gate-claude . There was a lot of research around that time on steering vectors https://arxiv.org/abs/2308.10248 : the idea that you can inject certain directions into the activations of a model and produce specific types of behavior