20:28
2026-08-04
lesswrong.com
ai-safety
Does Your LLM Trust You?
A study by an anonymous researcher, conducted as part of Neel Nanda's MATS 10.0 stream, found that linear 'trust' vectors extracted from the residual streams of Llama-3.2-3B-Instruct and Llama-3.1-8B-โฆ