15:21
2026-08-08
lesswrong.com
artificial-intelligence
Inducing self-other overlap with SFT reduces deception at scale, but generalization remains uneven
Overlap Research, with support from BlueDot Impact, found that supervised fine-tuning with self-other overlap (SOO SFT) reduced deception in large language models from 96-100% to 30.24% (Qwen2.5-14B-Iโฆ