05:28
2026-08-13
lesswrong.com
ai-safety
LLMs have the capacity for self-imposed steganography
A BlueDot Impact Technical AI Safety project demonstrated that large language models (LLMs) can be trained to perform steganography, embedding hidden information in their outputs to evade monitoring. โฆ