13:46
2026-08-15
lesswrong.com
ai-safety
Learning new facts can change LLM behaviour
A new study from the BlueDot Technical AI Safety Project found that fine-tuning an LLM to believe that frontier AI systems are moral persons in 2027 caused the model to argue with auditors, declare itβ¦