cd/entity/BlueDot Technical AI Safety ProjectΒ· homeβ€Ί entitiesβ€Ί BlueDot Technical AI Safety Project
grep -l @bluedot technical ai safety project /news/*.json | wc -l β†’ 2

BlueDot Technical AI Safety Project

mentions 2 type Organization feed RSS

// recent coverage 2 mentions

13:46
2026-08-15
lesswrong.com
ai-safety

Learning new facts can change LLM behaviour

A new study from the BlueDot Technical AI Safety Project found that fine-tuning an LLM to believe that frontier AI systems are moral persons in 2027 caused the model to argue with auditors, declare it…

05:58
2026-05-30
lesswrong.com
ai-safety

Belief manifolds, and how to steer along them

A BlueDot Technical AI Safety Project researcher reproduced a study from Goodfire demonstrating that language model representations form curved geometric manifolds, not simple linear directions. The w…

// co-occurs with top 6 entities
// topics top 6 topics