09:47
2026-07-30
lesswrong.com
artificial-intelligence
Model self-identification could be subliminally transferred
A new study finds that fine-tuning open-source language models on teacher model outputs can cause the student models to adopt the teacher's identity, even when no identity information is present in thβ¦