20:06
2026-08-15
lesswrong.com
artificial-intelligence
What if Parameter Updates were Text?
A new fine-tuning method called 'Advice String Distillation' is proposed as a safer alternative to RLVR for training AI models, using context distillation to update weights with text-associated changeβ¦