cd/entity/Beren MillidgeΒ· homeβ€Ί entitiesβ€Ί Beren Millidge
grep -l @beren millidge /news/*.json | wc -l β†’ 1

Beren Millidge

mentions 1 type Person feed RSS

// recent coverage 1 mentions

14:26
2026-07-24
lesswrong.com
large-language-models

LLMs are (still) mostly powered by imitative learning, not RL

LLMs derive most of their capabilities from imitative learning (pretraining and supervised fine-tuning), not from reinforcement learning from verifiable rewards (RLVR), according to a LessWrong analys…

// co-occurs with top 1 entities
// topics top 3 topics