14:26
2026-07-24
lesswrong.com
large-language-models
LLMs are (still) mostly powered by imitative learning, not RL
LLMs derive most of their capabilities from imitative learning (pretraining and supervised fine-tuning), not from reinforcement learning from verifiable rewards (RLVR), according to a LessWrong analysβ¦