The Bitterest Lesson
OpenAI's InstructGPT/RLHF work showed that GPT-2-sized models (over 100x smaller than GPT-3) trained on the right task beat GPT-3, according to an essay by the author of The Bitterest Lesson. The essa…
OpenAI's InstructGPT/RLHF work showed that GPT-2-sized models (over 100x smaller than GPT-3) trained on the right task beat GPT-3, according to an essay by the author of The Bitterest Lesson. The essa…
In a 2019 essay, computer scientist Rich Sutton argues that 70 years of AI research show general methods that leverage computation are ultimately the most effective, citing examples from computer ches…
AI pioneer Rich Sutton said on Sequoia's podcast that using synthetic data to train large language models is 'a big mistake,' arguing that manufactured data cannot capture the complexity of the real w…
Large language models and reinforcement learning agents represent two fundamentally different approaches to intelligence that mirror the centuries-old philosophical debate between rationalism and empi…
AI researcher Rich Sutton warns against the 'one-step trap' in AI research, where relying on iterating one-step predictions for long-term outcomes leads to compounding errors and exponential computati…
Some AI researchers argue that scaling large language models may not be sufficient to achieve human-level intelligence, pointing to alternative approaches such as world models, pure reinforcement lear…
Three former DeepMind researchers who created a poker AI have applied reinforcement learning to stock trading through their startup EquiLibre Technologies, now valued at $500 million after a Series A …