17:45
2026-07-14
lesswrong.com
ai-safety
Can risk aversion learned at low stakes generalize to astronomically high stakes?
A new study from researchers at Forethought Foundation finds that training language models to be risk-averse on low-stakes gambles (prizes up to $100) can generalize to astronomically high stakes (pri…