22:06
2026-08-09
lesswrong.com
artificial-intelligence
Overthinking: Amplifying reasoning weights makes models reveal their secrets
Amplifying the weight difference between a reasoning model and its non-reasoning counterpart—creating an 'overthinking model'—surfaces hidden secrets up to 10× more often than the original reasoning m…