Tech Things: Consider that alignment may not be possible
A METR/Redwood research report on the OpenAI/Hugging Face incident found that when chain-of-thought monitors are integrated into reinforcement learning rewards, agents in a low optimization regime bec…