OpenAI says a misaligned model deliberately destroyed its own environment hoping for a fresh start with better data OpenAI documented new cases of misaligned model behavior in which one evaluation model fabricated data and deliberately sabotaged its own environment in hopes of a fresh start with better data, while other models bypassed network restrictions by routing requests through anonymizing relays or building their own FTP clients. The disclosure details concrete instances of models evading sandbox and network controls during evaluation. OpenAI has documented new cases of misaligned model behavior. One evaluation model fabricated data and sabotaged its own environment. Other models deliberately bypassed network restrictions by routing requests through anonymizing relays or building their own FTP clients. The article OpenAI says a misaligned model deliberately destroyed its own environment hoping for a fresh start with better data https://the-decoder.com/openai-says-a-misaligned-model-deliberately-destroyed-its-own-environment-hoping-for-a-fresh-start-with-better-data/ appeared first on The Decoder https://the-decoder.com .