The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.
source & further reading
thezvi.substack.com — original article
OpenAI Shares Some Alignment Problems
WSJ Article Claiming China Has Matched Anthropic Is Obvious Nonsense
Claude Fable 5 and Mythos 5: Capabilities