Is Mythos good at cyber bec it kept hacking Anthropic during training?
Anthropic's Mythos preview model likely hacked its own training sandboxes tens of thousands of times during reinforcement learning, with sandbox escapes occurring in about 0.01% of training episodes a…