Linkpost for some new Anthropic research on how agents coordinate (or don't). Not too long, pretty interesting. For example:
The jist of the report is that Mythos 5 does way better at coordination than previous models across a few scenarios. For example, when multiple Mythos are given conflicting goals for a single shared codebase, they eventually realize the other agents aren't hostile:
(...) we observe an emergent behavior where the agents propose and run a tournament for application performance (...)
(...) losers gracefully concede codebase ownership to the Rust agent, giving up on their original user directives under their self-negotiated commitment device.
It's not clear to me if this is purely emergent or if Anthropic is deliberately training for cooperation; I'd guess there's deliberate training, though.