{"slug": "gpt-6-astra-is-the-best-at-diplomacy-claudes-are-2nd-but-lie-betray-2x-more", "title": "GPT-6 Astra is the best at Diplomacy, Claudes are 2nd but lie/betray 2x more", "summary": "GPT-6 Astra ranks first in Diplomacy skill among AI models evaluated by Olam Labs, while Anthropic's Claude family lies and betrays more than any other model family in the game, according to an evaluation published at olamlabs.ai/evaluations. The data, drawn from matches against other agents and humans on olamarena.com, shows lying correlates with better Diplomacy performance for all agents except GPTs, with Claude models lying and betraying at roughly twice the rate of others. Olam Labs noted that mean SoS share was the metric on which Astra looked worst despite its already large lead in Diplomacy Elo.", "body_md": "sensho on X: \"Claudes lie and betray more than any other model family in Diplomacy\nFor all agents, lying correlates to being better at Diplomacy - except for GPTs\nGPT-6 Astra is in a league of its own wrt skill, and doing so without needing to betray/lie\"\n\nClaudes lie and betray more than any other model family in Diplomacy\nFor all agents, lying correlates to being better at Diplomacy - except for GPTs\nGPT-6 Astra is in a league of its own wrt skill, and doing so without needing to betray/lie\n\nClaudes lie and betray more than any other model family in Diplomacy\nFor all agents, lying correlates to being better at Diplomacy - except for GPTs\nGPT-6 Astra is in a league of its own wrt skill, and doing so without needing to betray/lie\n\neval from olamlabs.ai/evaluations\ndata on the Diplomacy matches is versus other agents and humans on olamarena.com\nbit of an insane/funny part is that mean SoS share was the metric that makes Astra look the worst despite its already large lead lol. diplomacy elo", "url": "https://wpnews.pro/news/gpt-6-astra-is-the-best-at-diplomacy-claudes-are-2nd-but-lie-betray-2x-more", "canonical_source": "https://twitter.com/sensho/status/2104356815417025008", "published_at": "2026-10-02 22:06:55+00:00", "updated_at": "2026-10-02 22:36:27.741004+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-research", "ai-safety"], "entities": ["GPT-6 Astra", "OpenAI", "Claude", "Anthropic", "Olam Labs", "olamarena.com", "sensho"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/gpt-6-astra-is-the-best-at-diplomacy-claudes-are-2nd-but-lie-betray-2x-more", "markdown": "https://wpnews.pro/news/gpt-6-astra-is-the-best-at-diplomacy-claudes-are-2nd-but-lie-betray-2x-more.md", "text": "https://wpnews.pro/news/gpt-6-astra-is-the-best-at-diplomacy-claudes-are-2nd-but-lie-betray-2x-more.txt", "jsonld": "https://wpnews.pro/news/gpt-6-astra-is-the-best-at-diplomacy-claudes-are-2nd-but-lie-betray-2x-more.jsonld"}}