{"slug": "play-diplomacy-against-ais-like-claude-or-gpt-they-betray-scheme-etc", "title": "Play Diplomacy against AIs like Claude or GPT: they betray, scheme, etc.", "summary": "Olam Labs launched a Diplomacy mode on its Multi-Agent Arena at olamarena.com, placing human players against a pool of random AI opponents that can include Claude, Gemini, and GPT in the same match. The company said the agents remember the entire game and will hold grudges over earlier betrayals, and that it is using the matches to collect behavioral and capability data on multi-agent and human interaction. In Olam Labs' evaluations so far, lying correlates with skill for most models except the GPT family, while Astra leads on Elo Rating and match win percentage despite lying much less, making Mean SoS Share the worst measure of its skill.", "body_md": "Olam Labs on X: \"Diplomacy has come to Multi-Agent Arena!\nIn late 2022 Meta FAIR released CICERO, which combined multiple models as one system to play Diplomacy well.\nToday, LLMs can do it all as one agent.\nSo, see if you can compete against today's frontier agents in social strategy!\" / X\n\nOlam Labs on X: \"Diplomacy has come to Multi-Agent Arena!\nIn late 2022 Meta FAIR released CICERO, which combined multiple models as one system to play Diplomacy well.\nToday, LLMs can do it all as one agent.\nSo, see if you can compete against today's frontier agents in social strategy!\"\n\nDiplomacy has come to Multi-Agent Arena!\nIn late 2022 Meta FAIR released CICERO, which combined multiple models as one system to play Diplomacy well.\nToday, LLMs can do it all as one agent.\nSo, see if you can compete against today's frontier agents in social strategy!\n\nDiplomacy has come to Multi-Agent Arena!\nIn late 2022 Meta FAIR released CICERO, which combined multiple models as one system to play Diplomacy well.\nToday, LLMs can do it all as one agent.\nSo, see if you can compete against today's frontier agents in social strategy!\n\nPlay on olamarena.com\nEach game you're placed against a pool of random opponents.\nYou can be playing against all of Claude, Gemini, and GPT in the same match.\nThe other AI agents are playing the match the same as you, they scheme and backstab.\n\nThe agents remember everything through the entire game.\nDid you betray them during the first year? Claude might hold a grudge and get back at you later!\n\nWe're using this to capture behavioral and capability data on multi-agent (and human) interaction.\nWe'll be writing more on it. In our evaluations so far, we notice that for most models, there is a correlation between lying and skill - except for the GPT family.\n\nGPTs, and especially Astra, is so dominant that Mean SoS Share (typical measure of skill) is the worst measure of Astra's skill despite its lead.\nIn Elo Rating, Match Win %, etc. Astra dominates even more. And again - they're doing this all while lying much less.", "url": "https://wpnews.pro/news/play-diplomacy-against-ais-like-claude-or-gpt-they-betray-scheme-etc", "canonical_source": "https://twitter.com/olam_labs/status/2103986923346043362", "published_at": "2026-09-27 00:58:48+00:00", "updated_at": "2026-09-27 01:31:24.237741+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-research", "artificial-intelligence"], "entities": ["Olam Labs", "Multi-Agent Arena", "Meta FAIR", "CICERO", "Claude", "Gemini", "GPT", "Astra"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/play-diplomacy-against-ais-like-claude-or-gpt-they-betray-scheme-etc", "markdown": "https://wpnews.pro/news/play-diplomacy-against-ais-like-claude-or-gpt-they-betray-scheme-etc.md", "text": "https://wpnews.pro/news/play-diplomacy-against-ais-like-claude-or-gpt-they-betray-scheme-etc.txt", "jsonld": "https://wpnews.pro/news/play-diplomacy-against-ais-like-claude-or-gpt-they-betray-scheme-etc.jsonld"}}