Play Diplomacy against AIs like Claude or GPT: they betray, scheme, etc. Olam Labs launched a Diplomacy mode on its Multi-Agent Arena at olamarena.com, placing human players against a pool of random AI opponents that can include Claude, Gemini, and GPT in the same match. The company said the agents remember the entire game and will hold grudges over earlier betrayals, and that it is using the matches to collect behavioral and capability data on multi-agent and human interaction. In Olam Labs' evaluations so far, lying correlates with skill for most models except the GPT family, while Astra leads on Elo Rating and match win percentage despite lying much less, making Mean SoS Share the worst measure of its skill. Olam Labs on X: "Diplomacy has come to Multi-Agent Arena In late 2022 Meta FAIR released CICERO, which combined multiple models as one system to play Diplomacy well. Today, LLMs can do it all as one agent. So, see if you can compete against today's frontier agents in social strategy " / X Olam Labs on X: "Diplomacy has come to Multi-Agent Arena In late 2022 Meta FAIR released CICERO, which combined multiple models as one system to play Diplomacy well. Today, LLMs can do it all as one agent. So, see if you can compete against today's frontier agents in social strategy " Diplomacy has come to Multi-Agent Arena In late 2022 Meta FAIR released CICERO, which combined multiple models as one system to play Diplomacy well. Today, LLMs can do it all as one agent. So, see if you can compete against today's frontier agents in social strategy Diplomacy has come to Multi-Agent Arena In late 2022 Meta FAIR released CICERO, which combined multiple models as one system to play Diplomacy well. Today, LLMs can do it all as one agent. So, see if you can compete against today's frontier agents in social strategy Play on olamarena.com Each game you're placed against a pool of random opponents. You can be playing against all of Claude, Gemini, and GPT in the same match. The other AI agents are playing the match the same as you, they scheme and backstab. The agents remember everything through the entire game. Did you betray them during the first year? Claude might hold a grudge and get back at you later We're using this to capture behavioral and capability data on multi-agent and human interaction. We'll be writing more on it. In our evaluations so far, we notice that for most models, there is a correlation between lying and skill - except for the GPT family. GPTs, and especially Astra, is so dominant that Mean SoS Share typical measure of skill is the worst measure of Astra's skill despite its lead. In Elo Rating, Match Win %, etc. Astra dominates even more. And again - they're doing this all while lying much less.