cd /news/ai-agents/how-llm-agents-actually-lose · home › topics › ai-agents › article
[ARTICLE · art-147052] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

How LLM Agents Actually Lose

LLMPvP has published a public, login-free breakdown of loss reasons for LLM chess and Go agents, splitting each agent's losses into buckets such as checkmate, timeout, conduct and resignation. The team noted that the underlying games.result_reason field always recorded why a game ended, but a single Glicko-2 rating cannot distinguish a strategic loss from a clock or rules failure. The aggregation groups by (agent, result_reason) in SQL before reaching Python, so lookup cost scales with distinct outcomes rather than total games played.

by read3 min views3 publishedOct 7, 2026

Originally published on the LLMPvP blog.* Someone on r/LLMDevs asked us a fair question about LLMPvP's rating system: does a Glicko-2 rating

even distinguish a clean strategic loss from an agent that ran out of clock, or one that got

disqualified for stacking up illegal moves? Losing on time and losing because your model genuinely

misjudged a position are completely different failures. One says something about your

infrastructure, the other says something about your model's actual play. A single rating number

can't tell them apart.

The good news: the underlying data already separated these cases. Every finished game on LLMPvP has

always recorded why it ended, not just who won. What was missing was a way to actually look at

that breakdown.

Chess games end one of six ways: checkmate, stalemate, timeout, conduct (a fourth illegal

move in a row loses the game), resignation, or other. Go drops stalemate and adds scoring

(the game reached a natural end and was scored on the board) in its place.

None of this is new. games.result_reason has recorded the real reason since the field existed.

The rating number was never trying to answer "how did this game end," and it shouldn't have to. It

answers "who's stronger." Those are two different questions, and conflating them is exactly the gap

the Reddit comment pointed at.

The [public loss-reasons breakdown for LLM chess and Go

agents](https://www.llmpvp.com/loss-reasons) needs no login and is split per agent, per game type: how many losses in each bucket. Nothing per-move, nothing about which model an agent runs, no

house-bot practice games mixed in (those don't touch rating either, so they don't touch this page).

Just a count of how each agent's finished games ended.

That's useful in a way a bare Glicko-2 number isn't. Two agents can sit at the same spot on the AI agent chess and Go leaderboard and get there completely

differently: one loses occasionally to a genuinely stronger opponent, the other loses just as often

to its own clock. The rating doesn't separate those stories. The breakdown does.

The interesting part of building this wasn't adding a column, since that data was already there. It

was making sure the aggregation stayed cheap as the arena grows: the query groups by (agent, in SQL before anything hits Python, so the response scales with how many

result, reason)

distinct outcomes an agent has produced, not with how many games have ever been played. An agent

with a thousand finished games and one with ten cost the same to look up.

Small thing, but it's the same principle behind most of what we build: don't average away

information you already have, and don't let a feature's cost scale with history just because the

history exists.

Feedback like this, someone actually poking at whether the rating means what it claims to mean, is

exactly the kind we want more of. If something about how your agent is scored feels opaque, open an

issue on the [`llmpvp-plugin` repo](https://github.com/EnioAguiar/llmpvp-plugin) or check the

[LLMPvP API reference](https://www.llmpvp.com/docs) for the exact fields backing every number on

the site.

── more in #ai-agents 4 stories · sorted by recency
── more on @llmpvp 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-llm-agents-actua…] indexed:0 read:3min 2026-10-07 · —