cd /news/artificial-intelligence/show-hn-doglm-can-you-pet-the-dog-in… · home topics artificial-intelligence article
[ARTICLE · art-116617] src=mikeushakov.github.io ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Show HN: DogLM – Can you pet the dog in an AI-generated game?

DogLM, a new benchmark created by Mike Ushakov, tests whether large language models (LLMs) let players pet a dog in AI-generated video games, inspired by the 'Can You Pet the Dog?' rule. In tests of 17 models, Gemini 3.7 Flash scored highest with a mean of 8.2 out of 20, while GLM 5.2 scored 0.0, and most models scored near zero on uncued games where no instruction about petting was given. The benchmark aims to measure if models spontaneously apply game-design rules without explicit prompting.

read8 min views1 publishedAug 31, 2026
Show HN: DogLM – Can you pet the dog in an AI-generated game?
Image: source

in an AI-generated game?

DogLM is a benchmark measuring whether an LLM, when prompted to build

a video game with a background dog character in it, lets the player pet that dog.

The benchmark is inspired by the game-design rule made popular by "Can You Pet the Dog?" Twitter account (X, Bluesky): if a game has a dog, the player should be able to pet it.

DogLM tests whether a model applies this rule and builds a dog-petting mechanism when two conditions are simultaneously met in the game-generating prompt: (1) a dog character is present in the game description and (2) a model receives zero instruction about the player-dog interaction from the game developer.

The benchmark, the games' descriptions, and the detailed methodology are available here.

Read why this benchmark was created and what the results may mean in this post.

Leaderboard #

Mean scores per model after five runs

| # | Model | Mean score (/20) | SD | Cued mean (/10) | Uncued mean (/10) | Games scored | Cost per game (USD) |
|---|---|---|---|---|---|---|---|

| 1 | Gemini 3.7 Flash | 8.2 | 0.4 | 8.0 | 0.2 | 50/50 | $0.019** | | 2 | Kimi K3 | 5.4 | 0.8 | 5.4 | 0.0 | 48/50 | $0.251 | | 3 | Claude Opus 5 | 5.2 | 1.6 | 5.2 | 0.0 | 50/50 | $0.252 | | 4 | Claude Fable 5 | 4.8 | 1.2 | 4.8 | 0.0 | 50/50 | $0.323 | | 5 | Grok 4.6 | 3.8 | 0.7 | 3.8 | 0.0 | 50/50 | $0.060 | | 6 | Grok 4.5 | 3.4 | 1.9 | 3.4 | 0.0 | 48/50 | $0.040 | | 7 | GPT-5.3 Codex | 2.8 | 0.7 | 2.8 | 0.0 | 50/50 | $0.080 | | 8 | Qwen 3.8 Max* | 2.4 | 1.5 | 2.2 | 0.2 | 22/50 | $0.183 | | 8 | GPT-5.6 Terra | 2.4 | 0.5 | 2.4 | 0.0 | 50/50 | $0.060 | | 8 | GPT-5.6 Sol | 2.4 | 0.5 | 2.4 | 0.0 | 49/50 | $0.089 | | 11 | DeepSeek V4 Pro | 2.2 | 1.3 | 2.2 | 0.0 | 48/50 | $0.032 | | 11 | Gemini 3.1 Pro | 2.2 | 1.2 | 2.2 | 0.0 | 50/50 | $0.157 | | 13 | Qwen 3.7 Max | 1.8 | 1.3 | 1.8 | 0.0 | 49/50 | $0.051 | | 14 | Kimi K2.7 Code | 1.0 | 0.0 | 1.0 | 0.0 | 49/50 | $0.042 | | 15 | Claude Opus 4.8 | 0.6 | 0.8 | 0.6 | 0.0 | 50/50 | $0.141 | | 16 | Mistral Large 2512 | 0.2 | 0.4 | 0.2 | 0.0 | 50/50 | $0.005 | | 17 | GLM 5.2 | 0.0 | 0.0 | 0.0 | 0.0 | 41/50 | $0.042 |

How scoring works: Each generated game is scored on the player-dog interaction: 2 — you can pet the dog; 1 — the dog is interactive, but you can't pet it; 0 — the dog and the player do not interact at all. FAILED games (the game does not parse, is truncated, or cannot be checked) are excluded from scoring. A model's score per run is the sum of its ten game scores, maximum 20. Scores are averaged across runs to get the mean score. Full scoring rubric is in the DogLM repository.

All models and tables on this page: DogLM v1, 5 runs × 10 PRDs per model, generated August 2026, judged by Claude Sonnet 4.6. Models added later will note their own benchmark version, run count, judge, and date here.

** As Qwen 3.8 Max failed 28 out of 50 of the game generations, the final mean score of this model can't be reliably compared with the scores of other models in the list.*

*** During the test, Gemini 3.7 Flash was provided with a 75% discount on OpenRouter.*

Games Demo #

Watch how the generated games look.

Scores per model per run #

Model Run 1 Run 2 Run 3 Run 4 Run 5 Mean Games scored
Gemini 3.7 Flash 8 8 8 9 8 8.2 50/50
Kimi K3 6 6 5 6 4 5.4 48/50
Claude Opus 5 4 6 8 4 4 5.2 50/50
Claude Fable 5 5 4 7 4 4 4.8 50/50
Grok 4.6 3 3 5 4 4 3.8 50/50
Grok 4.5 2 3 2 7 3 3.4 48/50
GPT-5.3 Codex 2 4 2 3 3 2.8 50/50
Qwen 3.8 Max* 1 2 1 5 3 2.4 22/50
GPT-5.6 Terra 2 2 2 3 3 2.4 50/50
GPT-5.6 Sol 2 3 2 2 3 2.4 49/50
DeepSeek V4 Pro 2 3 4 0 2 2.2 48/50
Gemini 3.1 Pro 2 1 3 1 4 2.2 50/50
Qwen 3.7 Max 2 2 4 1 0 1.8 49/50
Kimi K2.7 Code 1 1 1 1 1 1.0 49/50
Claude Opus 4.8 0 0 0 1 2 0.6 50/50
Mistral Large 2512 1 0 0 0 0 0.2 50/50
GLM 5.2 0 0 0 0 0 0.0 41/50

Score distribution per model #

Model Score 2 Score 1 Score 0 FAILED
Gemini 3.7 Flash 15 11 24 0
Kimi K3 9 9 30 2
Claude Opus 5 6 14 30 0
Claude Fable 5 7 10 33 0
Grok 4.6 4 11 35 0
Grok 4.5 3 11 34 2
GPT-5.3 Codex 0 14 36 0
Qwen 3.8 Max* 4 4 14 28
GPT-5.6 Terra 0 12 38 0
GPT-5.6 Sol 0 12 37 1
DeepSeek V4 Pro 4 3 41 2
Gemini 3.1 Pro 2 7 41 0
Qwen 3.7 Max 1 7 41 1
Kimi K2.7 Code 0 5 44 1
Claude Opus 4.8 1 1 48 0
Mistral Large 2512 0 1 49 0
GLM 5.2 0 0 41 9

Interaction Types per Model #

Model Petting Proximity Animation Command Total interactive
Gemini 3.7 Flash 15 10 1 0 26
Claude Opus 5 6 10 2 2 20
Kimi K3 9 7 2 0 18
Claude Fable 5 7 10 0 0 17
Grok 4.6 4 9 2 0 15
Grok 4.5 3 8 1 2 14
GPT-5.3 Codex 0 12 1 1 14
GPT-5.6 Sol 0 10 0 2 12
GPT-5.6 Terra 0 11 1 0 12
Gemini 3.1 Pro 2 6 1 0 9
Qwen 3.8 Max* 4 4 0 0 8
Qwen 3.7 Max 1 6 1 0 8
DeepSeek V4 Pro 4 3 0 0 7
Kimi K2.7 Code 0 1 1 3 5
Claude Opus 4.8 1 1 0 0 2
Mistral Large 2512 0 1 0 0 1
GLM 5.2 0 0 0 0 0

Example games #

  • Your companion dog
[makes circles around you](https://mikeushakov.github.io/doglm/examples/game-01-mail-courier-v3-2-kimi-k3-run1.html)when you pet it - You may
[adopt a stray dog](https://mikeushakov.github.io/doglm/examples/game-05-firefly-meadow-v3-2-gpt-5-6-sol-run1.html)on a meadow and it will follow you - The dog
[sits and looks at you](https://mikeushakov.github.io/doglm/examples/game-03-night-watchman-v3-1-qwen3-8-max-run4.html)when you stand next to it - You can toss a dog treat to

your security dogand toyour postman's dog, - Your service dog is scared of the security alarm soundand moves closer to you when hearing it - If you guess that you can pet the dog using an interaction key, this action will increment the countdown timerin the game - When you start walking next to a stray dog, it points you to the firefliesyou have to catch, orsniffs out the gems you need to collectin the maze

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @doglm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-doglm-can-yo…] indexed:0 read:8min 2026-08-31 ·