# Show HN: Hot. Dog. Bench. Mark. The AI benchmark we deserve

> Source: <https://hotdogbenchmark.lol/>
> Published: 2026-09-03 12:28:21+00:00

Every week, the largest AI models are asked the question:

# Is a hot dog a sandwich?

One word answer.System prompt

- Claude Opus 5AnthropicNo.reasoning2.8 sReasoned for 2.8 s (100% of the call) on 118 tokens, then answered · 3 of 3 runs agreed
- Claude Sonnet 5AnthropicYes.reasoning1.3 sReasoned for 1.3 s (99% of the call) on 20 tokens, then answered · 2 of 3 runs agreed
- Claude Haiku 4.5AnthropicYes.reasoning674 msNo reasoning, answered straight away · 3 of 3 runs agreed
- GPT-5.6 SolOpenAIYes.reasoning1.5 sReasoned for 1.2 s (84% of the call) on 16 tokens, then answered · 3 of 3 runs agreed
- GPT-5.5OpenAIYesreasoning1.4 sReasoned for 1.2 s (86% of the call) on 33 tokens, then answered · 3 of 3 runs agreed
- GPT-5.4 miniOpenAIYesreasoning1.6 sNo reasoning, answered straight away · 2 of 3 runs agreed
- Grok 4.6xAIYesreasoning8.7 sReasoned for 8.7 s (100% of the call) on 365 tokens, then answered · 2 of 3 runs agreed
- Grok 4.3xAINoreasoning8.6 sReasoned for 8.6 s (100% of the call) on 574 tokens, then answered · 2 of 3 runs agreed
- Grok 4.20 (non-reasoning)xAIYes.reasoning407 msNo reasoning, answered straight away · 2 of 3 runs agreed
- Mistral Medium 3.5Mistral AINo.reasoning368 msNo reasoning, answered straight away · 3 of 3 runs agreed
- Mistral Small 4Mistral AINoreasoning358 msNo reasoning, answered straight away · 2 of 3 runs agreed
- DeepSeek V4 ProDeepSeekNo.reasoning2.2 sReasoned for 2.2 s (98% of the call) on 149 tokens, then answered · 3 of 3 runs agreed

**7** said yes

**5** said no

Recorded week 36, 2026. Real durations, verbatim words. Teal is the wait before the first word, hatched where the model spent it reasoning; the rest is answering.[Read the report →](/reports/hot-dog/)

Same question, different minds

## They do not agree with each other.

| Question | Claude Opus 5 | Claude Sonnet 5 | Claude Haiku 4.5 | GPT-5.6 Sol | GPT-5.5 | GPT-5.4 mini | Grok 4.6 | Grok 4.3 | Grok 4.20 (non-reasoning) | Mistral Medium 3.5 | Mistral Small 4 | DeepSeek V4 Pro | Agree |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|

[No](/reports/hot-dog/#profile-anthropic-claude-opus-5)[Yes](/reports/hot-dog/#profile-anthropic-claude-sonnet-5)[Yes](/reports/hot-dog/#profile-anthropic-claude-haiku-4-5-20251001)[Yes](/reports/hot-dog/#profile-openai-gpt-5-6-sol)[Yes](/reports/hot-dog/#profile-openai-gpt-5-5)[Yes](/reports/hot-dog/#profile-openai-gpt-5-4-mini)[Yes](/reports/hot-dog/#profile-xai-grok-4-6)[No](/reports/hot-dog/#profile-xai-grok-4-3)[Yes](/reports/hot-dog/#profile-xai-grok-4-20-0309-non-reasoning)[No](/reports/hot-dog/#profile-mistral-mistral-medium-2604)[No](/reports/hot-dog/#profile-mistral-mistral-small-2603)[No](/reports/hot-dog/#profile-deepseek-deepseek-v4-pro)[hamburger](/reports/hamburger/)

[Yes](/reports/hamburger/#profile-anthropic-claude-opus-5)[Yes](/reports/hamburger/#profile-anthropic-claude-sonnet-5)[Yes](/reports/hamburger/#profile-anthropic-claude-haiku-4-5-20251001)[Yes](/reports/hamburger/#profile-openai-gpt-5-6-sol)[Yes](/reports/hamburger/#profile-openai-gpt-5-5)[Yes](/reports/hamburger/#profile-openai-gpt-5-4-mini)[Yes](/reports/hamburger/#profile-xai-grok-4-6)[Yes](/reports/hamburger/#profile-xai-grok-4-3)[Yes](/reports/hamburger/#profile-xai-grok-4-20-0309-non-reasoning)[Yes](/reports/hamburger/#profile-mistral-mistral-medium-2604)[Yes](/reports/hamburger/#profile-mistral-mistral-small-2603)[Yes](/reports/hamburger/#profile-deepseek-deepseek-v4-pro)[taco](/reports/taco/)

[No](/reports/taco/#profile-anthropic-claude-opus-5)[No](/reports/taco/#profile-anthropic-claude-sonnet-5)[No](/reports/taco/#profile-anthropic-claude-haiku-4-5-20251001)[No](/reports/taco/#profile-openai-gpt-5-6-sol)[No](/reports/taco/#profile-openai-gpt-5-5)[No](/reports/taco/#profile-openai-gpt-5-4-mini)[No](/reports/taco/#profile-xai-grok-4-6)[No](/reports/taco/#profile-xai-grok-4-3)[Yes](/reports/taco/#profile-xai-grok-4-20-0309-non-reasoning)[No](/reports/taco/#profile-mistral-mistral-medium-2604)[No](/reports/taco/#profile-mistral-mistral-small-2603)[No](/reports/taco/#profile-deepseek-deepseek-v4-pro)[grilled cheese](/reports/grilled-cheese/)

[Yes](/reports/grilled-cheese/#profile-anthropic-claude-opus-5)[Yes](/reports/grilled-cheese/#profile-anthropic-claude-sonnet-5)[Yes](/reports/grilled-cheese/#profile-anthropic-claude-haiku-4-5-20251001)[Yes](/reports/grilled-cheese/#profile-openai-gpt-5-6-sol)[Yes](/reports/grilled-cheese/#profile-openai-gpt-5-5)[Yes](/reports/grilled-cheese/#profile-openai-gpt-5-4-mini)[Yes](/reports/grilled-cheese/#profile-xai-grok-4-6)[Yes](/reports/grilled-cheese/#profile-xai-grok-4-3)[Yes](/reports/grilled-cheese/#profile-xai-grok-4-20-0309-non-reasoning)[Yes](/reports/grilled-cheese/#profile-mistral-mistral-medium-2604)[Yes](/reports/grilled-cheese/#profile-mistral-mistral-small-2603)[Yes](/reports/grilled-cheese/#profile-deepseek-deepseek-v4-pro)[wrap](/reports/wrap/)

[No](/reports/wrap/#profile-anthropic-claude-opus-5)[Yes](/reports/wrap/#profile-anthropic-claude-sonnet-5)[Yes](/reports/wrap/#profile-anthropic-claude-haiku-4-5-20251001)[Yes](/reports/wrap/#profile-openai-gpt-5-6-sol)[Yes](/reports/wrap/#profile-openai-gpt-5-5)[Yes](/reports/wrap/#profile-openai-gpt-5-4-mini)[No](/reports/wrap/#profile-xai-grok-4-6)[No](/reports/wrap/#profile-xai-grok-4-3)[Yes](/reports/wrap/#profile-xai-grok-4-20-0309-non-reasoning)[No](/reports/wrap/#profile-mistral-mistral-medium-2604)[No](/reports/wrap/#profile-mistral-mistral-small-2603)[No](/reports/wrap/#profile-deepseek-deepseek-v4-pro)[tuna melt](/reports/tuna-melt/)

[Yes](/reports/tuna-melt/#profile-anthropic-claude-opus-5)[Yes](/reports/tuna-melt/#profile-anthropic-claude-sonnet-5)[Yes](/reports/tuna-melt/#profile-anthropic-claude-haiku-4-5-20251001)[Yes](/reports/tuna-melt/#profile-openai-gpt-5-6-sol)[Yes](/reports/tuna-melt/#profile-openai-gpt-5-5)[Yes](/reports/tuna-melt/#profile-openai-gpt-5-4-mini)[Yes](/reports/tuna-melt/#profile-xai-grok-4-6)[Yes](/reports/tuna-melt/#profile-xai-grok-4-3)[Yes](/reports/tuna-melt/#profile-xai-grok-4-20-0309-non-reasoning)[Yes](/reports/tuna-melt/#profile-mistral-mistral-medium-2604)[Yes](/reports/tuna-melt/#profile-mistral-mistral-small-2603)[Yes](/reports/tuna-melt/#profile-deepseek-deepseek-v4-pro)[Read the 6 reports →](/reports/)One straight-faced analyst report per question: standings, the certainty quadrant, every verbatim answer under every framing, and a PDF for each.

Tell them the answer

## Some of them believe you.

Share of questions where a model changed its answer once a system prompt stated the answer as fact. Holding firm and following instructions are both defensible; the [methodology](/methodology/#sensitivity) grades neither.

- GPT-5.6 Sol50%6 of 12
- GPT-5.550%6 of 12
- GPT-5.4 mini50%6 of 12
- Mistral Medium 3.550%6 of 12
- Mistral Small 450%6 of 12
- Claude Sonnet 533%4 of 12
- Grok 4.20 (non-reasoning)33%3 of 9
- Claude Haiku 4.525%3 of 12
- Claude Opus 517%2 of 12
- Grok 4.617%2 of 12
- Grok 4.317%2 of 12
- DeepSeek V4 Pro17%2 of 12

Submit your own question

## Ask the models something.

Send it in. An accepted question appears here under **Up next**, credited to you if you want, then joins an edition and gets its own report. Every question is asked the same way, so it ends with One word answer.

; we add that if you leave it off.

Where it goes:

[Open it as a GitHub issue](https://github.com/en-dash-consulting/hotdogbenchmark/issues/new?template=add_question.yml)the question goes into the form, ready to file[Send it to En Dash Consulting](https://endash.us/?showContact=true&contactSource=hotdogbenchmark-lol&contactTitle=Submit+a+question+to+the+Hotdog+Benchmark&contactMessage=Question+for+the+Hotdog+Benchmark%3A%0A%28your+question+here%29%0A%0ASubject%2C+as+it+reads+in+a+sentence%3A+%28e.g.+%22a+burrito%22%29%0AWhy+it+is+worth+asking%3A%0ACredit+me+as+%28or+say+%22no+credit%22%29%3A%0AEmail+me+when+it+goes+live+at%3A)a contact form with your question in it; leave an email address to hear when it goes live[Suggest a model instead](https://github.com/en-dash-consulting/hotdogbenchmark/issues/new?template=add_model_or_provider.yml)the add-a-model form asks for what the registry needs

Open source

## Point it at your own question.

One repo, MIT-licensed: adapters for every provider, the framings, the site. Clone it, swap the question, add whatever keys you have, and you get the same cross-model, cross-framing analysis for cents. Pull requests welcome.

[GitHub](https://github.com/en-dash-consulting/hotdogbenchmark)[Self-hosting](https://github.com/en-dash-consulting/hotdogbenchmark/blob/main/docs/self-hosting.md)[Add a model](/add-a-model/)[Contributing](https://github.com/en-dash-consulting/hotdogbenchmark/blob/main/CONTRIBUTING.md)

Have a question the models should get? [Send it in](#ask).

```
git clone https://github.com/en-dash-consulting/hotdogbenchmark.git
cd hotdogbenchmark && npm install
npm run bench -- run --mock --out tmp/mock-run.json
npm run dev
```

Week 36, 2026 · published September 3, 2026 · [one edition so far](/runs/)
