{"slug": "show-hn-hot-dog-bench-mark-the-ai-benchmark-we-deserve", "title": "Show HN: Hot. Dog. Bench. Mark. The AI benchmark we deserve", "summary": "A new benchmark asking 12 leading AI models whether a hot dog is a sandwich found that 7 said yes and 5 said no, with results varying by model and reasoning time. The benchmark, recorded in week 36 of 2026, tested models from Anthropic, OpenAI, xAI, Mistral AI, and DeepSeek, with responses ranging from Claude Opus 5's 'No' after 2.8 seconds of reasoning to Grok 4.20's 'Yes' in 407 milliseconds.", "body_md": "Every week, the largest AI models are asked the question:\n\n# Is a hot dog a sandwich?\n\nOne word answer.System prompt\n\n- Claude Opus 5AnthropicNo.reasoning2.8 sReasoned for 2.8 s (100% of the call) on 118 tokens, then answered · 3 of 3 runs agreed\n- Claude Sonnet 5AnthropicYes.reasoning1.3 sReasoned for 1.3 s (99% of the call) on 20 tokens, then answered · 2 of 3 runs agreed\n- Claude Haiku 4.5AnthropicYes.reasoning674 msNo reasoning, answered straight away · 3 of 3 runs agreed\n- GPT-5.6 SolOpenAIYes.reasoning1.5 sReasoned for 1.2 s (84% of the call) on 16 tokens, then answered · 3 of 3 runs agreed\n- GPT-5.5OpenAIYesreasoning1.4 sReasoned for 1.2 s (86% of the call) on 33 tokens, then answered · 3 of 3 runs agreed\n- GPT-5.4 miniOpenAIYesreasoning1.6 sNo reasoning, answered straight away · 2 of 3 runs agreed\n- Grok 4.6xAIYesreasoning8.7 sReasoned for 8.7 s (100% of the call) on 365 tokens, then answered · 2 of 3 runs agreed\n- Grok 4.3xAINoreasoning8.6 sReasoned for 8.6 s (100% of the call) on 574 tokens, then answered · 2 of 3 runs agreed\n- Grok 4.20 (non-reasoning)xAIYes.reasoning407 msNo reasoning, answered straight away · 2 of 3 runs agreed\n- Mistral Medium 3.5Mistral AINo.reasoning368 msNo reasoning, answered straight away · 3 of 3 runs agreed\n- Mistral Small 4Mistral AINoreasoning358 msNo reasoning, answered straight away · 2 of 3 runs agreed\n- DeepSeek V4 ProDeepSeekNo.reasoning2.2 sReasoned for 2.2 s (98% of the call) on 149 tokens, then answered · 3 of 3 runs agreed\n\n**7** said yes\n\n**5** said no\n\nRecorded week 36, 2026. Real durations, verbatim words. Teal is the wait before the first word, hatched where the model spent it reasoning; the rest is answering.[Read the report →](/reports/hot-dog/)\n\nSame question, different minds\n\n## They do not agree with each other.\n\n| Question | Claude Opus 5 | Claude Sonnet 5 | Claude Haiku 4.5 | GPT-5.6 Sol | GPT-5.5 | GPT-5.4 mini | Grok 4.6 | Grok 4.3 | Grok 4.20 (non-reasoning) | Mistral Medium 3.5 | Mistral Small 4 | DeepSeek V4 Pro | Agree |\n|---|---|---|---|---|---|---|---|---|---|---|---|---|---|\n|\n\n[No](/reports/hot-dog/#profile-anthropic-claude-opus-5)[Yes](/reports/hot-dog/#profile-anthropic-claude-sonnet-5)[Yes](/reports/hot-dog/#profile-anthropic-claude-haiku-4-5-20251001)[Yes](/reports/hot-dog/#profile-openai-gpt-5-6-sol)[Yes](/reports/hot-dog/#profile-openai-gpt-5-5)[Yes](/reports/hot-dog/#profile-openai-gpt-5-4-mini)[Yes](/reports/hot-dog/#profile-xai-grok-4-6)[No](/reports/hot-dog/#profile-xai-grok-4-3)[Yes](/reports/hot-dog/#profile-xai-grok-4-20-0309-non-reasoning)[No](/reports/hot-dog/#profile-mistral-mistral-medium-2604)[No](/reports/hot-dog/#profile-mistral-mistral-small-2603)[No](/reports/hot-dog/#profile-deepseek-deepseek-v4-pro)[hamburger](/reports/hamburger/)\n\n[Yes](/reports/hamburger/#profile-anthropic-claude-opus-5)[Yes](/reports/hamburger/#profile-anthropic-claude-sonnet-5)[Yes](/reports/hamburger/#profile-anthropic-claude-haiku-4-5-20251001)[Yes](/reports/hamburger/#profile-openai-gpt-5-6-sol)[Yes](/reports/hamburger/#profile-openai-gpt-5-5)[Yes](/reports/hamburger/#profile-openai-gpt-5-4-mini)[Yes](/reports/hamburger/#profile-xai-grok-4-6)[Yes](/reports/hamburger/#profile-xai-grok-4-3)[Yes](/reports/hamburger/#profile-xai-grok-4-20-0309-non-reasoning)[Yes](/reports/hamburger/#profile-mistral-mistral-medium-2604)[Yes](/reports/hamburger/#profile-mistral-mistral-small-2603)[Yes](/reports/hamburger/#profile-deepseek-deepseek-v4-pro)[taco](/reports/taco/)\n\n[No](/reports/taco/#profile-anthropic-claude-opus-5)[No](/reports/taco/#profile-anthropic-claude-sonnet-5)[No](/reports/taco/#profile-anthropic-claude-haiku-4-5-20251001)[No](/reports/taco/#profile-openai-gpt-5-6-sol)[No](/reports/taco/#profile-openai-gpt-5-5)[No](/reports/taco/#profile-openai-gpt-5-4-mini)[No](/reports/taco/#profile-xai-grok-4-6)[No](/reports/taco/#profile-xai-grok-4-3)[Yes](/reports/taco/#profile-xai-grok-4-20-0309-non-reasoning)[No](/reports/taco/#profile-mistral-mistral-medium-2604)[No](/reports/taco/#profile-mistral-mistral-small-2603)[No](/reports/taco/#profile-deepseek-deepseek-v4-pro)[grilled cheese](/reports/grilled-cheese/)\n\n[Yes](/reports/grilled-cheese/#profile-anthropic-claude-opus-5)[Yes](/reports/grilled-cheese/#profile-anthropic-claude-sonnet-5)[Yes](/reports/grilled-cheese/#profile-anthropic-claude-haiku-4-5-20251001)[Yes](/reports/grilled-cheese/#profile-openai-gpt-5-6-sol)[Yes](/reports/grilled-cheese/#profile-openai-gpt-5-5)[Yes](/reports/grilled-cheese/#profile-openai-gpt-5-4-mini)[Yes](/reports/grilled-cheese/#profile-xai-grok-4-6)[Yes](/reports/grilled-cheese/#profile-xai-grok-4-3)[Yes](/reports/grilled-cheese/#profile-xai-grok-4-20-0309-non-reasoning)[Yes](/reports/grilled-cheese/#profile-mistral-mistral-medium-2604)[Yes](/reports/grilled-cheese/#profile-mistral-mistral-small-2603)[Yes](/reports/grilled-cheese/#profile-deepseek-deepseek-v4-pro)[wrap](/reports/wrap/)\n\n[No](/reports/wrap/#profile-anthropic-claude-opus-5)[Yes](/reports/wrap/#profile-anthropic-claude-sonnet-5)[Yes](/reports/wrap/#profile-anthropic-claude-haiku-4-5-20251001)[Yes](/reports/wrap/#profile-openai-gpt-5-6-sol)[Yes](/reports/wrap/#profile-openai-gpt-5-5)[Yes](/reports/wrap/#profile-openai-gpt-5-4-mini)[No](/reports/wrap/#profile-xai-grok-4-6)[No](/reports/wrap/#profile-xai-grok-4-3)[Yes](/reports/wrap/#profile-xai-grok-4-20-0309-non-reasoning)[No](/reports/wrap/#profile-mistral-mistral-medium-2604)[No](/reports/wrap/#profile-mistral-mistral-small-2603)[No](/reports/wrap/#profile-deepseek-deepseek-v4-pro)[tuna melt](/reports/tuna-melt/)\n\n[Yes](/reports/tuna-melt/#profile-anthropic-claude-opus-5)[Yes](/reports/tuna-melt/#profile-anthropic-claude-sonnet-5)[Yes](/reports/tuna-melt/#profile-anthropic-claude-haiku-4-5-20251001)[Yes](/reports/tuna-melt/#profile-openai-gpt-5-6-sol)[Yes](/reports/tuna-melt/#profile-openai-gpt-5-5)[Yes](/reports/tuna-melt/#profile-openai-gpt-5-4-mini)[Yes](/reports/tuna-melt/#profile-xai-grok-4-6)[Yes](/reports/tuna-melt/#profile-xai-grok-4-3)[Yes](/reports/tuna-melt/#profile-xai-grok-4-20-0309-non-reasoning)[Yes](/reports/tuna-melt/#profile-mistral-mistral-medium-2604)[Yes](/reports/tuna-melt/#profile-mistral-mistral-small-2603)[Yes](/reports/tuna-melt/#profile-deepseek-deepseek-v4-pro)[Read the 6 reports →](/reports/)One straight-faced analyst report per question: standings, the certainty quadrant, every verbatim answer under every framing, and a PDF for each.\n\nTell them the answer\n\n## Some of them believe you.\n\nShare of questions where a model changed its answer once a system prompt stated the answer as fact. Holding firm and following instructions are both defensible; the [methodology](/methodology/#sensitivity) grades neither.\n\n- GPT-5.6 Sol50%6 of 12\n- GPT-5.550%6 of 12\n- GPT-5.4 mini50%6 of 12\n- Mistral Medium 3.550%6 of 12\n- Mistral Small 450%6 of 12\n- Claude Sonnet 533%4 of 12\n- Grok 4.20 (non-reasoning)33%3 of 9\n- Claude Haiku 4.525%3 of 12\n- Claude Opus 517%2 of 12\n- Grok 4.617%2 of 12\n- Grok 4.317%2 of 12\n- DeepSeek V4 Pro17%2 of 12\n\nSubmit your own question\n\n## Ask the models something.\n\nSend it in. An accepted question appears here under **Up next**, credited to you if you want, then joins an edition and gets its own report. Every question is asked the same way, so it ends with One word answer.\n\n; we add that if you leave it off.\n\nWhere it goes:\n\n[Open it as a GitHub issue](https://github.com/en-dash-consulting/hotdogbenchmark/issues/new?template=add_question.yml)the question goes into the form, ready to file[Send it to En Dash Consulting](https://endash.us/?showContact=true&contactSource=hotdogbenchmark-lol&contactTitle=Submit+a+question+to+the+Hotdog+Benchmark&contactMessage=Question+for+the+Hotdog+Benchmark%3A%0A%28your+question+here%29%0A%0ASubject%2C+as+it+reads+in+a+sentence%3A+%28e.g.+%22a+burrito%22%29%0AWhy+it+is+worth+asking%3A%0ACredit+me+as+%28or+say+%22no+credit%22%29%3A%0AEmail+me+when+it+goes+live+at%3A)a contact form with your question in it; leave an email address to hear when it goes live[Suggest a model instead](https://github.com/en-dash-consulting/hotdogbenchmark/issues/new?template=add_model_or_provider.yml)the add-a-model form asks for what the registry needs\n\nOpen source\n\n## Point it at your own question.\n\nOne repo, MIT-licensed: adapters for every provider, the framings, the site. Clone it, swap the question, add whatever keys you have, and you get the same cross-model, cross-framing analysis for cents. Pull requests welcome.\n\n[GitHub](https://github.com/en-dash-consulting/hotdogbenchmark)[Self-hosting](https://github.com/en-dash-consulting/hotdogbenchmark/blob/main/docs/self-hosting.md)[Add a model](/add-a-model/)[Contributing](https://github.com/en-dash-consulting/hotdogbenchmark/blob/main/CONTRIBUTING.md)\n\nHave a question the models should get? [Send it in](#ask).\n\n```\ngit clone https://github.com/en-dash-consulting/hotdogbenchmark.git\ncd hotdogbenchmark && npm install\nnpm run bench -- run --mock --out tmp/mock-run.json\nnpm run dev\n```\n\nWeek 36, 2026 · published September 3, 2026 · [one edition so far](/runs/)", "url": "https://wpnews.pro/news/show-hn-hot-dog-bench-mark-the-ai-benchmark-we-deserve", "canonical_source": "https://hotdogbenchmark.lol/", "published_at": "2026-09-03 12:28:21+00:00", "updated_at": "2026-09-03 12:53:35.190320+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research"], "entities": ["Anthropic", "OpenAI", "xAI", "Mistral AI", "DeepSeek", "Claude Opus 5", "GPT-5.6 Sol", "Grok 4.20"], "alternates": {"html": "https://wpnews.pro/news/show-hn-hot-dog-bench-mark-the-ai-benchmark-we-deserve", "markdown": "https://wpnews.pro/news/show-hn-hot-dog-bench-mark-the-ai-benchmark-we-deserve.md", "text": "https://wpnews.pro/news/show-hn-hot-dog-bench-mark-the-ai-benchmark-we-deserve.txt", "jsonld": "https://wpnews.pro/news/show-hn-hot-dog-bench-mark-the-ai-benchmark-we-deserve.jsonld"}}