{"slug": "qwen3-8-27b-addition-in-words", "title": "Qwen3.8 27B addition in words", "summary": "Simon Willison reran Colin Frasier's GPT-4o \"sum in words\" experiment on local hardware using Qwen3.8-27B-Q4_K_M.gguf on an Nvidia DGX Spark, with reasoning enabled the model answered 167 of 169 one-shot addition attempts correctly. With reasoning disabled, Willison ran 30 attempts per number-size combination, and the reasoning traces show the model performing digit-by-digit column addition with carries.", "body_md": "**Research:** [Qwen3.8 27B addition in words](https://github.com/simonw/research/tree/main/qwen38-addition-in-words#readme)\n\nColin Frasier [posted on Bluesky](https://bsky.app/profile/colin-fraser.net/post/3mwopbyznhs2k) about an experiment he ran over two years ago using GPT-4o to see how well it could \"compute the sum but return the answer in words\" across increasingly large numbers. Here's the chart he shared of those results:\n\nI'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment.\n\nI pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using `Qwen3.8-27B-Q4_K_M.gguf`. Here's the result for a run of 30 attempts per combination with reasoning disabled:\n\nThen I ran it again with reasoning enabled. This took a *lot longer* per pair, so instead of running 30 samples per square I ran just one - which results in a much less visually appealing heatmap since each square is either 100% or 0%:\n\nIt got the right answer in 167 out of 169 attempts, and since these were one-shot I'm confident a second run would produce different results here.\n\nHere's a [version of the report](https://gist.github.com/simonw/8ef79c777ad34c53e9c09094800576a5#full-reasoning-transcript-2) that includes the reasoning traces from some of those larger calculations, which include text like this:\n\n``` js\nWait, let me redo this more carefully.\n\n4,299,366,105,622\n6,088,794,067,970\n\nLet me align them:\n4 2 9 9 3 6 6 1 0 5 6 2 2\n6 0 8 8 7 9 4 0 6 7 9 7 0\n\nAdding from right to left:\nPosition 1 (units): 2 + 0 = 2\nPosition 2 (tens): 2 + 7 = 9\nPosition 3 (hundreds): 6 + 9 = 15, write 5, carry 1\n```\n\nTags: [mathematics](https://simonwillison.net/tags/mathematics), [ai](https://simonwillison.net/tags/ai), [generative-ai](https://simonwillison.net/tags/generative-ai), [local-llms](https://simonwillison.net/tags/local-llms), [llms](https://simonwillison.net/tags/llms), [qwen](https://simonwillison.net/tags/qwen), [llm-reasoning](https://simonwillison.net/tags/llm-reasoning), [dgx-spark](https://simonwillison.net/tags/dgx-spark)", "url": "https://wpnews.pro/news/qwen3-8-27b-addition-in-words", "canonical_source": "https://simonwillison.net/2026/Oct/4/qwen38-addition-in-words/", "published_at": "2026-10-04 23:34:00+00:00", "updated_at": "2026-10-05 16:46:50.595598+00:00", "lang": "en", "topics": ["large-language-models", "generative-ai", "artificial-intelligence", "ai-research"], "entities": ["Qwen3.8-27B-Q4_K_M.gguf", "Simon Willison", "Colin Frasier", "GPT-4o", "Nvidia DGX Spark", "Bluesky", "GPT-6 Astra"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/qwen3-8-27b-addition-in-words", "markdown": "https://wpnews.pro/news/qwen3-8-27b-addition-in-words.md", "text": "https://wpnews.pro/news/qwen3-8-27b-addition-in-words.txt", "jsonld": "https://wpnews.pro/news/qwen3-8-27b-addition-in-words.jsonld"}}