cd /news/large-language-models/qwen3-8-27b-addition-in-words · home › topics › large-language-models › article
[ARTICLE · art-145507] src=simonwillison.net ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Qwen3.8 27B addition in words

Simon Willison reran Colin Frasier's GPT-4o "sum in words" experiment on local hardware using Qwen3.8-27B-Q4_K_M.gguf on an Nvidia DGX Spark, with reasoning enabled the model answered 167 of 169 one-shot addition attempts correctly. With reasoning disabled, Willison ran 30 attempts per number-size combination, and the reasoning traces show the model performing digit-by-digit column addition with carries.

read2 min views1 publishedOct 4, 2026

Research: Qwen3.8 27B addition in words

Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results:

I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment.

I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf. Here's the result for a run of 30 attempts per combination with reasoning disabled:

Then I ran it again with reasoning enabled. This took a lot longer per pair, so instead of running 30 samples per square I ran just one - which results in a much less visually appealing heatmap since each square is either 100% or 0%:

It got the right answer in 167 out of 169 attempts, and since these were one-shot I'm confident a second run would produce different results here.

Here's a version of the report that includes the reasoning traces from some of those larger calculations, which include text like this:

Wait, let me redo this more carefully.

4,299,366,105,622
6,088,794,067,970

Let me align them:
4 2 9 9 3 6 6 1 0 5 6 2 2
6 0 8 8 7 9 4 0 6 7 9 7 0

Adding from right to left:
Position 1 (units): 2 + 0 = 2
Position 2 (tens): 2 + 7 = 9
Position 3 (hundreds): 6 + 9 = 15, write 5, carry 1

Tags: mathematics, ai, generative-ai, local-llms, llms, qwen, llm-reasoning, dgx-spark

── more in #large-language-models 4 stories · sorted by recency
── more on @qwen3.8-27b-q4_k_m.gguf 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/qwen3-8-27b-addition…] indexed:0 read:2min 2026-10-04 · —