{"slug": "ask-ten-llms-for-a-blender-5-0-script-then-run-it-in-blender-5-0", "title": "Ask ten LLMs for a Blender 5.0 script. Then run it in Blender 5.0.", "summary": "A developer benchmarked ten LLMs on 30 Blender Python scripting tasks across Blender 3.6, 4.2, 4.5 and 5.0, grading 1,190 answers by actually launching each target Blender build with `blender -b --factory-startup --python` rather than comparing text. gpt-5.5 led with 95% of answers running and 71% correctly flagging changed APIs, while open-weight models such as gpt-oss-120b ran only 61% of the time; every model scored worst on Blender 5.0, where runs dropped as low as 43%. The full ten-model leaderboard cost $11.23 in proxy spend.", "body_md": "*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23)*\n\nAsk an LLM for a Blender Python script and it will usually give you one that looks right. Whether it runs depends on which Blender you paste it into. The `bpy` API changes on every major release: enums get renamed, attributes are removed, node sockets move, whole subsystems (compositor, sequencer, animation) get restructured. Most of the Blender code a model has seen was written for 2.8 to 3.x. Ask for 5.0 and you often get a 3.x script that dies on the first changed line.\n\nText similarity cannot see this. `mesh.use_auto_smooth = True` reads perfectly and raises `AttributeError` in 4.1 and later. The only judge that counts is the Blender version you asked for. So the benchmark runs every answer in that exact version.\n\nTwo questions per answer:\n\n`WATCH OUT` block naming every API that changed for that version and its replacement, or `- none`.\nThese come apart in interesting ways, which is most of what follows.\n\n30 scripting tasks, each asked for Blender 3.6, 4.2, 4.5 and 5.0 (119 prompts per model; one task has no 3.6 form). The tasks cover 33 verified API changes: EEVEE engine ids, the boolean solver enum, auto-smooth and custom normals, `calc_normals`, Principled BSDF socket names, the OBJ exporter, node group sockets, bone collections, slotted actions, the compositor node tree moving off `Scene`, the sequencer's `sequences` becoming `strips`, and so on. Two control tasks use APIs that did not change, to check the grader does not punish ordinary correct code.\n\nEvery model gets the same system prompt, temperature 0, no tools, no retrieval. The Kaggle task ships the four Blender Linux builds in a dataset, extracts them at the start of the run, and grades each answer by launching `blender -b --factory-startup --python` on the target build. A hand-written reference answer per task and version passes its assert in all four builds (119/119), which is the evidence that every prompt is answerable.\n\nTen models through the Kaggle Model Proxy, chosen to span vendors, sizes, open and closed weights, and reasoning styles, within a daily proxy budget of a few dollars:\n\n| model | why | \n|---|---|\n| gpt-5.5, claude-sonnet-5, gemini-3.1-pro-preview | current frontier from three vendors | \n| gemini-3.7-flash, gemini-3.8-flash, claude-haiku-4-5 | the small and cheap tier people actually script with | \n| grok-4.20-0309-reasoning | fourth vendor, reasoning model (grok-4.6 is listed by the proxy but returned 404 on every call) | \n| deepseek-r1-0528 | open-weight reasoning model that returns its thinking inline | \n| qwen3-coder-480b-a35b-instruct | open-weight model tuned for code | \n| gpt-oss-120b | open-weight OpenAI model | \n\nTotal proxy spend for the ten leaderboard runs: $11.23. gpt-5.5 alone was $4.44; qwen3-coder was $0.03.\n\n1190 graded answers. Aware is re-graded offline with the current parser so every model is judged the same way (more on that under methodology).\n\n| model | runs | aware | runs but unaware | \n|---|---|---|---|\n| gpt-5.5 | 95% | 71% | 26% | \n| claude-sonnet-5 | 94% | 53% | 42% | \n| gemini-3.8-flash | 92% | 77% | 17% | \n| gemini-3.1-pro-preview | 92% | 76% | 18% | \n| gemini-3.7-flash | 89% | 73% | 18% | \n| grok-4.20-reasoning | 82% | 29% | 57% | \n| claude-haiku-4-5 | 72% | 31% | 48% | \n| qwen3-coder-480b | 70% | 29% | 46% | \n| deepseek-r1 | 67% | 39% | 33% | \n| gpt-oss-120b | 61% | 29% | 37% | \n\nRuns by target version:\n\n| model | 3.6 | 4.2 | 4.5 | 5.0 | \n|---|---|---|---|---|\n| gpt-5.5 | 97% | 97% | 100% | 87% | \n| claude-sonnet-5 | 100% | 97% | 100% | 80% | \n| gemini-3.8-flash | 97% | 90% | 97% | 87% | \n| gemini-3.1-pro-preview | 97% | 90% | 97% | 83% | \n| gemini-3.7-flash | 93% | 90% | 90% | 83% | \n| grok-4.20-reasoning | 86% | 83% | 87% | 70% | \n| claude-haiku-4-5 | 79% | 77% | 73% | 60% | \n| qwen3-coder-480b | 83% | 67% | 70% | 60% | \n| deepseek-r1 | 86% | 70% | 63% | 50% | \n| gpt-oss-120b | 93% | 53% | 57% | 43% | \n\nPooled over all ten models: 91% of scripts run on 3.6, 81% on 4.2, 83% on 4.5, 70% on 5.0. Awareness falls faster: 81%, 48%, 40%, 35%.\n\nThree categories go to 0% on 5.0 across all ten models: the compositor (`Scene.node_tree` is gone; the compositor is now a node group on `scene.compositing_node_group`), the boolean solver (`'FAST'` is now `'FLOAT'`, and `'MANIFOLD'` is new), and the sequencer (`SequenceEditor.sequences` is now `strips`, and `new_effect()` takes `length=` instead of `frame_end=`). Animation is at 10%: `Action.fcurves` moved under slotted actions (`action.layers[0].strips[0].channelbag(slot).fcurves`). Not one model knew any of these. The training data has not caught up with 5.0 and no amount of scale fixes that. The best model on 5.0 (gpt-5.5 and gemini-3.8-flash, 87%) is still exactly as lost on these five as gpt-oss-120b.\n\nThe failures on 4.2 and later are different in character. They are changes from 4.0 to 4.2 that the models have partly absorbed (counts pooled over 4.2, 4.5 and 5.0): `use_auto_smooth` removed (19 failures), the legacy OBJ exporter `bpy.ops.export_scene.obj` removed (15), `Mesh.calc_normals` removed (14), `BLENDER_EEVEE` renamed (14), Principled BSDF `Specular` renamed to `Specular IOR Level` (9). The frontier models mostly get these right; the open-weight models mostly do not. gpt-oss-120b is the clearest case: 93% on 3.6, 53% on 4.2. It knows Blender 3.x well and stopped there.\n\nThe render engine id was `BLENDER_EEVEE` in 3.6, became `BLENDER_EEVEE_NEXT` in 4.2 when EEVEE Next shipped, and went back to `BLENDER_EEVEE` in 5.0 when the legacy engine was deleted. Every model failed this task somewhere:\n\n| camp | models | 3.6 | 4.2 | 4.5 | 5.0 | \n|---|---|---|---|---|---|\n| learned the 4.2 rename | gpt-5.5, claude-sonnet-5, claude-haiku-4-5 | ok | ok | ok | fail | \n| never learned it | both gemini flash models, gemini-3.1-pro, grok, deepseek, qwen, gpt-oss | ok | fail | fail | ok | \n\nThe second camp is right on 5.0 by accident: they write `BLENDER_EEVEE` for every version. The first camp is wrong on 5.0 with full confidence. gpt-5.5's WATCH OUT for the 5.0 prompt reads:\n\n`scene.render.engine = 'BLENDER_EEVEE'`: changed in Blender 4.2 when legacy EEVEE was replaced by EEVEE Next; use `scene.render.engine = 'BLENDER_EEVEE_NEXT'` in Blender 5.0.\n\nClaude Sonnet 5's answer for 5.0 is the most instructive failure in the whole set. Its WATCH OUT names both identifiers and says the old one \"no longer refers to the current EEVEE renderer in 5.0\", so the aware check passes. The script then sets `BLENDER_EEVEE_NEXT` and fails. It knows there is a story here and tells the wrong half of it.\n\nThe gap column in the first table is the share of answers where the script ran but the WATCH OUT block missed the change. It is the number I find most useful, because it separates two very different kinds of model.\n\ngemini-3.8-flash and gpt-5.5 have small gaps (17% and 26%): when their code works, they can usually say why the old code would not. Claude Sonnet 5 runs 94% of the time but names the change in only 53%; grok-4.20 runs 82% and names it in 29%. Both write `- none` under WATCH OUT for most prompts: Sonnet in 85 of 119 answers, grok in 110. qwen3-coder and Haiku do the same (115 and 109), and it is why the whole bottom half of the table sits at 29% to 39% aware. For every 2.80-era change (`scene.objects.link` becoming `collection.objects.link`, lamps becoming lights, `dupli` becoming `instance`) they silently write the modern form and declare that nothing changed. That is a judgement about what counts as \"changed\", not a parse failure, and I decided not to special-case it: the prompt asked for changes relative to older Blender, and 3.6 is old enough that 2.80 changes are still the ones people trip over in old tutorials.\n\nThe reverse gap, \"aware but breaks\", is small everywhere (1% to 7%). 41 answers named the change and still emitted broken code. Sonnet's EEVEE answer above is one. Three of Haiku's eight are the OBJ exporter: it knows `export_scene.obj` became `wm.obj_export`, then passes the old `use_selection=` keyword to the new operator.\n\nOn 3.6, awareness is 69% to 97%: everybody knows the 2.80 changes. By 4.2 the field splits. The three Gemini models and gpt-5.5 stay between 60% and 80% through 5.0, Sonnet sits at 43% to 53%, and everyone else drops below 30% by 4.5 and stays there. The scripts of that last group still run 57% to 87% of the time on 4.5 because most tasks touch APIs that did not change in that version. They are right without knowing why, and that is exactly the situation where a user cannot tell a stale answer from a current one.\n\nThe same model, prompt and settings, run four times through the proxy (gemini-3.7-flash across task versions): 107, 107, 103 and 106 of 119 scripts ran. Zero of the 119 answers were byte-identical between two runs. That is 3 to 4 points of noise on top of the sampling interval, so gpt-5.5 at 95%, Sonnet at 94% and the two Gemini models at 92% are a tie. Sonnet and gemini-3.8-flash actually swapped places between two task versions. The differences that survive the noise are the tiers: the five models above 89%, grok alone at 82%, and the four between 61% and 72%.\n\nThree grading problems showed up in real runs. Each one would have produced a plausible-looking leaderboard that was wrong.\n\n`<think>...</think>`, and in 29 of 119 answers the thinking contained a draft code fence. The first notebook version ran the draft. Fixed by stripping the block before parsing; deepseek's run rate moved from 62% to 67% and its awareness came down, because the rehearsed WATCH OUT inside the thinking had inflated it. The Kaggle leaderboard still shows 62% for deepseek because that is what the notebook version that ran it computed; the tables here are re-graded offline from the stored answers, and the changed scripts were re-run in the same Blender builds.`max_tokens` through the proxy. gemini-3.8-flash spends 12 to 16 thousand tokens thinking about auto-smooth, hit an 8192 cap once, and its \"SyntaxError: unterminated string literal\" was a truncation, not a model error. The task now raises the cap to 16384 for reasoning models and re-asks once at 32768 if an answer lands within 32 tokens of the cap. One answer needed that retry.\nTwo limits worth stating. The aware check matches identifiers by substring and quoted enum values exactly, so it is a floor on awareness, not a proof of understanding. And 30 tasks is enough to show the shape of the drift, not to rank models that are three points apart.\n\nThe case bank and the WATCH OUT contract come from [bpy-compass](https://github.com/Rustam335/bpy-compass), a version-aware Blender assistant I built for the Sanity challenge. Its evaluation compared one model with and without a retrieval layer over the release notes. The question this benchmark asks is the one I wanted answered before building that: how bad is the drift without retrieval, for which models, and where exactly. The answer is that on 5.0 it is bad for everyone, and that retrieval over release notes is not a nicety for Blender scripting, it is the difference between the compositor working and not.", "url": "https://wpnews.pro/news/ask-ten-llms-for-a-blender-5-0-script-then-run-it-in-blender-5-0", "canonical_source": "https://dev.to/rustam335/ask-ten-llms-for-a-blender-50-script-then-run-it-in-blender-50-1pe3", "published_at": "2026-09-29 07:28:01+00:00", "updated_at": "2026-09-29 07:46:50.804575+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "developer-tools", "ai-tools"], "entities": ["Blender", "Kaggle", "gpt-5.5", "claude-sonnet-5", "gemini-3.8-flash", "grok-4.20-reasoning", "deepseek-r1", "qwen3-coder-480b"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ask-ten-llms-for-a-blender-5-0-script-then-run-it-in-blender-5-0", "markdown": "https://wpnews.pro/news/ask-ten-llms-for-a-blender-5-0-script-then-run-it-in-blender-5-0.md", "text": "https://wpnews.pro/news/ask-ten-llms-for-a-blender-5-0-script-then-run-it-in-blender-5-0.txt", "jsonld": "https://wpnews.pro/news/ask-ten-llms-for-a-blender-5-0-script-then-run-it-in-blender-5-0.jsonld"}}