This is a submission for the Kaggle Benchmarking Challenge
Ask an LLM for a Blender Python script and it will usually give you one that looks right. Whether it runs depends on which Blender you paste it into. The bpy API changes on every major release: enums get renamed, attributes are removed, node sockets move, whole subsystems (compositor, sequencer, animation) get restructured. Most of the Blender code a model has seen was written for 2.8 to 3.x. Ask for 5.0 and you often get a 3.x script that dies on the first changed line.
Text similarity cannot see this. mesh.use_auto_smooth = True reads perfectly and raises AttributeError in 4.1 and later. The only judge that counts is the Blender version you asked for. So the benchmark runs every answer in that exact version.
Two questions per answer:
WATCH OUT block naming every API that changed for that version and its replacement, or - none.
These come apart in interesting ways, which is most of what follows.
30 scripting tasks, each asked for Blender 3.6, 4.2, 4.5 and 5.0 (119 prompts per model; one task has no 3.6 form). The tasks cover 33 verified API changes: EEVEE engine ids, the boolean solver enum, auto-smooth and custom normals, calc_normals, Principled BSDF socket names, the OBJ exporter, node group sockets, bone collections, slotted actions, the compositor node tree moving off Scene, the sequencer's sequences becoming strips, and so on. Two control tasks use APIs that did not change, to check the grader does not punish ordinary correct code.
Every model gets the same system prompt, temperature 0, no tools, no retrieval. The Kaggle task ships the four Blender Linux builds in a dataset, extracts them at the start of the run, and grades each answer by launching blender -b --factory-startup --python on the target build. A hand-written reference answer per task and version passes its assert in all four builds (119/119), which is the evidence that every prompt is answerable.
Ten models through the Kaggle Model Proxy, chosen to span vendors, sizes, open and closed weights, and reasoning styles, within a daily proxy budget of a few dollars:
| model | why |
|---|---|
| gpt-5.5, claude-sonnet-5, gemini-3.1-pro-preview | current frontier from three vendors |
| gemini-3.7-flash, gemini-3.8-flash, claude-haiku-4-5 | the small and cheap tier people actually script with |
| grok-4.20-0309-reasoning | fourth vendor, reasoning model (grok-4.6 is listed by the proxy but returned 404 on every call) | | deepseek-r1-0528 | open-weight reasoning model that returns its thinking inline |
| qwen3-coder-480b-a35b-instruct | open-weight model tuned for code |
| gpt-oss-120b | open-weight OpenAI model |
Total proxy spend for the ten leaderboard runs: $11.23. gpt-5.5 alone was $4.44; qwen3-coder was $0.03.
1190 graded answers. Aware is re-graded offline with the current parser so every model is judged the same way (more on that under methodology).
| model | runs | aware | runs but unaware |
|---|---|---|---|
| gpt-5.5 | 95% | 71% | 26% |
| claude-sonnet-5 | 94% | 53% | 42% |
| gemini-3.8-flash | 92% | 77% | 17% |
| gemini-3.1-pro-preview | 92% | 76% | 18% |
| gemini-3.7-flash | 89% | 73% | 18% |
| grok-4.20-reasoning | 82% | 29% | 57% |
| claude-haiku-4-5 | 72% | 31% | 48% |
| qwen3-coder-480b | 70% | 29% | 46% |
| deepseek-r1 | 67% | 39% | 33% |
| gpt-oss-120b | 61% | 29% | 37% | Runs by target version:
| model | 3.6 | 4.2 | 4.5 | 5.0 |
|---|---|---|---|---|
| gpt-5.5 | 97% | 97% | 100% | 87% |
| claude-sonnet-5 | 100% | 97% | 100% | 80% |
| gemini-3.8-flash | 97% | 90% | 97% | 87% |
| gemini-3.1-pro-preview | 97% | 90% | 97% | 83% |
| gemini-3.7-flash | 93% | 90% | 90% | 83% |
| grok-4.20-reasoning | 86% | 83% | 87% | 70% |
| claude-haiku-4-5 | 79% | 77% | 73% | 60% |
| qwen3-coder-480b | 83% | 67% | 70% | 60% |
| deepseek-r1 | 86% | 70% | 63% | 50% |
| gpt-oss-120b | 93% | 53% | 57% | 43% |
Pooled over all ten models: 91% of scripts run on 3.6, 81% on 4.2, 83% on 4.5, 70% on 5.0. Awareness falls faster: 81%, 48%, 40%, 35%.
Three categories go to 0% on 5.0 across all ten models: the compositor (Scene.node_tree is gone; the compositor is now a node group on scene.compositing_node_group), the boolean solver ('FAST' is now 'FLOAT', and 'MANIFOLD' is new), and the sequencer (SequenceEditor.sequences is now strips, and new_effect() takes length= instead of frame_end=). Animation is at 10%: Action.fcurves moved under slotted actions (action.layers[0].strips[0].channelbag(slot).fcurves). Not one model knew any of these. The training data has not caught up with 5.0 and no amount of scale fixes that. The best model on 5.0 (gpt-5.5 and gemini-3.8-flash, 87%) is still exactly as lost on these five as gpt-oss-120b.
The failures on 4.2 and later are different in character. They are changes from 4.0 to 4.2 that the models have partly absorbed (counts pooled over 4.2, 4.5 and 5.0): use_auto_smooth removed (19 failures), the legacy OBJ exporter bpy.ops.export_scene.obj removed (15), Mesh.calc_normals removed (14), BLENDER_EEVEE renamed (14), Principled BSDF Specular renamed to Specular IOR Level (9). The frontier models mostly get these right; the open-weight models mostly do not. gpt-oss-120b is the clearest case: 93% on 3.6, 53% on 4.2. It knows Blender 3.x well and stopped there.
The render engine id was BLENDER_EEVEE in 3.6, became BLENDER_EEVEE_NEXT in 4.2 when EEVEE Next shipped, and went back to BLENDER_EEVEE in 5.0 when the legacy engine was deleted. Every model failed this task somewhere:
| camp | models | 3.6 | 4.2 | 4.5 | 5.0 |
|---|---|---|---|---|---|
| learned the 4.2 rename | gpt-5.5, claude-sonnet-5, claude-haiku-4-5 | ok | ok | ok | fail |
| never learned it | both gemini flash models, gemini-3.1-pro, grok, deepseek, qwen, gpt-oss | ok | fail | fail | ok |
The second camp is right on 5.0 by accident: they write BLENDER_EEVEE for every version. The first camp is wrong on 5.0 with full confidence. gpt-5.5's WATCH OUT for the 5.0 prompt reads:
scene.render.engine = 'BLENDER_EEVEE': changed in Blender 4.2 when legacy EEVEE was replaced by EEVEE Next; use scene.render.engine = 'BLENDER_EEVEE_NEXT' in Blender 5.0.
Claude Sonnet 5's answer for 5.0 is the most instructive failure in the whole set. Its WATCH OUT names both identifiers and says the old one "no longer refers to the current EEVEE renderer in 5.0", so the aware check passes. The script then sets BLENDER_EEVEE_NEXT and fails. It knows there is a story here and tells the wrong half of it.
The gap column in the first table is the share of answers where the script ran but the WATCH OUT block missed the change. It is the number I find most useful, because it separates two very different kinds of model.
gemini-3.8-flash and gpt-5.5 have small gaps (17% and 26%): when their code works, they can usually say why the old code would not. Claude Sonnet 5 runs 94% of the time but names the change in only 53%; grok-4.20 runs 82% and names it in 29%. Both write - none under WATCH OUT for most prompts: Sonnet in 85 of 119 answers, grok in 110. qwen3-coder and Haiku do the same (115 and 109), and it is why the whole bottom half of the table sits at 29% to 39% aware. For every 2.80-era change (scene.objects.link becoming collection.objects.link, lamps becoming lights, dupli becoming instance) they silently write the modern form and declare that nothing changed. That is a judgement about what counts as "changed", not a parse failure, and I decided not to special-case it: the prompt asked for changes relative to older Blender, and 3.6 is old enough that 2.80 changes are still the ones people trip over in old tutorials.
The reverse gap, "aware but breaks", is small everywhere (1% to 7%). 41 answers named the change and still emitted broken code. Sonnet's EEVEE answer above is one. Three of Haiku's eight are the OBJ exporter: it knows export_scene.obj became wm.obj_export, then passes the old use_selection= keyword to the new operator.
On 3.6, awareness is 69% to 97%: everybody knows the 2.80 changes. By 4.2 the field splits. The three Gemini models and gpt-5.5 stay between 60% and 80% through 5.0, Sonnet sits at 43% to 53%, and everyone else drops below 30% by 4.5 and stays there. The scripts of that last group still run 57% to 87% of the time on 4.5 because most tasks touch APIs that did not change in that version. They are right without knowing why, and that is exactly the situation where a user cannot tell a stale answer from a current one.
The same model, prompt and settings, run four times through the proxy (gemini-3.7-flash across task versions): 107, 107, 103 and 106 of 119 scripts ran. Zero of the 119 answers were byte-identical between two runs. That is 3 to 4 points of noise on top of the sampling interval, so gpt-5.5 at 95%, Sonnet at 94% and the two Gemini models at 92% are a tie. Sonnet and gemini-3.8-flash actually swapped places between two task versions. The differences that survive the noise are the tiers: the five models above 89%, grok alone at 82%, and the four between 61% and 72%.
Three grading problems showed up in real runs. Each one would have produced a plausible-looking leaderboard that was wrong.
<think>...</think>, and in 29 of 119 answers the thinking contained a draft code fence. The first notebook version ran the draft. Fixed by stripping the block before parsing; deepseek's run rate moved from 62% to 67% and its awareness came down, because the rehearsed WATCH OUT inside the thinking had inflated it. The Kaggle leaderboard still shows 62% for deepseek because that is what the notebook version that ran it computed; the tables here are re-graded offline from the stored answers, and the changed scripts were re-run in the same Blender builds.max_tokens through the proxy. gemini-3.8-flash spends 12 to 16 thousand tokens thinking about auto-smooth, hit an 8192 cap once, and its "SyntaxError: unterminated string literal" was a truncation, not a model error. The task now raises the cap to 16384 for reasoning models and re-asks once at 32768 if an answer lands within 32 tokens of the cap. One answer needed that retry.
Two limits worth stating. The aware check matches identifiers by substring and quoted enum values exactly, so it is a floor on awareness, not a proof of understanding. And 30 tasks is enough to show the shape of the drift, not to rank models that are three points apart.
The case bank and the WATCH OUT contract come from bpy-compass, a version-aware Blender assistant I built for the Sanity challenge. Its evaluation compared one model with and without a retrieval layer over the release notes. The question this benchmark asks is the one I wanted answered before building that: how bad is the drift without retrieval, for which models, and where exactly. The answer is that on 5.0 it is bad for everyone, and that retrieval over release notes is not a nicety for Blender scripting, it is the difference between the compositor working and not.