Best AI Coding Assistant for Godot in 2026: 9 Compared GPT-6 Astra in Codex scored 68.8% pass@1 and Claude Fable 5 in Claude Code scored 67.3% pass@1 on 333 GameDevBench Godot 4 tasks, according to the GameDevBench leaderboard from Carnegie Mellon and Princeton researchers presented at ICML 2026. The overlapping 95% confidence intervals of ±5.0 for both entries mean the benchmark does not establish a clear winner among Codex, Claude Code and Muse Code, whose muse-spark-1.2 entry scored 61.0%. A Godot Community Poll 2026 found 63.2% of respondents mainly use Godot's built-in code editor, with prices and versions checked September 19, 2026. Best AI Coding Assistant for Godot in 2026: 9 Compared For terminal work, choose GPT-6 Astra in Codex or Claude Fable 5 in Claude Code: these model-agent pairs score 68.8% and 67.3% on 333 GameDevBench Godot 4 tasks https://waynechi.com/gamedevbench . For an assistant inside Godot, consider Ziva, which we make, or AI Assistant Hub. Prices and versions checked September 19, 2026. In the self-selected Godot Community Poll 2026, 63.2% of respondents https://godotengine.org/storage/data/GodotCommunityPoll2026.csv mainly use Godot’s built-in code editor. Cursor, Devin Desktop formerly Windsurf , Copilot’s official clients and JetBrains AI work outside that editor. For more plugins, see Best AI Tools for Godot in 2026 https://ziva.sh/blogs/best-ai-tools-for-godot-2026 . The scores come from GameDevBench, and the features and prices come from product documentation. We checked sample Godot 3 scripts with Godot 4.7.2’s command-line parser on Linux. We did not test assistants, local setups, project instructions or the prompts below. TL;DR | If you… | Pick | |---|---| | code in Godot’s built-in editor | Ziva https://ziva.sh/docs free Hobby plan https://ziva.sh/ pricing , AI Assistant Hub free, MIT https://github.com/FlamxGames/godot-ai-assistant-hub , or a terminal agent beside Godot | | want autocomplete in a separate editor | Copilot Free https://docs.github.com/en/copilot/get-started/plans , Devin Desktop Free https://devin.ai/pricing or Cursor Pro https://cursor.com/docs/models-and-pricing | | want multi-file agents | GPT-6 Astra in Codex or Claude Fable 5 in Claude Code; add a Godot MCP server, a bridge that lets agents use editor tools | | write C in Godot https://ziva.sh/blogs/gdscript-vs-csharp | Rider https://www.jetbrains.com/lp/rider-godot/ free for non-commercial use with Copilot https://blog.jetbrains.com/dotnet/2026/07/22/rider-2026-2-release/ needs a Copilot subscription or JetBrains AI | | want to run local models | AI Assistant Hub, Ziva https://ziva.sh/docs/local-providers or Claude Code through Ollama https://docs.ollama.com/integrations/claude-code | Ziva requires internet for AI features https://ziva.sh/docs/installation ; its local-model docs do not say whether those models work offline. For autocomplete inside Godot, a community Copilot plugin https://godotengine.org/asset-library/asset/4800 requires Node.js 20.8+ and a Copilot subscription, including Free. The official godot-tools https://github.com/godotengine/godot-vscode-plugin extension adds GDScript support to VS Code and is also on Open VSX https://open-vsx.org/extension/geequlim/godot-tools , an extension registry. Cursor installs extensions from Open VSX https://cursor.com/help/customization/extensions . What GameDevBench shows GameDevBench https://arxiv.org/html/2602.11103v2 , by Carnegie Mellon and Princeton researchers at ICML 2026, derives its tasks from Godot 4 tutorials. Its leaderboard https://raw.githubusercontent.com/waynchi/gamedevbench/main/results/leaderboard.csv mixes models, agent software and submission dates, so the scores cannot isolate the assistant’s contribution. These are the highest listed entries for each agent. Missing ranks are other entries from agents already listed, mostly Codex and Claude Code. The score is pass@1: the percentage of tasks solved on the first attempt. The 95% confidence interval shows the uncertainty around that score; high and xhigh are reasoning effort settings. | Rank | Agent | Model entry | pass@1, 95% CI | |---|---|---|---| | 1 | Codex | gpt-6-astra high | 68.8% ±5.0 | | 2 | Claude Code | claude-fable-5 xhigh | 67.3% ±5.0 | | 5 | Muse Code | muse-spark-1.2 high | 61.0% ±5.2 | | 7 | Kimi Code | kimi-k3 | 58.0% ±5.3 | | 10 | Gemini CLI | gemini-3-pro-preview | 53.8% ±5.4 | | 14 | OpenCode | glm-5.2 | 38.4% ±5.2 | | 16 | OpenHands | kimi-k2.5 | 20.7% ±4.4 | The overlapping confidence intervals do not establish a clear winner among Codex, Claude Code and Muse Code. Each entry uses its model’s best agent software and feedback setup, which can include editor screenshots and gameplay video. OpenHands reached 38.4% with GPT-5.4 Mini without visual feedback, above its 20.7% leaderboard entry with Kimi K2.5. Even the leader failed 104 of 333 tasks. Common failures involve scene structure, signals messages between objects and resources data objects used by nodes and scripts . Treat the rankings as a shortlist, then check the assistant against your own project. The benchmark excludes Copilot, Cursor, Devin Desktop, JetBrains AI and in-editor plugins, including Ziva and AI Assistant Hub. It uses Godot 4.4.1 https://github.com/waynchi/gamedevbench , while 63.8% of the community poll’s respondents mainly use 4.7. Disclosure: GPT-6 Astra wrote this post from a Claude draft, with research and fact-checking by Claude agents. GPT-6 Astra and Claude Fable 5 rank first and second. Can the assistant see your editor and game? The paper’s Table 2 compares visual feedback across eleven model-agent pairs. Eight scores rose with editor screenshots plus gameplay video; three fell. Most changes are small compared with the confidence intervals. GPT-5.4 showed a larger gain, from 41.1% to 52.0%. Researchers supplied screenshots through their own MCP server and instructions for recording gameplay video. The leaderboard instead keeps each model’s best feedback setup. | Agent | Model | Without visual feedback | Screenshots + video | |---|---|---|---| | Codex | GPT-5.4 | 41.1% | 52.0% | | Claude Code | Sonnet 4.5 | 28.8% | 34.8% | | Gemini CLI | Gemini 3 Pro Preview | 50.1% | 53.8% | | Claude Code | Haiku 4.5 | 13.8% | 16.5% | | Codex | GPT-5.4 Mini | 36.9% | 39.0% | | OpenHands | Kimi K2.5 | 18.9% | 20.7% | | OpenHands | Qwen 3.5 397B A17B | 5.4% | 5.1% | | Gemini CLI | Gemini 3 Flash Preview | 45.4% | 44.1% | | OpenHands | Haiku 4.5 | 15.6% | 17.7% | | OpenHands | GPT-5.4 Mini | 38.4% | 36.9% | | OpenHands | Gemini 3 Flash Preview | 30.3% | 31.8% | The paper tested its own server, not these public bridges. Per its docs, Coding-Solo/godot-mcp https://github.com/Coding-Solo/godot-mcp runs projects and captures debug output. The third-party open-source MCP plugin hi-godot/godot-ai https://github.com/hi-godot/godot-ai connects Claude Code and Codex to a live editor to edit scenes, nodes and scripts. Ziva’s MCP server https://ziva.sh/docs/ziva-mcp-server lets external agents inspect the scene tree the hierarchy of nodes , read errors and run games. AI Assistant Hub enables agent tools only with Ollama or llama.cpp. Its tools scan scenes and manage nodes https://store.godotengine.org/asset/flamxgames/ai-assistant-hub/ . Isaac Dedini https://vivecuervo7.github.io/dev-blog/p/claude-code-godot/ built his card-game UI entirely through Claude Code, opening the Godot editor only once. His custom test runner let Claude compare screenshots against the intended UI. Avoiding Godot 3 code in Godot 4 Developer reports vary. In March, Ariarule https://news.ycombinator.com/item?id=47407325 reported good GDScript from Opus 4.5, 4.6 and Sonnet 4.6 using CLAUDE.md . Ariarule still saw occasional Godot 3 output. In September, ieishi https://forum.godotengine.org/t/143691/52 reported mixed syntax despite requesting Godot 4, without naming the tool. Put your engine version and project conventions in a root AGENTS.md . Codex, Cursor and Ziva https://ziva.sh/docs/agents-md read it. For Claude Code, add a line containing @AGENTS.md to CLAUDE.md . This imports your shared instructions; see the memory docs https://code.claude.com/docs/en/memory for automatic loading rules. Gemini CLI uses GEMINI.md https://github.com/google-gemini/gemini-cli . Copilot chat in VS Code https://code.visualstudio.com/docs/agent-customization/custom-instructions reads AGENTS.md , but inline suggestions ignore it. Include rules specific to your game. For example, MrPhil’s Stellar Throne instructions https://www.mrphilgames.com/blog/claude-md-for-game-devs prohibit await in manager ready functions to prevent load-order bugs. Godot 4 changed classes, signals, Tween and the tool keyword https://docs.godotengine.org/en/stable/tutorials/migrating/upgrading to godot 4.html . It replaced yield, export and onready with await and annotations https://docs.godotengine.org/en/stable/tutorials/scripting/gdscript/gdscript basics.html . Use this as a reference when writing your AGENTS.md rules: | Godot 3 | Godot 4 | |---|---| | KinematicBody2D | CharacterBody2D | | KinematicBody | CharacterBody3D | | Spatial | Node3D | | yield ... | await | | export var | @export var | | onready var | @onready var | | instance | instantiate | | Tween node | create tween https://docs.godotengine.org/en/stable/classes/class node.html | | connect "sig", obj, "method" | sig.connect callable | | tool | @tool | A five-minute test 1. Ask for a player controller matching your engine version and naming conventions. Check for the old forms on the left above. 2. Use Godot’s parse-only check https://docs.godotengine.org/en/stable/tutorials/editor/command line tutorial.html : godot --headless --check-only --script