CreateGame.ai

Reports

Best AI for Coding Games: Claude vs GPT vs Gemini and the Open Models

Picking the best AI for coding games in October 2026 is less about one winner than about which benchmark you trust and what you plan to build. We reviewed the published leaderboards, papers and public build write-ups to work out which models write game code that runs, which are cheap enough to iterate with, and which open-weights options are worth putting on your own machine.

Key takeaways

  • On the one leaderboard where people vote on generated web apps (LMArena WebDev), Anthropic's Claude Opus 5.5 leads with 1813 Elo, ahead of OpenAI's GPT-6 Astra (1786) and Claude Sonnet 5.5 (1774). The top five are all Claude or GPT.
  • On Artificial Analysis's Coding Agents Index, Claude Sonnet 5.5 inside Claude Code scores highest (0.684), but the index measures model plus harness, and the gap to the next five rows is small.
  • Cost per task varies by an order of magnitude among near-equal scores. GPT-6.1 Sol in Codex reached 0.629 at a reported $1.04 per task, against $14.19 for Sonnet 5.5 in Claude Code at 0.684.
  • The best open-weights model on WebDev, Kimi K3 (1654), trails the leader by about 160 Elo but is far weaker on terminal-style agent work. Small open models like MiMo-V2.6-Flash get surprisingly close for web code at a fraction of the price.
  • Nobody has published a clean, current, per-model benchmark for Unity C#, Godot GDScript or shaders. The one game-engine benchmark we found (GameDevBench) is Godot-only, and its headline result is that engine work is still hard: the best agent solved 53.8% of tasks.
  • For most people the practical answer is a $20 to $100 subscription to one of the three big agent products, plus a cheap model on the API for repetitive work.

What "coding games" actually asks of a model

Game code is an awkward test case, and that is why we think it deserves its own post. A to-do app has a spec you can verify with unit tests. A game has a spec that includes feel: whether a jump arc is satisfying, whether enemies are fair, whether the frame rate holds at 60. Models are good at the first kind of requirement and mixed on the second.

It helps to split the work into five jobs, because the models rank differently on each.

Browser games (HTML5 canvas, Three.js, Phaser, WebGL) are the easiest target. Training data is abundant, there is no engine to configure, and the model can often run its own output in a headless browser. This is where the evidence is strongest, and it is the use case LMArena's WebDev board measures most directly.

Game logic (state machines, inventory, combat maths, procedural generation, save systems) is ordinary software engineering with unusual constraints. Agentic coding benchmarks are the best proxy.

Engine scripting in Unity C# or Godot GDScript adds an engine the model cannot always see. The scene tree, inspector-assigned references and asset import settings live outside the text of the code. Models do well on pure script files and worse when the bug is in the scene.

Shaders (GLSL, HLSL, WGSL, Godot's shading language) are a low-resource language where the output is an image. A shader can compile and still look wrong.

Tooling and glue (build scripts, asset pipelines, editor extensions) looks like the Terminal-Bench style of work, where an agent runs commands and recovers from errors.

No single score covers all five. We use four different instruments below and say what each one can and cannot tell you.

The three scoreboards that matter

Artificial Analysis Intelligence Index: the general baseline

Artificial Analysis (AA) publishes a composite Intelligence Index. The current version, v4.3.2 as of 9 October 2026, averages ten evaluations, including Terminal-Bench 4.0, SciCode, GDPval-AA and Humanity's Last Exam [1][2]. It is not a coding index, and AA no longer publishes a standalone one in the data we read. But it is a useful sanity check on general capability, and it shows the shape of the market.

At the top, Claude Opus 5.5 scores 57.6, Claude Sonnet 5.5 scores 56.0, Claude Fable 5.1 scores 53.4, GPT-6 Astra 52.7, Gemini 4 Argon 52.6 and GPT-6.1 Sol 51.8 [1]. Among models flagged open-weights, MiMo-V2.6-Pro leads at 46.3, followed by GLM-5.3 (44.8) and Kimi K3 (43.6) [1]. Two warnings apply. AA rescales the index between versions: cached search snippets showed GPT-6 Astra between 55 and 60 while the live data we parsed gives 52.7, so any number you read elsewhere may use a different version. And the index folds in plenty that has nothing to do with games, such as scientific reasoning and long-document recall.

Coding Agents Index: closest to "can it build my project"

The more relevant AA product is the Coding Agents Index. It is the mean of three evaluations (DeepSWE v1.1, SWE-Atlas QnA and Terminal-Bench v4), each run with the model inside a coding agent [3]. DeepSWE tests multi-file software engineering, SWE-Atlas QnA tests answering questions about a codebase, and Terminal-Bench tests command-line task completion.

The catch, and it is a large one, is that each row is an agent plus a model. Claude models run in Claude Code, GPT models in Codex, Gemini in Google's Antigravity CLI, Grok in Grok Build and Meta's Muse Spark in Muse Code [3]. There are only eight rows. So when Sonnet 5.5 outscores GPT-6.1 Sol, we cannot say how much belongs to the model and how much to the harness. We cannot say it about any pair that crosses vendors.

Coding Agents Index: model plus harnessMean of DeepSWE v1.1, SWE-Atlas QnA and Terminal-Bench v4. Each bar is a model run inside its own vendor's agent, so it is not a pure model score.
Coding Agents Index: model plus harness00.10.20.30.40.50.60.7AnthropicGoogleOpenAIxAIMetaSonnet 5.5 (max) (Claude Code)0.68Opus 5.5 (max) (Claude Code)0.66Gemini 4 Argon (Antigravity CLI)0.64GPT-6.1 Sol (xhigh) (Codex)0.63Fable 5.1 (max) (with fallback) (Claude Code)Fable 5.1 (max) (with fallback) (Claud…0.62GPT-6 Astra (max) (Codex)0.62Grok 4.7 (xhigh) (Grok Build)0.56Muse Spark 1.3 (max) (Muse Code)0.54Agent / model

Eight rows only. Harness differs per row (Claude Code, Codex, Antigravity CLI, Grok Build, Muse Code). Gemini 4 Argon is pre-release. Source: Artificial Analysis Coding Agents Index, accessed 2026-10-09.

The ordering is Claude Sonnet 5.5 (0.684), Claude Opus 5.5 (0.660), Gemini 4 Argon (0.638), GPT-6.1 Sol (0.629), Claude Fable 5.1 (0.622), GPT-6 Astra (0.616), Grok 4.7 (0.563) and Muse Spark 1.3 (0.543) [3]. The top six sit within 0.07 of each other. We would not read a meaningful ranking into that spread, especially since AA publishes no confidence interval alongside it in the data we saw.

Some structure inside the components is more interesting than the totals. Gemini 4 Argon has the highest DeepSWE score of the eight at 0.788, but a mid-table 0.561 on Terminal-Bench v4 [3]. Grok 4.7 and Muse Spark 1.3 are competitive on DeepSWE (0.726 and 0.717) and collapse on Terminal-Bench v4 (0.333 and 0.318) [3]. Put plainly: some models write good code in a repository and are poor at driving a terminal, which matters if your workflow is "run the build, read the error, fix it".

A caveat on Gemini 4 Argon: AA and LMArena both flag it pre-release, and Google announced it on 30 September with limited access [1][4]. Press reports say the introductory price of $2 and $10 per million tokens will rise; we have not read Google's own post and treat that as unconfirmed. Separately, one outlet reports a Google claim of 77.9% on DeepSWE v1.1, where AA measured 78.8%, a small gap that is probably a matter of setup but is a reminder that vendor and independent numbers differ [23].

Terminal-Bench: the harness effect, in numbers

Terminal-Bench is the benchmark most often cited for agentic coding, and it is also where the harness effect is easiest to see. The tbench.ai leaderboard for version 2.0, mirrored by Epoch AI, lists the best submission per model and the agent that produced it [5][6]. For Claude Opus 4.6 there are two entries: 79.8% with the ForgeCode agent and 69.9% with the Droid agent [5]. They may differ in settings beyond the harness, so we do not claim a clean ten-point harness effect, but a gap that size between submissions for one model family says the scaffolding matters at least as much as small differences between frontier models.

That leaderboard has not been refreshed for the newest models. Its top entries are from the GPT-5.5 era (84.7%), and AA now uses Terminal-Bench 4.0, a different test [5]. Do not compare numbers across Terminal-Bench 2.0, 2.1 and 4.0. On the AA version inside the Intelligence Index, the snapshot we use has Claude Sonnet 5.5 at 0.636, Opus 5.5 at 0.596, GPT-6 Astra at 0.591, Gemini 4 Argon at 0.571 and GPT-6.1 Sol at 0.561 [1]. Among open models, GLM-5.3 reaches 0.419 and Kimi K3 only 0.126 [1]. We will come back to why Kimi's number matters.

LMArena WebDev: where the votes are about the thing you want

LMArena's WebDev board is the closest thing to a direct measurement of "which model makes the better small web app or game". Users ask for a front-end build, see two outputs and vote. The page showed 852,934 votes across 142 models in a snapshot dated 8 October 2026 [4].

LMArena WebDev: people voting on generated web appsElo from pairwise votes on front-end builds (raw, no style control). Open-weights models are marked.
LMArena WebDev: people voting on generated web apps05001,0001,5002,000AnthropicOpenAIGoogleAlibabaMetaMoonshotxAIXiaomiZ.aiDeepSeekClaude Opus 5.51,813GPT-6 Astra1,786Claude Sonnet 5.51,774GPT-6.1 Sol1,755Claude Fable 5.11,744Gemini 4 Argon1,678Qwen3.8 Max (0902)1,674Muse Spark 1.31,657MKimi K3 (open)1,654Grok 4.71,639MiMo-V2.6-Flash (open)1,637Qwen3.8-Flash-Next (open)1,633MiMo-V2.6-Pro (open)1,629ZGLM-5.3 (open)1,622DDeepSeek V4.1 Flash (open)1,619ZGLM-5.3-Flash (open)1,609Qwen3.8 27B (open)1,593Model

Arena Elo is relative to this board only and must not be compared with AA scores or with LMArena's text Elo. Gemini 4 Argon is flagged preliminary. GLM-5.3 open status conflicts between sources. Source: LMArena WebDev leaderboard, accessed 2026-10-09.

The top of the table: Claude Opus 5.5 (1813), GPT-6 Astra (1786), Claude Sonnet 5.5 (1774), GPT-6.1 Sol (1755), Claude Fable 5.1 (1744) [4]. Claude variants take five of the top seven spots on the page. The WebDev board lists effort variants separately: Sonnet 5.5 at its high setting is 1715, well below the xhigh variant at 1774 [4]. That 59-point gap is a good reminder that "Sonnet 5.5" is a range of behaviours, not a single score.

What does 1813 versus 1774 mean? Arena confidence bounds for those two are about plus or minus 15 each [4]. Under the standard Elo convention, a 400-point gap means the stronger side wins ten votes in eleven; a 39-point gap implies roughly a 56% win rate. The best cheap models are further back. Claude Haiku 5.5 (1587) and GPT-6 Luna (1581) sit about 226 points below Opus 5.5, which implies the top model wins close to 79% of head-to-head votes against them [4]. That is a clear preference, but the majority of the time the two outputs are close enough that a voter picked the cheaper one 21% of the time.

Three weaknesses of WebDev as a games benchmark. First, voters judge a rendered page, so visual polish and apparent completeness are rewarded, and a game that looks right but plays badly can win. Second, prompts are often short and one-shot, while real game projects are many iterations. Third, "web dev" includes plenty of dashboards and landing pages, not just games. We treat it as strong evidence about browser prototypes and weaker evidence about anything else. Its Elo scale also cannot be compared with AA's index or with LMArena's text Elo [1].

Claude, GPT and Gemini compared

Claude: the web-prototype leader, and the expensive one at the top

Anthropic holds the top WebDev spot and has the highest Coding Agents Index. Pricing is the awkward part. Opus 5.5 costs $4 per million input tokens and $20 per million output (a 3:1 blend of $8), Sonnet 5.5 costs $2 and $10 (blend $4), and the new Haiku 5.5 costs $0.10 and $0.50 for prompts up to 100,000 tokens [16]. Claude Fable 5.1 sits above them at $10 and $50 and scores lower on both WebDev (1744) and the Coding Agents Index (0.622) than the cheaper Opus and Sonnet, so for game code we see no reason to pay for it [1][16].

The notable result is that Sonnet 5.5 beats the more expensive Opus 5.5 on the Coding Agents Index (0.684 versus 0.660) and on AA's Terminal-Bench 4.0 (0.636 versus 0.596) while costing half as much per token [1][3]. Opus keeps the WebDev lead (1813 versus 1774). We read that as Opus having an edge on one-shot visual builds and Sonnet being the better default for iterative agent work. Both are close enough that your prompt and harness settings will matter more.

Cost per task is where Claude Code's high scores come with a bill. AA reports a mean of $14.19 per Coding Agents task for Sonnet 5.5 and $13.04 for Opus 5.5 [3]. These are API-equivalent costs; if you use Claude Code through a subscription the plan caps usage rather than billing per task, which we cover below.

GPT: strongest value in the agent table

OpenAI's GPT-6 Astra is the number-two WebDev model (1786) and costs $10 and $50 per million tokens, the same as Fable [1]. GPT-6.1 Sol is the more interesting one for budgets. It lists at $2 and $10, matching Sonnet, scores 0.629 on the Coding Agents Index inside Codex, and, per AA, costs a mean of $1.04 per task [1][3]. That is 7% of Sonnet 5.5's reported cost for 92% of its index score.

We do not know why the cost difference is so large given matching list prices. The likeliest explanations are that Sol uses far fewer tokens per task, that caching works differently, or that the xhigh and max effort settings AA ran consume different amounts of reasoning. AA does not break it down in the data we read, and OpenAI's pricing page returned an HTTP 403 when we tried to verify prices directly, so the $2 and $10 figure comes from AA and LMArena, which agree [1][4]. If the number holds up in your own use, it is the single most relevant fact in this post for anyone paying per token.

A word on OpenAI's smaller models. GPT-6 Luna costs $0.10 and $0.50, scores 38.1 on AA's index and 1581 on WebDev, effectively matching Haiku 5.5 on price and WebDev [1]. GPT-5.6 Terra, an older mid-tier at $2 and $12, scores 1522 on WebDev, well behind both [1].

Gemini: a strong coder with a launch caveat

Google's current flagship, Gemini 4 Argon, is the hardest to assess. It ranks ninth on WebDev at 1678 but is flagged preliminary, and AA has no speed data for it [1][4]. In the agent table it takes third place with 0.638 at $5.84 per task, between Claude and GPT on cost [3]. Its DeepSWE score of 0.788 is the highest we saw [3]. We found no public build write-up using Argon specifically; the first Gemini game-building material we found is from the Gemini 3 era, such as Google's own codelab for a match-3 game built with Antigravity [22]. Treat Argon's numbers as promising and provisional. Access was limited at launch.

The cheaper Gemini 3.8 Flash ($0.75 and $3.75 per million tokens through 2026, rising to $1.50 and $7.50 from 2027 according to Google's pricing page) scores 40.9 on AA's index [1]. It has no WebDev entry in the snapshot, so we cannot place it against Haiku or Luna for web games.

What the other labs offer

Grok 4.7 and Meta's Muse Spark 1.3 trail in the agent table (0.563 and 0.543) and are weak on Terminal-Bench v4, so we would not make either a first choice for game tooling [3]. Alibaba's closed Qwen3.8 Max (1674 on WebDev, $2 and $6) is the best non-Anthropic, non-OpenAI, non-Google model for web code, though preliminary [1][4].

What a coding-agent task costs vs how well it scoresAA mean cost per task (USD) against Coding Agents Index. Up and left is better.
What a coding-agent task costs vs how well it scores0.540.560.580.60.620.640.660.680.70246810121416Coding Agents Index (0-1)Mean cost per task (USD)Sonnet 5.5 (max) (Claude Code)Opus 5.5 (max) (Claude Code)Gemini 4 Argon (Antigravity CLI)GPT-6.1 Sol (xhigh) (Codex)Fable 5.1 (max) (with fallback) (Claude Code)GPT-6 Astra (max) (Codex)Grok 4.7 (xhigh) (Grok Build)Muse Spark 1.3 (max) (Muse Code)AnthropicGoogleOpenAIxAIMeta

Cost depends on tokens used per task, caching and effort setting as well as list price; AA does not break this down in the data we read. Costs are API-equivalent, not subscription prices. Source: Artificial Analysis Coding Agents Index, accessed 2026-10-09.

What people have actually built

Benchmarks only go so far, so we looked at public write-ups and talks. We did not run these builds ourselves, and most are self-reported. Where a source is promotional we say so.

Claude of Duty (Matt Shumer, July 2026). The most prominent recent example is a browser first-person shooter built with Claude Opus 5 in Claude Code. According to Decrypt's report, Shumer used a three-paragraph prompt asking for quality comparable to a recent Call of Duty, told the model to fan out subagents, and had separate critic agents loop on each piece. The result runs on Three.js and plain WebGL2, with about 55,000 lines across 11 subsystems and all assets generated procedurally at load time [8]. Skeptics questioned whether there was hidden manual coding, and Shumer published the prompt and code. Decrypt also notes the obvious contamination worry: Three.js pointer-lock controls and many shooter tutorials are in training data, and none of the reruns we read about published a contamination check [8]. A rerun by another developer, according to the same article, took about ten hours and 1.3 million tokens [8]. We take it as proof that a strong agent can assemble a large, playable browser game in a well-documented genre, not as evidence about novel game design.

Eight payment-themed games (Dodo Payments). A company engineering blog describes building eight browser games for its open-source game collection in a single session with a multi-agent setup, routing Claude to orchestration and GPT to deeper reasoning [11]. The post is promotional, and the author says the games had working loops, scoring and responsive design while feel, such as control tightness and difficulty pacing, still needed human tuning. It gives no cost figure for the project [11]. That human-tuning caveat shows up in almost every account we read.

Manager 11 and Era Online (Code with Claude 2026). Two conference talks describe larger solo projects: a football management game finished by its creator after years of attempts, and the revival of a 1999 Visual Basic MMORPG [12][13]. We only had the session pages, which are abstracts without cost or time figures, so we cannot say how much code Claude wrote or which models were used. They are evidence that people are shipping real projects with agents, not evidence of how well any model performs.

Codex in Godot (OpenAI community forum, July 2025). An older but instructive thread: a developer building a Godot 4.4 game with Codex CLI said it was the tool whose work they trusted most, but it could not look at screenshots at the time, so they used Claude and Gemini for visual checks of scenes; they also reported that a rules file standardising tabs over spaces made output far more consistent, and that large tasks needed splitting to avoid hallucinated code [9]. That is a year old, and image input in Codex may well have changed since. We include it because the two practical lessons, give the agent a style file and break up big tasks, recur in every other source.

Engine-specific integrations. A 2026 guide to Codex CLI for game teams argues it works well for Unity and Godot when connected to engine MCP servers (tool bridges that let an agent inspect and edit a running editor) and given project guidance in an AGENTS.md file [10]. A vendor comparison claims Cursor edits Godot scripts well but cannot see the scene tree or run scenes; the vendor sells a competing product, so weigh that accordingly [24]. The consistent finding is that access to the running game decides the result as much as the base model.

Engine work and shaders: the weakest evidence

If you build in Unity or Godot, the most honest answer is that the benchmarking has not caught up. We found one purpose-built benchmark and one on shaders.

GameDevBench (Carnegie Mellon and Princeton, arXiv 2602.11103) tests agents on 333 Godot tasks drawn from tutorials. In the camera-ready version, the best agent and method solved 53.8% of tasks. Success fell from 51.4% on gameplay tasks to 33.0% on 2D graphics tasks, and adding image and video feedback raised GPT-5.4 from 41.1% to 52.0% [7]. Two readings follow. Engine work is still hard for agents even though gameplay logic is tractable. And letting the agent see the game matters about as much as which model you choose. The paper uses models from the early 2026 generation, so the absolute numbers will be higher for Opus 5.5 or GPT-6.1 Sol, but we have no published result to show by how much. Third-party leaderboards for GameDevBench exist; we did not verify them.

For shaders, the only systematic benchmark we found is ShaderMatch, presented at the 2025 LLM4Code workshop. It takes real Shadertoy fragment shaders and asks a model to complete a function from a header and comments, then compares the rendered frame. Even the best of more than 20 open code models failed to produce working code 31% of the time, which the authors attribute to GLSL being rare in training data [14][15]. This is old, covers open models only, and tests function completion rather than whole effects. A current frontier model will do better, and we would still expect shaders to be the area where "compiles, looks wrong" is most common. Our practical advice is to treat shader output as a first draft, render it and iterate with a screenshot.

A vendor blog also warns that Godot 4's syntax changes (annotations, await replacing yield) trip up models trained on Godot 3 material, and recommends pinning the engine version in your project guidance [24]. It is a plausible and cheap precaution, so we pass it on; we have not seen a measurement of how often it happens with current models.

The open-weights options

Open-weights models matter to game developers for three reasons: they can be run locally (no per-token bill, no data leaving your machine), they can be hosted by many providers at low prices, and some licences permit commercial use without conditions. They also trail the closed frontier. By our reading of AA's data, the best open model, MiMo-V2.6-Pro at 46.3, is 11.3 index points behind Opus 5.5, and by Epoch's Capabilities Index Kimi K3 is about 9.9 points behind [1][5].

Open-weights models: WebDev Elo vs output priceHosted API output price per 1M tokens against LMArena WebDev Elo. Closed reference points included for scale.
Open-weights models: WebDev Elo vs output price$1,400$1,500$1,600$1,700$1,800$1,90005101520WebDev EloOutput price (USD per 1M tokens)Claude Opus 5.5Claude Sonnet 5.5GPT-6.1 SolKimi K3GLM-5.3Qwen3.8 27BDeepSeek V4 Pro (0813)MiniMax-M3InklingMoonshotZ.aiXiaomiDeepSeekAlibabaMiniMaxAnthropicOpenAIOpen weightsClosed reference

Prices are hosted API list prices from AA (not vendor-verified except Anthropic and Google). GLM-5.3 open status conflicts between sources. Licences vary (MIT, Apache 2.0, custom). Source: Artificial Analysis (via CreateGame.ai snapshot), accessed 2026-10-09.

The most useful table for a game coder is the one below. WebDev Elo is a sensible proxy for browser-game output, and Terminal-Bench 4.0 is a proxy for agent robustness. All figures are from the snapshot at 9 October 2026 [1][4]. Prices are hosted API prices per million tokens as listed by AA; self-hosting is a different calculation.

Model License WebDev Elo AA index Terminal-Bench 4.0 (AA) Hosted $ in / out
Kimi K3 Kimi K3 license 1654 43.6 0.126 3.00 / 15.00
MiMo-V2.6-Flash MIT 1637 37.9 0.227 0.14 / 0.28
Qwen3.8-Flash-Next qwen-community-1.0 1633 39.8 0.253 0.15 / 0.47
MiMo-V2.6-Pro MIT 1629 46.3 0.348 0.435 / 0.87
GLM-5.3 MIT (source conflict) 1622 44.8 0.419 1.40 / 4.40
DeepSeek V4.1 Flash MIT 1619 39.5 0.268 0.30 / 1.20
Qwen3.8 27B Apache 2.0 1593 33.7 0.056 0.50 / 3.00

Three observations. First, on web output the open and closed models are not far apart in the cheap tier: MiMo-V2.6-Flash at 1637 beats Claude Haiku 5.5 (1587) and GPT-6 Luna (1581) on WebDev at a lower output price ($0.28 versus $0.50) [1][4]. Second, the open leader on web output, Kimi K3, is poor on Terminal-Bench 4.0 at 0.126, a gap that would show up as an agent that writes plausible code but struggles to drive a build [1]. Third, GLM-5.3 has the best Terminal-Bench figure of the open group (0.419), close to the closed Qwen3.8 Max at 0.389 [1].

On the open status of GLM-5.3: AA and LMArena list it as open weights under MIT, while Epoch's capabilities data labels it "Closed weights" [1][5]. We could not settle that conflict, so check the Hugging Face release before building a pipeline around it.

Running a model locally

The reference local option is Qwen3.8 27B (Apache 2.0, dense). Community guides put the 4-bit quantised model at about 16 to 18 GB, which fits on a 24 GB card, and a 128K-context setup near 20 GB of VRAM [21]. Reported speed on an RTX 5090 is around 142 tokens per second in a single DataCamp tutorial using NVFP4 weights and speculative decoding, with lower figures for other setups [21]. We would not plan around that; independent testers on other hardware reported 15 to 30 tokens per second [21].

How good is it as a coder? Qwen3.8 27B scores 1593 on WebDev and 33.7 on AA's index, comparable to a Haiku-class model rather than a frontier one [1]. For autocomplete, refactors and boilerplate, that is usable. For driving an agent through a multi-hour build it is the wrong tool: its Terminal-Bench 4.0 score is 0.056 [1]. A local model is a privacy and cost choice, not a quality one. We cover the economics in our cost breakdown.

OpenAI's gpt-oss-120b (Apache 2.0, August 2025) and Meta's Llama 4 Maverick (April 2025) are still the newest in their families on AA, scoring 11.6 and 10.0 [1]. They are included for reference; we would not start a new project on either.

Practical recommendations

All recommendations come with the same disclaimer: they are based on published benchmarks and write-ups, not our own head-to-head tests, and the leaderboards are days or weeks old at the time of writing. Several of the models above were released in the past six weeks.

By budget

About $20 a month. Every major vendor sells an entry plan: Claude Pro at $20 monthly ($17 on annual billing), ChatGPT Plus at $20 with Codex included, and Cursor's $20 individual tier [17][18][20]. Claude Code is included in Pro and Max per Anthropic's pricing page; one third-party article we found claimed Anthropic pulled Claude Code from Pro in April 2026, which contradicts the page we read on 9 October, and we trust the page [17]. Limits are the constraint. Anthropic describes five-hour session windows plus weekly limits that depend on conversation length and model; OpenAI publishes estimates for Plus, for example 15 to 160 local messages per five hours on GPT-6.1 Sol [17][18]. A small browser game or game-jam entry is feasible in this tier.

$100 to $200 a month. Claude Max starts at $100 (5x and 20x Pro usage), ChatGPT Pro has $100, $200 and $500 tiers, and Google AI Ultra has $100 and $200 tiers according to press coverage of Google I/O 2026 [17][18][19]. Cursor's higher tiers are reported at $60 and $200 by third parties; Cursor's own page did not show those prices when we read it [20]. This is where agentic work with large contexts becomes comfortable. We cannot compute a break-even between plans and API usage because the plans publish no token allowances, which is one of the larger unknowns in this market.

Pay per token. If you want to know exactly what a session costs, use the API. Our worked estimate in the cost article puts a heavy weekend session at about $11.70 on Sonnet 5.5 and $11.20 on GPT-6.1 Sol, using stated token assumptions, and under $1.50 on cheap models like DeepSeek V4.1 Flash. AA's reported mean Coding Agents task cost ($1.04 to $14.19) is consistent with that range.

Local and open. Run Qwen3.8 27B or another open model for private, unmetered completions, and keep a hosted frontier model for the hard problems. This hybrid is the arrangement most of the write-ups we read converge on.

By use case

If you are building Start with Why Cheaper fallback
A browser game or jam prototype Claude Opus 5.5 or Sonnet 5.5 Top WebDev Elo (1813, 1774); 1M context GPT-6.1 Sol, then MiMo-V2.6-Flash
A large multi-system web game Sonnet 5.5 in Claude Code, or GPT-6.1 Sol in Codex Highest Coding Agents Index; Sol cheapest per task Gemini 4 Argon once access opens
Godot GDScript Any frontier agent with screenshot feedback GameDevBench: visual feedback adds about 11 points Qwen3.8 27B for local edits
Unity C# Codex or Claude Code with a Unity MCP server Tool access to the editor matters most [10] Cursor for pure script editing
Shaders Frontier model, render and iterate Low-resource language; ShaderMatch shows high failure rates [15] Re-prompt with the failed frame
Tool scripts and pipelines Sonnet 5.5 or GPT-6.1 Sol Terminal-Bench strength GLM-5.3 (open)
High-volume repetitive edits Haiku 5.5, GPT-6 Luna, MiMo-V2.6-Flash $0.10 to $0.28 per million input or output Local Qwen3.8 27B

Habits that matter more than the model

Across the sources, five habits repeat. Write a project guidance file (AGENTS.md, CLAUDE.md or GEMINI.md) with the engine version, folder layout, naming and test commands. Give the agent a way to run the game and look at the result; GameDevBench's 11-point gain from visual feedback is the cleanest evidence we have for this [7]. Split big tasks, as both the forum thread and the Dodo write-up recommend [9][11]. Use a second model or fresh agent as reviewer, which is the core of Shumer's loop [8]. And keep a human in charge of feel: control tightness, difficulty curves and pacing are what every account says still needs a person.

Where this is heading

Frontier coding models have clustered: six models land within 0.07 of each other in the agent table, and four vendors hold the top ten on WebDev. Prices at a given capability are falling fast (our analysis of how fast AI is improving covers the rates), but the non-code half of game development is not improving as visibly, as GameDevBench's graphics numbers show. Engine vendors may address it first; see our look at Unity and Unreal's AI tools.

If you want to see what whole-game generation looks like when the model also handles worlds and story rather than only code, our comparison of AI game makers covers the products. We are building CreateGame.ai on the premise that one sentence should be enough to get a playable browser game, and we are honest that today's evidence says the code is the easier half. For the broader model picture across writing, art and dialogue, start with the pillar guide to AI models for game development.

How we reviewed this

We reviewed published sources only. We did not run any of these models or agents ourselves and make no claim of hands-on testing.

FAQ

What is the best AI for coding games right now? For one-shot browser games, Claude Opus 5.5 leads LMArena's WebDev board (1813 Elo). For agentic work on larger projects, Claude Sonnet 5.5 in Claude Code leads AA's Coding Agents Index (0.684), with GPT-6.1 Sol in Codex close behind at a much lower reported cost per task.

Is Claude or GPT better for game development? Claude holds more of the top WebDev spots, GPT-6 Astra is second there, and the Coding Agents Index puts both in the same band. The harness differs for each (Claude Code versus Codex), so the data cannot separate model from tool. Try both on your own project.

Can AI write Unity C# and Godot GDScript? Yes for scripts, with caveats for the scene. The only dedicated benchmark we found, GameDevBench for Godot, shows the best agent solving 53.8% of tasks and graphics tasks much harder than logic. Giving the agent screenshots or video of the running game helped substantially.

What is the best free or open-source AI for coding games? Among open-weights models, Kimi K3 leads WebDev (1654) and MiMo-V2.6-Pro leads AA's index among open models (46.3). Qwen3.8 27B (Apache 2.0) is the practical choice to run on a single consumer GPU, but it is a Haiku-class coder, not a frontier one.

Can AI write shaders? It can write them, but they are the hardest area. A 2025 benchmark found even strong open code models failed to produce working GLSL 31% of the time. Always render the result and feed the image back to the model.

How much does AI game coding cost? Between a $20 monthly subscription and a few dollars of API tokens per session for cheap models, up to roughly $10 to $25 for a heavy session on a frontier model. See our full cost breakdown.

If you want to see where this goes next, join the waitlist at https://creategame.ai.

Sources

  1. Artificial Analysis, model leaderboard and Intelligence Index v4.3.2, accessed 2026-10-09: https://artificialanalysis.ai/leaderboards/models
  2. Artificial Analysis, Claude Opus 5.5 model page (index version and component evals), accessed 2026-10-09: https://artificialanalysis.ai/models/claude-opus-5-5
  3. Artificial Analysis, Coding Agents Index, accessed 2026-10-09: https://artificialanalysis.ai/agents/coding-agents
  4. LMArena (arena.ai), WebDev leaderboard, page dated 2026-10-08, accessed 2026-10-09: https://arena.ai/leaderboard/code/webdev
  5. Epoch AI, benchmark data hub (ECI, SWE-bench Verified, Terminal-Bench 2.0 mirror), accessed 2026-10-09: https://epoch.ai/benchmarks
  6. Terminal-Bench 2.0 leaderboard (tbench.ai): https://www.tbench.ai/leaderboard/terminal-bench/2.0
  7. Chi et al., GameDevBench: Evaluating Agentic Capabilities Through Game Development, arXiv 2602.11103: https://arxiv.org/abs/2602.11103
  8. Decrypt, "Dumbest AI prompt: Claude beat careful game design", 28 July 2026: https://decrypt.co/374560/dumbest-ai-prompt-claude-beat-careful-game-design
  9. OpenAI Developer Community, "Codex CLI programming game in Godot", July 2025: https://community.openai.com/t/codex-cli-programming-game-in-godot/1316806
  10. "Codex CLI for Game Development Teams: Unity MCP, Godot MCP, and Agent-Driven Game Workflows", 27 April 2026: https://codex.danielvaughan.com/2026/04/27/codex-cli-game-development-unity-godot-mcp-agent-driven-workflows/
  11. Dodo Payments engineering blog, "AI agents build 8 games" (promotional): https://dodopayments.com/engineering/ai-agents-build-8-games
  12. Code with Claude 2026 session page, "Manager 11" (abstract only): https://claude.com/code-with-claude/session/tyo-ext-manager-11
  13. Code with Claude 2026 session page, "Era Online: Resurrecting a 1999 MMORPG with Claude Code" (abstract only): https://claude.com/it/code-with-claude/session/sf-ext-era-online-resurrecting-a-1999-mmorpg-with-claude-code
  14. ShaderMatch benchmark card and leaderboard (Vipitis), Hugging Face: https://huggingface.co/spaces/Vipitis/shadermatch
  15. Evaluating Language Models for Computer Graphics Code Completion (ShaderMatch), LLM4Code 2025: https://conf.researchr.org/details/icse-2025/llm4code-2025-papers/13/Evaluating-Language-Models-for-Computer-Graphics-Code-Completion
  16. Anthropic, API pricing, accessed 2026-10-09: https://platform.claude.com/docs/en/about-claude/pricing
  17. Claude pricing page (individual plans), accessed 2026-10-09: https://claude.com/pricing
  18. ChatGPT/Codex plans and usage limits, accessed 2026-10-09: https://learn.chatgpt.com/docs/pricing
  19. Google, "Everything new in our Google AI subscriptions, fresh from I/O 2026" (read via search summary): https://blog.google/products-and-platforms/products/google-one/google-ai-subscriptions
  20. Cursor pricing page, accessed 2026-10-09: https://cursor.com/pricing
  21. DataCamp, "How to run Qwen3.8-27B locally" and related community guides (read via search summary): https://www.datacamp.com/tutorial/how-to-run-qwen3-8-27b-locally
  22. Google Codelabs, "Build a Match 3 Arcade Game With Gemini and Antigravity": https://codelabs.developers.google.com/gemini-match3-golang
  23. TestingCatalog, "Google prepares Antigravity for Gemini 4 Argon and Concierge": https://testingcatalog.com/google-prepares-antigravity-for-gemini-4-argon-and-concierge
  24. Summer Engine, "Cursor + Godot plugin vs Summer Engine" (vendor-written) and Seeles, "AI coding agents for games" (vendor-written): https://www.summerengine.com/blog/cursor-plus-godot-vs-summer-engine and https://www.seeles.ai/resources/blogs/ai-game-coding-agents-for-games