CreateGame.ai

Reports

Best AI for Game Development in 2026: We Reviewed the Models and Benchmarks

There is no single best AI for game development, and the leaderboards agree on that more than their headlines suggest: six models released in September 2026 sit within six points of each other, and the right pick depends on whether you are writing code, voicing a shopkeeper or drawing a sprite sheet. This is the hub of a 13-part series; below is where the market stands on 9 October 2026, which model we would reach for in each game-dev job, and how far to trust the numbers behind that.

Key takeaways

  • The top is crowded and brand new. On Artificial Analysis's Intelligence Index (v4.3.2, 9 Oct 2026), Claude Opus 5.5 leads at 57.6, followed by Claude Sonnet 5.5 (56.0), Claude Fable 5.1 (53.4), GPT-6 Astra (52.7), Gemini 4 Argon (52.6, pre-release) and GPT-6.1 Sol (51.8). All six shipped between 1 and 30 September.
  • For code, the evidence points at Claude and GPT. The top five on LMArena's WebDev board are Claude or GPT models (Opus 5.5 at 1813 Elo), and in AA's Coding Agents Index Sonnet 5.5 inside Claude Code leads at 0.684. That index measures model plus harness, so it is not a clean model ranking.
  • Price per token and price per task are different things. Opus 5.5 and Sonnet 5.5 list at $8 and $4 blended per million tokens, yet AA reports $5.98 and $5.46 to run its index, against $0.72 for GPT-6.1 Sol, which scores 51.8.
  • Open weights are about 11 index points back, and cheap. MiMo-V2.6-Pro (46.3) is the best open-weights model on AA's index, 80% of Opus 5.5's score at roughly one-fifteenth of its blended price. Epoch AI's index shows a similar gap of about 10 points.
  • Small models cover a lot of game work. Claude Haiku 5.5 scores 43.4 at $0.10/$0.50 per million tokens, and fast non-reasoning variants of cheap models start answering in under a second, which NPC dialogue needs.
  • No benchmark measures games. The ten evaluations behind AA's index cover work tasks, science and terminals, not level design or character voice. Our picks are reasoned from published evidence, not our own tests.

The state of the model market on 9 October 2026

Everything below comes from a shared data snapshot we built on 9 October 2026 from Artificial Analysis (AA), LMArena and Epoch AI, plus vendor pricing pages where we could read them. We cite the index and its version each time because AA rescales its index between versions: cached copies of its pages showed GPT-6 Astra between 55 and 60 while the live data we parsed gave 52.7, and we use the live figure [1][2]. For how those sources work and where they fail, jump to how the series uses benchmarks.

General intelligence: a tight top six, and Anthropic holds the first three places

AA's Intelligence Index v4.3.2 combines ten evaluations into one number [1][2]. Claude Opus 5.5 (released 22 September) scores 57.6. Claude Sonnet 5.5 is 1.6 points behind at 56.0, which is 97% of the leader's score. Then come Claude Fable 5.1 (53.4), GPT-6 Astra (52.7), Gemini 4 Argon (52.6) and GPT-6.1 Sol (51.8). Meta's Muse Spark 1.3 is seventh at 48.1.

Artificial Analysis Intelligence Index v4.3.2, top 16 modelsAll six leaders were released between 1 and 30 September 2026. Open-weights models in the list: MiMo-V2.6-Pro, GLM-5.3 (status disputed), Kimi K3, GLM-5.3-Flash.
Artificial Analysis Intelligence Index v4.3.2, top 16 models0102030405060AnthropicOpenAIGoogleMetaxAIXiaomiAlibabaZ.aiStepFunMoonshotClaude Opus 5.557.6Claude Sonnet 5.556Claude Fable 5.153.4GPT-6 Astra52.7Gemini 4 Argon52.6GPT-6.1 Sol51.8Muse Spark 1.348.1Grok 4.746.4MiMo-V2.6-Pro46.3Qwen3.8 Max (0902)45.4ZGLM-5.344.8SStep 5 (Preview)43.7MKimi K343.6Claude Haiku 5.543.4GPT-5.6 Terra42.1ZGLM-5.3-Flash41.8Model

Highest-effort variant of each model as AA lists it. Gemini 4 Argon is pre-release. AA rescales the index between versions, so cite the version with the number. Does not measure game-specific work. Source: Artificial Analysis, Intelligence Index v4.3.2, accessed 2026-10-09.

Two features of that list matter for anyone choosing a model. First, the whole top six was released within a single month (Fable 5.1 on 1 September, GPT-6 Astra on 3 September, Opus 5.5 on 22 September, Sonnet 5.5 on 28 September, GPT-6.1 Sol on 29 September, Gemini 4 Argon on 30 September) [1]. The leaderboard is weeks old, and today's leader may not lead in November. Second, Gemini 4 Argon is a pre-release: Google announced it on 30 September with limited access, and both AA and LMArena flag it as pre-release [1][4][5]. Its place on any chart is provisional.

Epoch AI publishes a separate composite, the Epoch Capabilities Index (ECI), built from different benchmarks and on a different scale. We treat it as a second opinion, never as the same axis. It also has Opus 5.5 first (167.3), but the order beneath it differs: GPT-6 Astra (166.5), GPT-6.1 Sol (166.1), Sonnet 5.5 (165.0) and Fable 5.1 (164.7) [6]. Five models fall within 2.6 ECI points. GPT-6 Astra is fourth on AA and second on Epoch, while Sonnet 5.5 is second on AA and fourth on Epoch. Gemini 4 Argon is not yet in Epoch's data [6].

We do not know why the two indexes order OpenAI's flagships differently; AA's index leans on agentic and terminal work, but we did not audit Epoch's weighting, so treat that as a guess. The practical reading is that the top five or six are tied for most purposes, and the choice should rest on price, speed, harness and the kind of work you do.

Coding agents: Claude and GPT lead, but the harness is half the score

For game developers the most relevant AA product is the Coding Agents Index. It averages three evaluations (DeepSWE v1.1, SWE-Atlas QnA and Terminal-Bench v4), each run with a model inside a vendor's own coding agent [3]. That is both a feature and a flaw: it tells you what you get from Claude Code or Codex, but not how much is the model and how much the tooling. It has eight rows and no open-weights model.

Agent (harness) Model Index (0 to 1) Mean cost per task
Claude Code Sonnet 5.5 (max) 0.684 $14.19
Claude Code Opus 5.5 (max) 0.660 $13.04
Antigravity CLI Gemini 4 Argon 0.638 $5.84
Codex GPT-6.1 Sol (xhigh) 0.629 $1.04
Claude Code Fable 5.1 (max) 0.622 $12.39
Codex GPT-6 Astra (max) 0.616 $7.47
Grok Build Grok 4.7 (xhigh) 0.563 $8.82
Muse Code Muse Spark 1.3 (max) 0.543 $3.98

Source: Artificial Analysis Coding Agents Index, accessed 9 October 2026 [3].

The first six rows are within 0.07 of each other, and the cost column spans an order of magnitude: GPT-6.1 Sol in Codex reaches 0.629 for a reported $1.04 per task, while Sonnet 5.5 in Claude Code reaches 0.684 for $14.19. Whether the extra 0.055 is worth thirteen times the cost depends on your project. Our post on the best AI for coding games goes through the components, including the fact that Gemini 4 Argon has the highest DeepSWE score of the eight (0.788) despite ranking third overall.

The second coding instrument is LMArena's WebDev board, where people vote on which of two model-built web apps they prefer. Claude Opus 5.5 leads with 1813 Elo, ahead of GPT-6 Astra (1786), Claude Sonnet 5.5 (1774), GPT-6.1 Sol (1755) and Claude Fable 5.1 (1744). Gemini 4 Argon follows at 1678 [4]. This is the closest published measure to "make me a small browser game that looks good", and also the one most exposed to taste and visual polish.

AA's older standalone Coding Index no longer exists in its data, and the familiar SWE-bench Verified and LiveCodeBench scores are empty for every current model in our snapshot. Epoch's SWE-bench Verified runs stop at about 83.5% for models up to roughly April 2026, near saturation, so we do not mix them with September models [1][6].

Human preference: the one board where Google leads

LMArena's text board ranks models by blind pairwise votes on general chat prompts, with style control on. Gemini 4 Argon is first at 1525 Elo, with Claude Opus 5.5 at 1507, Claude Fable 5.1 at 1501, Gemini 3.8 Flash at 1497 and Muse Spark 1.3 at 1494 [5]. Several of the intelligence leaders are further down: Sonnet 5.5 is 32nd at 1476 and GPT-6 Astra 35th at 1475 [5].

A board that rewards answers people like reading measures something different from one that scores terminal tasks, and for dialogue and flavor text it is arguably closer to what you need. It is still a general chat board, not a game-writing one, and Argon's lead comes from a model most readers cannot use yet. Elo values from different boards, and from AA's boards, are on different scales and we never put them on one axis.

Price: a spread of about 100 times, and list price misleads

The cheapest and dearest models in our snapshot with an index of 20 or more differ in list price by a factor of about 100 (Claude Fable 5.1 at $20 blended per million tokens, against MiMo-V2.6-Flash at $0.175) [1]. The more useful view is price against intelligence.

Price against intelligence: what each point of the index costsBlended list price per million tokens (3 input : 1 output, log scale) against AA Intelligence Index. Up and left is better.
Price against intelligence: what each point of the index costs$20$30$40$50$60$0.1$0.3$1$3$10$30AA Intelligence IndexBlended price, $ per 1M tokens (log scale)Claude Opus 5.5Claude Sonnet 5.5Claude Fable 5.1GPT-6 AstraGemini 4 ArgonGPT-6.1 SolMuse Spark 1.3Grok 4.7MiMo-V2.6-ProGLM-5.3Step 5 (Preview)Kimi K3Claude Haiku 5.5GPT-5.6 TerraGLM-5.3-FlashGemini 3.8 FlashQwen3.8 2.4T A95BDeepSeek V4.1 FlashMistral Large 4 (Preview)MiMo-V2.6-FlashDeepSeek V4 Pro (0813)Qwen3.8 27BMiniMax-M3InklingNemotron 3 Ultra 550B A55BGemini 3.5 Flash-LiteAnthropicOpenAIGooglexAIMetaAlibabaMistralStepFunMoonshotZ.aiXiaomiDeepSeekMiniMaxNVIDIAClosed weightsOpen weights (AA flag)

Models with an index of 20 or more and a list price. Prices are list prices on 9 Oct 2026 and include introductory pricing (Gemini 4 Argon). Gemini 3.8 Flash rises to $1.50/$7.50 in 2027. GLM-5.3 is flagged open by AA but closed by Epoch. Price per token is not cost per task: reasoning models spend more tokens. Source: Artificial Analysis, Intelligence Index v4.3.2, accessed 2026-10-09.

Three comparisons from that chart. Sonnet 5.5 scores 97% of Opus 5.5 at half the blended price ($4 against $8) [1][7]. Claude Haiku 5.5, released 7 October, scores 43.4, which is 75% of Opus 5.5's 57.6, at $0.20 blended: one-fortieth of the price (8 / 0.2 = 40). Note that Anthropic charges five times more for Haiku prompts over 100,000 tokens ($0.50/$2.50), which matters for anyone stuffing a lore bible into it [7]. And MiMo-V2.6-Pro, an open-weights model, scores 46.3 at $0.544 blended: 80% of Opus's score for 1/14.7 of the price. Claude Fable 5.1, at $10/$50, is the outlier in the wrong direction: it scores below both Opus 5.5 and Sonnet 5.5 and costs 2.5 to 5 times as much per token [1][7].

List price per token is still not what you pay. Reasoning models bill their hidden thinking as output tokens, so a model that thinks longer costs more per answer. AA reports the cost of running its whole index, and that column reorders the table. Claude Opus 5.5 costs $5.98 per run and Sonnet 5.5 $5.46; GPT-6.1 Sol, which has the same $2/$10 list price as Sonnet 5.5, costs $0.72 [1]. Same price per token, about 7.5 times the cost (5.461 / 0.724), which implies Sonnet consumed several times more tokens on AA's tasks. That is our inference from the arithmetic, not an AA statement. The cheapest runs among models scoring 35 or more are MiMo-V2.6-Flash ($0.06), GPT-6 Luna ($0.07) and MiMo-V2.6-Pro ($0.13). Our cost post turns this into bills for prototyping sessions, NPC dialogue at scale and art.

Gemini 4 Argon's $2/$10 is introductory, and secondary sources report $4/$20 later, though we did not find Google's own statement. Gemini 3.8 Flash costs $0.75/$3.75 through 2026 and $1.50/$7.50 from 2027, per Google's pricing page [8]. OpenAI's pricing page returned an HTTP 403 when we tried to read it, so its prices are AA's and LMArena's, which agree [1][4].

Speed and context: fast small models, uniform million-token windows

For a coding agent that runs for minutes, output speed is a convenience. For an NPC that answers a player in real time, time to first token is a design constraint.

Output speed against priceMedian output tokens per second against blended list price (log scale). Up and left is faster and cheaper; the frontier models sit at the bottom right.
Output speed against price01002003004005000.10.3131030Output speed, tokens per secondBlended price, $ per 1M tokens (log scale)Gemini 3.5 Flash-LiteClaude Haiku 5.5DeepSeek V4.1 FlashInklingClaude Sonnet 5.5Nemotron 3 Ultra 550B A55BGPT-6 LunaMuse Spark 1.3Gemini 3.8 FlashGPT-5.6 TerraMiniMax-M3Claude Opus 5.5GLM-5.3Grok 4.7Claude Fable 5.1MiMo-V2.6-FlashGLM-5.3-FlashQwen3.8-Flash-NextGPT-6 AstraQwen3.8 27BMiMo-V2.6-ProKimi K3Qwen3.8 2.4T A95BAnthropicOpenAIGooglexAIMetaAlibabaStepFunMoonshotZ.aiXiaomiDeepSeekMiniMaxNVIDIAClosed weightsOpen weights (AA flag)

Models with an index of 20 or more. Speeds are medians on AA's hosted endpoints and depend on the provider; self-hosted speed differs. Gemini 4 Argon and Mistral Large 4 (Preview) have no speed data. Output speed says nothing about time to first token, which includes thinking for reasoning models. Source: Artificial Analysis, speed and price (Intelligence Index v4.3.2 snapshot), accessed 2026-10-09.

On median output speed, AA has Gemini 3.5 Flash-Lite at 408.6 tokens per second (index 22.2), Claude Haiku 5.5 at 240.4, DeepSeek V4.1 Flash at 219.7, Claude Sonnet 5.5 at 141.8 and GPT-6 Luna at 138.6. The flagships are slower: Opus 5.5 at 97.2, Claude Fable 5.1 at 71.4 and GPT-6 Astra at 47.3 [1]. The slowest top-six model with speed data is roughly one-ninth as fast as the fastest small one (408.6 / 47.3 = 8.6). Speeds depend on the hosting provider.

Be careful with AA's time-to-first-token column. For reasoning models at maximum effort it includes thinking time, which is why Opus 5.5 shows a median of about 683 seconds. That measures a long reasoning run on AA's test prompts, not the wait on a short question, and it is useless for NPC planning unless you pick the non-reasoning or low-effort variant. Our NPC post does that homework: GPT-6 Luna's non-reasoning variant starts answering in 0.83 seconds, against 98.8 seconds at maximum effort.

Context windows have stopped being a differentiator at the top. Twenty-two of the 33 models in our snapshot list a window of exactly 1,000,000 tokens, including 14 of the 17 models above 40 on the index; the exceptions are Grok 4.7 (500,000), Qwen3.8 Max (983,616) and Kimi K3 (1,048,576) [1]. Advertised size is not usable memory, though. The binding limit for a long game is consistency across a large input, which is what AA's long-context test (AA-LCR) probes: Kimi K3 scores highest at 0.887, ahead of Step 5 Preview (0.883), MiMo-V2.6-Pro (0.863) and Fable 5.1 (0.853), while Opus 5.5 is 0.847 [1]. Large windows therefore do not decide the choice between frontier and mid-priced models; our posts on AI RPGs and AI world building argue that state kept outside the model matters more than window size.

Open weights versus closed: a persistent gap of months, not years

The best open-weights model on AA's index is MiMo-V2.6-Pro at 46.3 (released 21 September, MIT license on LMArena), 11.3 points behind Opus 5.5. Next come GLM-5.3 at 44.8 and Kimi K3 at 43.6 [1]. The line chart shows the shape: both series rise steeply, and the open line trails the closed one by a few months.

Best model to date versus best open-weights model, December 2024 to October 2026Running maximum by release date on AA Intelligence Index v4.3.2. The open line trails by months, not by years.
Best model to date versus best open-weights model, December 2024 to October 20260102030405060Jan 2025Apr 2025Jul 2025Oct 2025Jan 2026Apr 2026Jul 2026AA Intelligence IndexBest model (any)Best open weights

Points before about December 2025 are AA back-filled estimates. Each step is a new record, so the lines only rise. MiMo-V2.6-Pro and GLM-5.3 open status should be checked against their licenses. Our how-fast-is-ai-improving post shows the full history back to 2022. Source: Artificial Analysis, Intelligence Index v4.3.2 (running maximum derived by CreateGame.ai), accessed 2026-10-09.

Epoch's index shows a similar picture on a different scale. Its best open-weights model with a clear label is Kimi K3 at 157.5, which is 9.9 ECI points below Opus 5.5 at 167.3 [6]. Measured in time rather than points, our how-fast-is-ai-improving analysis puts the lag at roughly three to four months, with a median of about 4.3 months over the past two years on AA's index; Epoch's own figure is about four months since January 2026 (see our AI progress analysis and [14]). In points the gap looks bigger than before because the scale stretched, while in months it has held or shrunk.

Open does not always mean free to ship. AA and LMArena call GLM-5.3 open (MIT per LMArena), but Epoch labels it closed weights, so we never call it open without that caveat. Kimi K3 uses a custom license, MiniMax-M3 a community license, Qwen3.8 27B is Apache 2.0, and DeepSeek V4.1 Flash and the MiMo models are MIT [1]. Read the license before shipping a model inside a game. The largest open models are also far too big for a consumer GPU, so "open" mostly means "rentable from several hosts at lower prices", with local use reserved for smaller models like Qwen3.8 27B.

The best AI for game development, job by job

The caveat comes first: none of the benchmarks above is a game benchmark, and we ran no head-to-head tests. Every pick is our reading of published results, and each links to the deep dive with the evidence.

Job Our pick Runner-up Budget or open-weights option Deep dive
Game code (web and general) Claude Opus 5.5, or Sonnet 5.5 for volume GPT-6.1 Sol in Codex MiMo-V2.6-Flash or Kimi K3 Coding
Engine scripting (Unity, Godot, Unreal) A frontier agent wired to the editor Gemini 4 Argon once open Qwen3.8 27B locally Engine tools
NPC dialogue GPT-6 Luna, non-reasoning Claude Sonnet 5.5 at low effort DeepSeek V4.1 Flash NPCs
Story and RPG writing Claude Opus 5.5 or Sonnet 5.5 Kimi K3 MiMo-V2.6-Pro RPGs
World building Claude Sonnet 5.5 Claude Opus 5.5 for consistency passes MiMo-V2.6-Pro World building
Art GPT Image 2.5 for shipping art, Nano Banana 2.1 for exploration MAI-Image-2.6 None commercially usable that we found (Qwen-Image-2.1 is research-licensed) Art
Playtesting and agents GPT-6 Astra or Claude Opus 5.5 with a harness Claude Fable 5.1 Cheap model for scripted checks Playing games

Game code

If you want the model that people prefer when it builds a playable web game, the answer on the published data is Claude Opus 5.5: first on WebDev at 1813, first on AA's index, and, by the same margin, the preferred model in the head-to-head votes on generated web apps [1][4]. If you want the best agent on a big, multi-file project, Sonnet 5.5 in Claude Code has the highest Coding Agents Index (0.684) and the top AA Terminal-Bench 4.0 score (0.636 against Opus's 0.596) at half the list price of Opus [1][3]. Our habit would be Sonnet 5.5 for day-to-day work and Opus 5.5 for the hard problem. The runner-up is GPT-6.1 Sol in Codex, which reaches 0.629 on the agents index for a reported $1.04 per task, and the case for it is cost per finished task rather than raw score [3].

For the budget slot, our coding post points at MiMo-V2.6-Flash: 1637 Elo on WebDev, 17 points behind Kimi K3 (1654, the best open-weights model on that board) and 176 behind Opus 5.5, at $0.14/$0.28 per million tokens [1][4]. Both are far weaker on terminal-style agent work: the best open-weights Terminal-Bench 4.0 score in our data is GLM-5.3 at 0.419, against 0.520 to 0.636 for the top six [1]. A sensible hybrid is a frontier model for architecture and bug-hunting and a cheap one for repetitive edits.

Engine scripting

This is the weakest area for evidence, and we say so plainly. We found no clean, current, per-model benchmark for Unity C#, Godot GDScript, Unreal C++ or shaders. The one engine benchmark we know of, GameDevBench (arXiv 2602.11103), covers Godot only, and its headline result is that the best agent solved 53.8% of tasks, with graphics tasks well below gameplay tasks [9]. The gain from giving the agent screenshots was large (GPT-5.4 went from 41.1% to 52.0%, per our coding and progress posts).

So our pick is less a model than a setup: a frontier coding agent (Claude Code or Codex) connected to the editor through an MCP server, able to run the game and look at the output. Both major engines are heading that way, as we describe in our look at Unity and Unreal's AI tools: Unity ships a metered in-editor assistant in beta, and Unreal 5.8 has an experimental MCP server for any model you bring. Gemini 4 Argon is the one to watch, because it scored the best DeepSWE result in AA's table (0.788) at $5.84 per task, but it is pre-release [3]. For private, unmetered local edits, Qwen3.8 27B (Apache 2.0, index 33.7) is the usual suggestion.

NPC dialogue

Live dialogue has three constraints at once: start answering in about a second, stay cheap when the game succeeds, and obey the character sheet. Time to first token filters the field. In AA's 9 October data, GPT-6 Luna's non-reasoning variant starts in 0.83 seconds and costs about $0.16 per 1,000 lines on our assumptions (1,200 input tokens and 80 output tokens per line), while Claude Sonnet 5.5 at low effort starts in 0.88 seconds and costs about $3.20 per 1,000 lines (see our NPC latency post). That is a twentyfold cost difference for a similar response time.

Our pick is GPT-6 Luna in non-reasoning mode for bark-style lines and small talk, with Sonnet 5.5 at low effort for the few characters whose lines must be better. DeepSeek V4.1 Flash (MIT, $0.30/$1.20, about $0.46 per 1,000 lines) is the open-weights option, and the NPC post lists smaller open models for those who want to host. The caveats are real: AA's index gives Luna's reasoning-off variant about 18, and no benchmark measures character believability, so this pick rests on latency and price, not on a test of dialogue quality. Voice raises the bar again.

Story and RPG writing

Prose quality has no authoritative leaderboard, and the boards disagree. On a 31 August 2026 comparison by Digital Applied, EQ-Bench's creative-writing list put Claude Opus 5 first, Kimi K3 second and GLM-5.3 third, while arena.ai's creative category put Claude Fable 5 first and Kimi K3 twenty-third [10]. Surge AI's study with professional writers argued that LLM-judged boards reward heavy use of literary devices [11].

Our reading is that Claude models are the safe pick when human preference matters, so Opus 5.5 where the writing carries the experience and Sonnet 5.5 for the volume of an ongoing campaign. Kimi K3 is the runner-up and the best open-weights choice on the LLM-judged board, with the highest long-context score in our data. For cost, an hour of play at 40 turns with an 8,000-token context comes to about $0.76 on Sonnet 5.5, $1.52 on Opus 5.5 and about $0.15 on MiMo-V2.6-Pro, by the arithmetic in our RPG post. Memory, not eloquence, is the real constraint on a 50-hour campaign.

World building

World building is closer to editing a large, internally consistent document than to writing a scene. The relevant measures are long-context accuracy and how rarely the model invents facts. On AA-LCR, Kimi K3 (0.887) and MiMo-V2.6-Pro (0.863) are ahead of Opus 5.5 (0.847), and on AA-Omniscience Opus 5.5 leads at 46.4 against 19.7 for Kimi K3 [1]. We did not open Omniscience's methodology page; we understand it to reward correct answers and penalise confident wrong ones, which is why we treat it as a rough hallucination signal and not more.

Our pick is Claude Sonnet 5.5 for the working draft and Opus 5.5 for consistency passes over the finished bible, because resending a 100,000-word world bible costs only about $0.27 per request on Sonnet 5.5 uncached by the world-building post's estimate, so a stronger model for a final audit is cheap. MiMo-V2.6-Pro is the open-weights option. The process matters more than the pick; see our world-building analysis.

Art

Image models are a separate market with a separate leaderboard. On AA's text-to-image arena (Elo, not comparable with LMArena's), GPT Image 2.5 Sunburst leads at 1198 with Flare at 1191, followed by GPT Image 2 (1172), Nano Banana 2.1 (1160) and Grok Imagine Image 2.0 (1156) [15]. LMArena agrees on the top three and reorders the rest [16]. OpenAI's top models cost $0.211 per image on AA's price; Nano Banana 2.1 costs $0.034, about one-sixth (0.034 / 0.211 = 0.16), for 38 Elo points less.

Our picks follow from that price gap and from our art deep dive: GPT Image 2.5 for key art you ship, Nano Banana 2.1 (or MAI-Image-2.6) for exploration and volume. The open-weights column is thin. The best open model on AA's arena, Qwen-Image-2.1, sits 20th at 1035, and the art post's reading of its license limits it to research and evaluation. Arenas also do not test what sprite work needs, such as consistency across a character sheet or a true pixel grid, and 3D asset generation has no benchmark-quality source at all.

Playtesting and agent play

Language models can now finish a long turn-based game with a purpose-built harness, and they are still near zero on real-time games from raw pixels. Our post on whether AI can play video games reports that the top language model on the BALROG board is GPT-6-Astra-Max (68.3% progress), that Claude models have completed Pokémon Red and FireRed with their harnesses, and that the best result on the real-time VideoGameBench was 0.48% of the full benchmark.

For testing, the practical use is narrower than "AI plays the game": scripted playthroughs, regression checks, balance simulations and bug-hunting in turn-based or menu-driven logic. We would use GPT-6 Astra or Claude Opus 5.5 with a harness, Fable 5.1 as the runner-up, and a cheap model for repetitive scripted checks. Our snapshot has no game-playing results for current open-weights models.

A whole game from one sentence

The jobs above assume you assemble the game yourself. A separate category generates the whole thing, from hosted prompt-to-playable platforms to engine assistants; our comparison of AI game makers sorts it out. It is also the category CreateGame.ai is aimed at, though we are pre-launch with only a landing page and a waitlist today.

What would change these picks

Gemini 4 Argon reaching general availability. It already leads LMArena text (1525), has the best DeepSWE score in the agent table and sits at 52.6 on AA's index. Access, final pricing and independent runs will decide whether it joins the coding recommendation [1][3][5].

A price change. Argon's introductory price and Gemini 3.8 Flash's 2027 step-up show that today's rates are not promises; check before building a live service on them.

A benchmark that measures games. Engine work, level layout, game feel and character voice have little coverage. Our world-models post covers research that might eventually blur the line between model and engine, and the web-games post explains why browser-first development suits AI-generated code.

The next release round. The best model on the index went from 49.6 (Claude Fable 5, 9 June) to 57.6 (Opus 5.5, 22 September) in about three and a half months, and our AI progress analysis reports a gain of 32.7 points over twelve months against 13.5 the year before. Expect this page to be stale by early 2027.

How the series uses benchmarks

Every post in the series leans on the same handful of public measures. Here is what each tells you, and where it misleads.

Measure What it measures Known weaknesses
AA Intelligence Index v4.3.2 A composite of ten evaluations (AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, SciCode, HLE, GDP.pdf, CritPt, AA-Omniscience, AA-LCR) [2] Rescaled between versions; tilted toward work, science and agent tasks, not games or creative writing; pre-2026 scores are back-filled estimates
AA Coding Agents Index Model plus its vendor's coding agent on DeepSWE, SWE-Atlas QnA and Terminal-Bench v4 [3] Mixes model and harness; eight rows only; no open-weights model
LMArena (text, WebDev, image) Elo from blind human pairwise votes [4][5] Reflects taste and polish; Elo scales are not comparable between boards or with AA; WebDev here is the raw, no-style-control view; some entries are variants or pre-release
Epoch Capabilities Index (ECI) A composite built by Epoch AI from many benchmarks, on its own scale [6] Excludes newer models until evaluated; open-weights labels can differ from AA's
SWE-bench Verified Fixing real GitHub issues in repositories Near saturation (Epoch's runs top out at 83.5%); null for every current model in our data
Terminal-Bench Agent tasks in a terminal Several versions in circulation (2.0, 2.1, Hard, 4.0); depends on the agent
GameDevBench Godot game-development tasks for coding agents [9] One engine; three revisions with different headline figures (54.5%, 53.8%, 49.0%)
EQ-Bench Creative Writing Prose scored by LLM judges Judges can favor their own style; disagrees with human-preference boards [10][11]

How we apply them

We rarely compare across instruments. AA's index, ECI and every Elo board answer different questions on different scales. We use them to ask whether independent rulers agree, and when they do not (OpenAI's flagships sit lower on AA than on Epoch; Gemini 4 Argon is first on LMArena text and fifth on AA), we say so and do not average them.

We use the default variant AA lists, and say when that is not what you would run. The snapshot holds the highest-effort variant of each model. Lower-effort variants score lower and cost less per task: Sonnet 5.5 at low effort scores 36, against 56 at maximum, according to our NPC post.

We treat leaderboards as dated measurements. Each number carries an access date, because prices, rankings and even index values change.

We distrust saturated and contaminated tests. Epoch's staff have said most benchmarks saturate within months, and a February 2026 preprint found nearly half of 60 major benchmarks showing saturation [12][13]. That is why SWE-bench Verified, once the default coding yardstick, has faded from our comparisons.

We separate evidence from judgment. Picks are labelled as our reasoning from published numbers, arithmetic is shown and estimates are marked.

How we reviewed this

We reviewed published sources only: leaderboards, benchmark pages, pricing pages, papers, and the sibling posts in this series that existed when we wrote this, whose numbers this pillar reuses. We ran no model, agent or tool ourselves.

The cross-model data comes from the shared snapshot we assembled on 9 October 2026 from Artificial Analysis, LMArena and Epoch AI (ECI, CC BY 4.0). The running-maximum series and the price comparisons are our own derivations.

What we could not verify: OpenAI's pricing page returned HTTP 403, so OpenAI prices are AA's and LMArena's. Several vendors' pages other than Anthropic's and Google's were not checked. Gemini 4 Argon's later price rise comes from secondary reports. GLM-5.3's open-weights status conflicts between sources. We did not read the methodology pages for AA-Briefcase, AutomationBench-AA, GDP.pdf or AA-Omniscience. The EQ-Bench versus arena.ai comparison is secondhand, from one analysis. The Epoch 4-month lag and the saturation findings were read through our sibling post, which in turn cites search summaries for some of them. The art picks combine our snapshot's image-arena data with the art post's license reading.

FAQ

What is the best AI model for game development in October 2026? There is no single best. Claude Opus 5.5 leads AA's Intelligence Index (57.6) and LMArena WebDev (1813), Claude Sonnet 5.5 leads the Coding Agents Index (0.684), and GPT-6.1 Sol is cheapest per task in that table. Pick by the job.

Which AI is best for writing game code? On published data, Claude Opus 5.5 and Sonnet 5.5, with GPT-6.1 Sol in Codex as the cost-efficient alternative. See best AI for coding games for the components and caveats.

Is there a good open-source AI model for game development? Yes, with limits. MiMo-V2.6-Pro (46.3), GLM-5.3 (44.8, disputed status) and Kimi K3 (43.6) are the strongest open-weights models on AA's index, roughly 11 to 14 points behind the leader, and they are good at web code. They are weaker at terminal-style agent work. Check each license before shipping.

Which model is fast and cheap enough for NPC dialogue? Non-reasoning variants of small models, such as GPT-6 Luna (0.83 seconds to first token, about $0.16 per 1,000 lines on our assumptions). Details are in our NPC post.

How much does it cost to build a game with AI? Our estimate for a heavy weekend of agentic prototyping is about $11 to $12 on Sonnet 5.5 or GPT-6.1 Sol and under $1.50 on cheap models, from the cost post. Subscriptions and live-dialogue costs are covered there.

Can AI make a whole game from one prompt? Small browser games, yes; large polished ones, not reliably. AI game makers compared covers the tools.

Sources

  1. Artificial Analysis, model leaderboard (Intelligence Index v4.3.2, prices, speed, context), accessed 2026-10-09: https://artificialanalysis.ai/leaderboards/models
  2. Artificial Analysis, Claude Opus 5.5 model page (index version and component evaluations), accessed 2026-10-09: https://artificialanalysis.ai/models/claude-opus-5-5
  3. Artificial Analysis, Coding Agents Index, accessed 2026-10-09: https://artificialanalysis.ai/agents/coding-agents
  4. LMArena (arena.ai), WebDev leaderboard, accessed 2026-10-09: https://arena.ai/leaderboard/code/webdev
  5. LMArena (arena.ai), text leaderboard, accessed 2026-10-09: https://arena.ai/leaderboard/chat/text
  6. Epoch AI, benchmark data hub (Epoch Capabilities Index, SWE-bench Verified runs; CC BY 4.0), accessed 2026-10-09: https://epoch.ai/benchmarks
  7. Anthropic, API pricing, accessed 2026-10-09: https://platform.claude.com/docs/en/about-claude/pricing
  8. Google, Gemini API pricing, accessed 2026-10-09: https://ai.google.dev/gemini-api/docs/pricing
  9. Chi et al., "GameDevBench: Evaluating Agentic Capabilities Through Game Development," arXiv 2602.11103: https://arxiv.org/abs/2602.11103
  10. Digital Applied, "Which AI model writes best? Leaderboards disagree," 31 Aug 2026: https://www.digitalapplied.com/blog/which-ai-model-writes-best-leaderboards-disagree
  11. Surge AI, Hemingway-bench: https://surgehq.ai/blog/hemingway-bench-ai-writing-leaderboard
  12. Epoch AI, "Are AI benchmarks doomed?" (podcast; read via our sibling post): https://epoch.ai/epoch-after-hours/are-ai-benchmarks-doomed
  13. Benchmark saturation preprint, arXiv 2602.16763 (read via our sibling post): https://arxiv.org/pdf/2602.16763v1
  14. Epoch AI, open versus closed ECI gap (read via our sibling post): https://epoch.ai/data-insights/open-closed-eci-gap
  15. Artificial Analysis, text-to-image leaderboard, accessed 2026-10-09: https://artificialanalysis.ai/image/leaderboard/text-to-image
  16. LMArena (arena.ai), text-to-image leaderboard, accessed 2026-10-09: https://arena.ai/leaderboard/image/text-to-image
  17. CreateGame.ai series posts used for cross-checked numbers (relative links above): /blog/best-ai-for-coding-games/, /blog/ai-npcs/, /blog/ai-rpg/, /blog/ai-world-building/, /blog/cost-of-making-a-game-with-ai/, /blog/how-fast-is-ai-improving/, /blog/can-ai-play-video-games/, /blog/unity-unreal-ai-tools/

Building with these models yourself is one route; if you would rather type a sentence and get a playable game, join the CreateGame.ai waitlist at https://creategame.ai.