How Much Does It Cost to Make a Game with AI? Token Prices Compared
The cost of AI game development in October 2026 is small next to the cost of your own time, and very uneven between uses: an agentic coding session costs tens of dollars, 10,000 players chatting with AI characters can cost anywhere from $400 to $38,000 a month depending on the model, and a sprite sheet costs a few dollars to a few hundred. We worked through the published prices, plan terms and benchmark data to show the arithmetic, and we mark every estimate as an estimate.
Key takeaways
- List prices span a factor of about 100. Claude Opus 5.5 costs $4 per million input tokens and $20 per million output; Claude Haiku 5.5 costs $0.10 and $0.50; MiMo-V2.6-Flash, an open-weights model, costs $0.14 and $0.28 on AA's hosted listing.
- Our estimate for a heavy agentic prototyping session (a weekend game) is roughly $11 to $12 on Claude Sonnet 5.5 or GPT-6.1 Sol, $23 on Opus 5.5, and under $1.50 on cheap models. A subscription at $100 a month pays for itself in about nine such sessions by our assumptions.
- Live AI dialogue is where costs scale. For 10,000 daily players with 20 exchanges each, we estimate $960 per month on a Haiku-class model and $19,200 on Sonnet 5.5, before caching. Hidden reasoning tokens can more than double the small-model figure.
- Image generation is priced per image: from $0.01 (Meta's Muse Image) to $0.211 (OpenAI's GPT Image 2.5). Generating 1,200 sprites costs between $12 and $253.
- The price of a fixed capability level has been falling by tens of times per year, according to Epoch AI and a16z and to our own rough series. The price of the best model is not falling: top-tier blended prices are still $8 to $20 per million tokens.
- Self-hosting rarely beats the API on dollars. A rented H100 at about $2.50 an hour costs roughly $1,825 a month if left on, and a high-end consumer GPU now costs several times its launch price. It wins on privacy and on shipping a model inside the game.
Token prices in October 2026
Every large language model API bills by the token, roughly three-quarters of an English word. Prices are quoted per million tokens, split between input (what you send, including the whole conversation or codebase context) and output (what the model writes). Output costs several times more, typically four to five times. To compare models on one number, Artificial Analysis (AA) publishes a blended price weighted three parts input to one part output [1]. We use the same convention: blended = (3 x input + output) / 4.
Gemini 4 Argon $2/$10 is introductory and pre-release; press reports $4/$20 later. Gemini 3.8 Flash rises to $1.50/$7.50 in 2027. Open-weights prices are hosted-API prices from AA. OpenAI prices not verified at the vendor (pricing page returned HTTP 403). Source: Artificial Analysis; Anthropic and Google pricing pages, accessed 2026-10-09.
The table below gives the list prices we used throughout this post, from AA's leaderboard of 9 October 2026, checked against Anthropic's and Google's own pricing pages where we could [1][2][3]. OpenAI's pricing page returned an HTTP 403 when we tried to read it, so OpenAI prices come from AA and LMArena, which agree [1].
| Model | Input $/1M | Output $/1M | Blended $/1M | Cached input $/1M | Notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 | 4.00 | 20.00 | 8.00 | 0.20 | AA Intelligence Index 57.6 |
| Claude Sonnet 5.5 | 2.00 | 10.00 | 4.00 | 0.10 | Index 56.0 |
| Claude Haiku 5.5 | 0.10 | 0.50 | 0.20 | 0.01 | Prompts over 100k tokens: $0.50 / $2.50 |
| Claude Fable 5.1 | 10.00 | 50.00 | 20.00 | 0.25 | Index 53.4 |
| GPT-6 Astra | 10.00 | 50.00 | 20.00 | 1.00 | Index 52.7 |
| GPT-6.1 Sol | 2.00 | 10.00 | 4.00 | 0.10 | Index 51.8 |
| GPT-6 Luna | 0.10 | 0.50 | 0.20 | 0.01 | Index 38.1 |
| Gemini 4 Argon | 2.00 | 10.00 | 4.00 | 0.10 | Pre-release; introductory price |
| Gemini 3.8 Flash | 0.75 | 3.75 | 1.50 | 0.075 | Rises to $1.50 / $7.50 in 2027 |
| Gemini 3.5 Flash-Lite | 0.30 | 2.50 | 0.85 | 0.03 | Index 22.2 |
| Kimi K3 (open) | 3.00 | 15.00 | 6.00 | 0.30 | Hosted price |
| GLM-5.3 (open, see note) | 1.40 | 4.40 | 2.15 | 0.26 | Open status disputed |
| DeepSeek V4.1 Flash (open) | 0.30 | 1.20 | 0.525 | 0.006 | MIT |
| MiMo-V2.6-Flash (open) | 0.14 | 0.28 | 0.175 | 0.0028 | MIT |
Source: Artificial Analysis, accessed 2026-10-09 [1]; Anthropic pricing page [2]; Google pricing page [3]. Index values are AA Intelligence Index v4.3.2.
Four pricing mechanics change the real bill more than the headline rate.
Caching. Agentic coding resends the same large context on every turn. Cached input is billed at a fraction of the normal rate: 5% of base input on Opus 5.5 and Sonnet 5.5 ($0.20 and $0.10 per million), about 10% on most other models in AA's data [1][2]. Anthropic charges 1.25 times the base price to write to the 5-minute cache [2]. Almost all of an agent session's input is cache reads, which is why session costs are far lower than naive token arithmetic suggests.
Batch. Anthropic's Batch API gives 50% off input and output for non-urgent work, for example Opus 5.5 at $2 and $10 [2]. That suits overnight jobs such as generating item descriptions or localising text; it is no use for live chat.
Long context and fast modes. Claude Haiku 5.5 charges five times more ($0.50 and $2.50) when a prompt exceeds 100,000 tokens, while Opus and Sonnet 5.5 include a 1M-token window at standard rates [2]. Anthropic's fast mode for Opus 5.5 is priced at $8 and $40 [2]. OpenAI's credits for Codex multiply 2x for fast mode and 6x for ultrafast [5].
Reasoning tokens. Reasoning models bill their hidden thinking as output tokens. A model that writes a 60-token reply after 400 tokens of thought bills 460 output tokens. AA's listed prices are per token, not per answer, so the same model can cost several times more per task at higher effort settings.
What we assumed, and why
Every worked example below follows the same discipline: state the assumption, show the multiplication, and flag the result as an estimate. These are our assumptions, not measurements. We did not run these workloads. When you scale them, change the assumptions first and the prices second, because token volume varies more than price.
Example 1: prototyping a small browser game with an agentic coder
An agentic coding tool such as Claude Code or Codex works in a loop: it reads project files, writes a change, runs the game or tests, reads the output and repeats. Each turn resends the conversation, so input volume grows with the number of turns, but most of it is served from cache.
Our assumed weekend session: 200 agent turns. Each turn re-reads about 60,000 tokens of context from cache, adds about 5,000 tokens of fresh input (new file contents, tool output) and produces about 4,000 output tokens including thinking. That gives:
- Cached reads: 200 x 60,000 = 12 million tokens
- Fresh input: 200 x 5,000 = 1 million tokens
- Output: 200 x 4,000 = 0.8 million tokens
On Claude Sonnet 5.5: 12M x $0.10 = $1.20 in cache reads; 1M x $2.50 (the 1.25x cache-write rate) = $2.50; 0.8M x $10 = $8.00. Total about $11.70. On Opus 5.5: $2.40 + $5.00 + $16.00 = $23.40. On GPT-6.1 Sol, which has no separate cache-write charge: $1.20 + $2.00 + $8.00 = $11.20. On DeepSeek V4.1 Flash: $0.07 + $0.30 + $0.96 = $1.33. On MiMo-V2.6-Flash: $0.03 + $0.14 + $0.22 = $0.40.
CreateGame.ai editorial estimate, not a measurement. Token volumes are our assumption; real sessions vary by several times. Anthropic cache writes at 1.25x input; other vendors' fresh input at list price. Excludes subscription plans. Source: Vendor list prices via CreateGame.ai snapshot; Anthropic pricing page, accessed 2026-10-09.
The full set of estimates: Claude Opus 5.5 $23.40, Kimi K3 $18.60, Claude Sonnet 5.5 $11.70, GPT-6.1 Sol $11.20, GLM-5.3 $8.04, Gemini 3.8 Flash $4.65, DeepSeek V4.1 Flash $1.33, Claude Haiku 5.5 $0.65, GPT-6 Luna $0.62 and MiMo-V2.6-Flash $0.40. They tell you that output tokens dominate (68% of the Sonnet bill) and that a model that talks less is cheaper at the same price.
Two sanity checks make us comfortable with the order of magnitude, though not the exact figure. First, AA's Coding Agents Index reports a mean cost per task of $13.04 for Opus 5.5 in Claude Code, $14.19 for Sonnet 5.5, $5.84 for Gemini 4 Argon in the Antigravity CLI and $1.04 for GPT-6.1 Sol in Codex [11]. A benchmark task is smaller than a weekend project, so those are costs for one task, not one game. Second, a public rerun of the "Claude of Duty" browser shooter by another developer was reported at about ten hours and 1.3 million tokens [13]. If that token count means total tokens at a blended Sonnet rate of $4 per million, it is $5.20 (and $10.40 on Opus 5.5 at $8 blended). We do not know how that source counted tokens, and it probably excluded cache reads; our 13.8 million total is about ten times larger. So treat $5 to $25 as the plausible range for a weekend on a frontier model and higher for sloppy prompting or a large codebase.
The model choice, then, is about quality more than price. Our post on the best AI for coding games looks at the benchmarks: GPT-6.1 Sol delivered 92% of Sonnet 5.5's Coding Agents Index at about 7% of its reported task cost, and we cannot yet explain why from AA's published data [11].
The bigger numbers come from running many sessions. A month of evening work at 20 sessions would cost about $234 on Sonnet 5.5 at our assumptions (20 x $11.70), against $100 for a Claude Max plan, which is where subscriptions start to make sense.
Subscriptions versus the API
Subscriptions trade a flat fee for usage limits. They are priced for individuals, and for sustained agent use they are almost always cheaper than the API, with the caveat that you cannot see exactly what you are buying.
| Plan | Price | What the vendor says | Source |
|---|---|---|---|
| Claude Pro | $20 monthly, $17 annual | Includes Claude Code; five-hour windows and weekly limits | [4] |
| Claude Max | From $100 (5x or 20x Pro usage) | Includes Claude Code; shared pool across web, desktop and Claude Code | [4] |
| ChatGPT Plus | $20 | Codex included; e.g. GPT-6.1 Sol 15 to 160 local messages per 5 hours (estimate) | [5] |
| ChatGPT Pro | $100, $200 or $500 | No five-hour limit at present | [5] |
| Google AI Ultra | $100 and $200 tiers (reported) | 5x or 20x higher limits in the Gemini app and Antigravity than Pro | [6] |
| Cursor | $20 individual (Pro; higher tiers reported by third parties at $60 and $200) | Includes a set amount of model usage; on-demand billed after | [7] |
Plan prices as listed on vendor pages we read on 9 October 2026; Google's tiers come from press coverage of its I/O 2026 announcement [6], which we read only through search summaries.
Comparing plans to the API is hard because plan allowances are given in windows, not tokens. OpenAI's pages show the nearest thing to a conversion: extra Codex usage is bought as credits at a per-million-token rate, with GPT-6.1 Sol at 50 credits per million input tokens and 250 per million output tokens, and GPT-6 Luna at 2.5 and 12.5 [5]. The ratio between those two models (20 times) matches the ratio of their API list prices ($2 versus $0.10 input), so credits scale with list price; the page we read did not give us the dollar value of a credit, so we do not convert.
Our rule of thumb from the estimates above: if you do fewer than about eight serious agent sessions a month on a frontier model, pay per token (8.5 sessions x $11.70 is about $100). Above that, a $100 plan is cheaper at list prices, assuming it does not throttle you mid-session. For API-style use in a game itself, such as NPC dialogue, there is no subscription alternative; you pay per token.
One conflict to flag: a third-party article we found said Anthropic removed Claude Code from the $20 Pro plan in April 2026. Anthropic's own pricing page, read on 9 October 2026, lists it as included [4]. We went with the vendor page. Plans change often, and plan limits are the least stable number in this post.
Example 2: AI NPC dialogue for N players
Dialogue is the cost that grows with your audience. Our assumed scenario: 10,000 daily active players, each having 20 exchanges per day with AI characters. Each exchange sends 1,200 input tokens (character prompt, memory summary, recent lines and the player's message) and gets about 80 tokens back (two short sentences). That is 10,000 x 20 = 200,000 exchanges per day, or 6 million per 30 days.
For Claude Haiku 5.5: input 1,200 x $0.10 / 1M = $0.00012; output 80 x $0.50 / 1M = $0.00004; $0.00016 per exchange. 200,000 x $0.00016 = $32 per day, or $960 per 30 days. That is about $0.096 per daily player per month.
CreateGame.ai editorial estimate. Reasoning models bill hidden thinking tokens as output, which can multiply the output share several times; use non-reasoning or low-effort settings. Prices at 9 Oct 2026. Source: Artificial Analysis list prices via CreateGame.ai snapshot, accessed 2026-10-09.
Across models, all at 30 days with no caching: MiMo-V2.6-Flash $1,142; GLM-5.3-Flash $1,320; Gemini 3.5 Flash-Lite $3,360; DeepSeek V4.1 Flash $2,736; Claude Sonnet 5.5 $19,200; Claude Opus 5.5 $38,400. GPT-6 Luna matches Haiku at $960. The spread between the cheapest and dearest is 40 times, and it is almost entirely a design decision: how smart should a shopkeeper be?
Three adjustments change the numbers materially.
Caching. The character prompt is repeated on every request. If 1,000 of the 1,200 input tokens are cached on Haiku 5.5 at $0.01 per million, the input cost falls to (1,000 x $0.01 + 200 x $0.10) / 1M = $0.00003, and the exchange to $0.00007. Over 200,000 exchanges that is $14 per day, $420 per month, a 56% cut. We ignored cache-write fees, which are small by comparison.
Hidden reasoning. If a model thinks for 400 tokens before an 80-token reply, output becomes 480 tokens. On Haiku 5.5: $0.00012 + 480 x $0.50 / 1M ($0.00024) = $0.00036 per exchange, or $2,160 per month, 2.25 times the baseline. The model you pick for dialogue must be run at low or no reasoning effort. AA's time-to-first-token for reasoning models includes thinking and runs to hundreds of seconds for the top models at maximum effort (Opus 5.5 shows 683 seconds), so those figures are useless for NPC latency; we cover speed in our post on AI NPCs [1].
Quality. The cheapest models are not the same quality, and the models in the 40-ish range on AA's index (Haiku 5.5 at 43.4, MiMo-V2.6-Flash at 37.9) are good at short in-character lines and weaker at long-term consistency. The cost of consistency is a bigger prompt or a bigger model.
What does this mean for a business? A premium game sold once cannot afford open-ended per-token costs for every player forever: at $0.096 per daily player per month, a title with 10,000 daily players sold at $20 needs the equivalent of about 48 new sales per month (at $20 each, $960) just to cover chat. We do not know the revenue per player of any particular game, so we leave this as a formula: monthly dialogue cost per daily player, multiplied by expected lifetime months, must be a small share of revenue per player. A free-to-play game with strong monetisation can afford more; a one-time purchase usually needs a cheap model, a message cap or on-device inference (see below).
Example 3: generating art with image models
Image models bill per image, usually by resolution. AA's text-to-image leaderboard lists API prices per 1,000 images alongside an Elo rating from head-to-head votes [12]. Elo here is AA's scale and not comparable to LMArena's.
| Model | AA Elo | $ per image | 1,200 images |
|---|---|---|---|
| GPT Image 2.5 Sunburst (max) | 1198 | 0.211 | $253 |
| Nano Banana 2.1 (Google) | 1160 | 0.034 | $41 |
| Grok Imagine Image 2.0 | 1156 | 0.060 | $72 |
| MAI-Image-2.6 (Microsoft) | 1151 | 0.039 | $47 |
| FLUX 3 Image (Black Forest Labs) | 1117 | 0.048 | $58 |
| Muse Image (Meta) | 1115 | 0.010 | $12 |
| MAI-Image-2.6-Flash | 1105 | 0.020 | $24 |
Source: Artificial Analysis text-to-image leaderboard, accessed 9 October 2026 [12]. Google's own pricing page lists Nano Banana 2.1 at about $0.0336 per 1K image, $0.0504 at 2K and $0.113 at 4K [3]. GPT Image 2.5's price is AA's number; OpenAI's pricing page was unreachable.
Our assumed project: a small 2D game with 300 distinct sprites and tiles. We assume four generations per kept asset (one keeper in four), which is 300 x 4 = 1,200 generations. Multiplying: $253 at GPT Image 2.5, $41 at Nano Banana 2.1, $12 at Muse Image. The fourfold retry rate is our guess; for consistent style across a sprite sheet it could be higher, and generation is only the first step before cleanup, background removal and animation, which we do not price. Cost per usable asset is then price divided by hit rate: at 25% hit rate, Nano Banana 2.1 costs $0.136 per kept sprite.
The interesting number is the ratio. GPT Image 2.5 Sunburst leads the AA board by 38 Elo points over Nano Banana 2.1 and costs 6.2 times more ($0.211 versus $0.034). Whether those 38 points justify the premium depends on whether your game's look depends on them. For production art we would test two candidates on your style before buying volume. Our comparison of image models for game art is in the game art post. Open-weights image models such as Qwen-Image-2.1 (AA Elo 1035) have no per-image fee but sit about 160 points below the top proprietary models, and you pay for the GPU [12].
Example 4: running models yourself
Open-weights models let you pay for hardware instead of tokens. The question is whether the hardware ever comes out cheaper.
Renting a data-center GPU. Specialised clouds list H100 80GB rentals at roughly $2 to $4 per hour in 2026, depending on provider and tier, with community marketplaces lower and hyperscalers higher; the sources we found disagree and are mostly comparison blogs, so check live pricing [14]. At $2.50 an hour: 730 hours x $2.50 = $1,825 per month if always on, or 176 hours (8 hours x 22 days) x $2.50 = $440 if you start it only when working. For comparison, $1,825 would buy about 3.5 billion tokens at DeepSeek V4.1 Flash's blended hosted price of $0.525 per million ($1,825 / 0.525 = 3,476 million). A rented GPU only beats that if it generates more than 3.5 billion useful tokens a month, which requires keeping it saturated with requests. We found no throughput numbers we trust for the largest open models on one GPU, so we cannot say whether that is realistic. For a small team the hosted API usually wins.
Buying a consumer GPU. An RTX 5090 launched at $1,999 in January 2025, but 2026 street prices we found ranged from about $2,400 to over $5,000 across sources; the most specific US retailer scan we saw (1 September 2026) found a lowest new price of $4,999.99 [15]. Assuming $4,500 amortised over 36 months is $125 per month. Add electricity at an assumed 450 W average for 4 hours a day: 0.45 kW x 4 h x 30 days = 54 kWh, at $0.17 per kWh about $9. That is roughly $134 a month, compared with $100 for a Claude Max plan, and what you can run on it is a smaller model: the community guides we found put Qwen3.8 27B at about 16 to 18 GB at 4-bit, with reported speeds around 142 tokens per second in a favourable single-user setup and 15 to 30 on other hardware [16]. On WebDev Qwen3.8 27B scores 1593 against Opus 5.5's 1813 and has a Terminal-Bench 4.0 score of 0.056 [1][17]. So local models are not a cheaper way to get frontier coding; they are a way to get private, unmetered, good-enough assistance if you own the card already.
Shipping a model inside the game. The case where open-weights economics change is dialogue on the player's own machine. A small model that runs on the player's GPU or CPU moves inference cost from your invoice to their hardware and their electricity. We have not seen reliable data on how well a 1-to-4-billion-parameter model sustains a character, and it forces hardware requirements, so we treat it as promising and unproven. Models like Google's Gemma 4 31B are open and list at $0 on AA because AA tracks no hosted price [1].
Two further facts help in comparing. First, open-weights models are sold by many hosts, and the prices differ: AA lists a single hosted price per model, but competition between hosts is a main reason open models are cheap. Second, licences differ: Kimi K3 uses a custom licence and MiniMax-M3 a community licence, while Qwen3.8 27B is Apache 2.0 and DeepSeek V4.1 Flash is MIT [1]. Check them before shipping a model with a game. GLM-5.3 is marked open by AA and LMArena but "closed weights" in Epoch's data, which we could not reconcile [1].
How fast prices are falling
Anyone budgeting a game that ships in a year should ask what these prices will be then. The published research agrees on the direction and disagrees on the rate.
Epoch AI measured the price to reach a given benchmark level and found declines between 9 and 900 times per year depending on the milestone, with roughly 40 times per year for GPT-4-level performance on GPQA Diamond; its data ran to February 2025, and the authors warn the fastest declines were recent and may not persist [8]. a16z's Guido Appenzeller calls it "LLMflation": for equivalent performance the cost fell about 10 times per year, with MMLU 42 costing $60 per million tokens with GPT-3 in November 2021 and $0.06 with Llama 3.2 3B three years later, a drop of around 1,000 times; he also cautions the method is imperfect and the rate may slow [9]. A November 2025 paper, "The Price of Progress" (arXiv 2511.23455), reports a figure of roughly 5 to 10 times per year for frontier models at fixed benchmark performance. We have only seen a search summary of that paper and did not open it, so treat the number as unconfirmed [10].
We built our own rough series from AA's data. For each AA Intelligence Index threshold, we took the lowest blended list price (3:1) among models released up to each date. This is an editorial estimate, with serious limitations: it uses today's list prices for old models, which understates what they cost at launch, and the index values for pre-2026 models are AA back-fills, with the index rescaled between versions.
Our own estimate. Uses CURRENT list prices of old models (understates launch prices) and AA index values that are back-filled for pre-2026 models; AA rescales the index between versions. Illustration of the trend, not a measured price history. Source: Artificial Analysis leaderboard, derived by CreateGame.ai, accessed 2026-10-09.
What the series shows, for what it is worth. At index 20 or above, the cheapest qualifying model went from $3.50 blended (o3, April 2025) to $0.10 (Hy3-preview, April 2026): 35 times in 1.02 years. At index 30 or above, from $4.81 (GPT-5.2 Xhigh, December 2025) to $0.175 (MiMo-V2.6-Flash, September 2026): 27.5 times in 0.78 years. At index 10 or above: $28.88 (o1-preview, September 2024) to $0.06 (Qwen3.5 4B, March 2026): 481 times over 1.47 years. Annualised, those work out to roughly 33, 71 and 67 times per year. At index 40 or above, the cheapest model went from $10 (Claude Opus 4.7, April 2026) to $0.20 (Claude Haiku 5.5, October 2026): 50 times in about six months, which we do not annualise because a window that short is mostly noise. These rates sit inside Epoch's 9-to-900 range, and above a16z's 10, which is the expected outcome given that all three studies cover different benchmarks and periods. We would not forecast from our series; we offer it as consistent evidence that capability gets cheap quickly.
The flip side is that the price of the best model is not falling. Opus 4.7 was $10 blended in April 2026 by our series and Opus 5.5 is $8 now; Fable 5.1 and GPT-6 Astra sit at $20 blended [1]. Vendors have used the savings to sell more capability at the same price rather than the same capability for less. For a game developer that means two strategies: pin a capability level and wait for it to get cheap, or budget for the top model at today's prices and treat savings as upside. If you want the broader picture of progress, see how fast AI is improving.
One caution about promotional prices. Gemini 4 Argon's $2 and $10 is introductory and press reports a later $4 and $20 (we did not verify that against Google's blog), and Gemini 3.8 Flash's price doubles in 2027 per Google's page [3]. A price you build a business model on can move against you as well as for you.
Putting a budget together
Putting the pieces together for three imaginary projects. These are illustrations built from our assumptions above, not benchmarks, and none is a quote for a real game.
| Project | AI cost elements | Estimated total |
|---|---|---|
| Weekend game-jam browser game | One agentic session on Sonnet 5.5 or Sol ($11 to $12), 100 images on a mid-priced model (about $3 to $4 at $0.034 to $0.039 per image, four tries per keeper) | Under $30 |
| Month of evening work on a small 2D game | $100 coding plan, 1,200 image generations ($12 to $253) | $112 to $353 |
| Same game plus live NPC chat for 10,000 daily players | Add $420 to $1,320 a month on a cheap model, with caching and low reasoning effort | $530 to $1,670 first month |
The conclusion we draw from the three: the AI cost of building is rarely the problem; the AI cost of running is. A solo developer can build for a few hundred dollars. The cost that grows with success is runtime inference, and it can be managed by picking the cheapest model that holds a character together, caching the persona prompt, capping message counts, and moving trivial barks to pre-written lines.
What is missing from all of the above is your own time, review of generated output, art direction, testing and the cost of rework when a model's code is wrong. Every account of building games with agents that we read says that a human still tunes game feel and checks the result [13]. We have no cost figure for that and are cautious about anyone who gives you one.
For context, here is where we sit. CreateGame.ai is a pre-launch product that aims to turn one sentence into a playable browser game, and cost per game is a question we think about because the inference behind a world, a story and a playable build has to be cheap enough to offer. We have no numbers to share yet.
For model-level detail, see the pillar post on the best AI models for game development and our review of AI game makers, which includes what some products charge.
How we reviewed this
We reviewed published sources: pricing pages, leaderboards, papers and public write-ups. We did not run the workloads, so every dollar figure in the worked examples is an editorial estimate from stated token assumptions at list prices, not a bill from a real project. The assumptions (200 turns, 60,000 cached tokens per turn, 1,200-token NPC prompts, a 25% image hit rate) are ours and can be changed in the arithmetic we show.
- Prices are from Artificial Analysis's model leaderboard as of 9 October 2026, checked against Anthropic's and Google's pricing pages, which we read directly [1][2][3]. OpenAI's page returned HTTP 403; OpenAI prices are AA's and LMArena's, which agree. Prices for Chinese open-weights models (Kimi, GLM, DeepSeek, MiMo, Qwen) are AA's hosted listings, not checked at vendor pages.
- Gemini 4 Argon is pre-release with introductory pricing; its later price is from secondary reports.
- Plan prices are from vendor pages we read on 9 October 2026, except Google AI Ultra and Cursor's higher tiers, which come from press and third-party summaries. Plan limits are not expressed in tokens, so we cannot convert them to API-equivalent dollars.
- Epoch AI and a16z pages were read; the arXiv paper (2511.23455) was not opened. Epoch's page text was truncated in our reader, so we cite only the figures we saw (the 9-to-900 range and the 40 times for GPQA Diamond).
- Our price series uses current list prices and AA back-filled index values; it is an illustration.
- GPU prices come from comparison blogs and retailer scans that disagree with each other; the RTX 5090 range is wide for that reason. GPU throughput for large models was not available.
- Local-run speeds for Qwen3.8 27B come from community and tutorial write-ups read via search summaries; one tutorial's 142 tokens per second is a best case.
- AA mean cost per task for coding agents is API-equivalent and the data does not break it down by token type.
- GPT-6.1 Sol's low cost per task is as reported by AA and unexplained.
FAQ
How much does it cost to make a game with AI? The AI costs for a small game are usually tens to a few hundred dollars: our estimate is $11 to $25 for a heavy agentic coding session, $12 to $253 for 1,200 image generations, and a $20 to $200 monthly plan if you code regularly. Costs that grow with players come from live features such as NPC chat.
How much does an LLM API cost per million tokens? In October 2026, from $0.10 input and $0.50 output (Claude Haiku 5.5, GPT-6 Luna) to $10 and $50 (Claude Fable 5.1, GPT-6 Astra). Mid-tier models like Claude Sonnet 5.5 and GPT-6.1 Sol are $2 and $10.
Is a Claude Code or Codex subscription cheaper than the API? For heavy use, usually yes at list prices: by our estimate, a $100 plan equals about nine API sessions of $11.70 each. Plan limits are given in time windows, not tokens, so the break-even is approximate.
How much does AI NPC dialogue cost? By our estimate, 10,000 daily players with 20 exchanges each cost about $960 per month on a Haiku-class model, $420 with prompt caching, and $19,200 on Sonnet 5.5. Reasoning tokens can more than double the cheap-model figure.
Are AI prices dropping? For a fixed level of capability, yes: Epoch reports 9 to 900 times per year depending on the benchmark and a16z about 10 times per year. The price of the best available model has not dropped.
Is it cheaper to self-host an open model? Rarely for coding. A rented H100 is about $1,825 a month if left on, and a consumer GPU now costs several times its launch price, for a weaker model. It makes sense for privacy or for running dialogue on the player's own hardware.
If you want to follow what we are building, join the waitlist at https://creategame.ai.
Sources
- Artificial Analysis, model leaderboard (prices, Intelligence Index v4.3.2), accessed 2026-10-09: https://artificialanalysis.ai/leaderboards/models
- Anthropic, API pricing (model prices, caching, batch, long context, fast mode), accessed 2026-10-09: https://platform.claude.com/docs/en/about-claude/pricing
- Google, Gemini API pricing, accessed 2026-10-09: https://ai.google.dev/gemini-api/docs/pricing
- Claude pricing page (individual plans), accessed 2026-10-09: https://claude.com/pricing
- ChatGPT and Codex plans, usage limits and credits, accessed 2026-10-09: https://learn.chatgpt.com/docs/pricing
- Google, "Everything new in our Google AI subscriptions, fresh from I/O 2026" (via search summary): https://blog.google/products-and-platforms/products/google-one/google-ai-subscriptions
- Cursor pricing page, accessed 2026-10-09: https://cursor.com/pricing
- Epoch AI, "LLM inference prices have fallen rapidly but unequally across tasks", 12 March 2025: https://epoch.ai/data-insights/llm-inference-price-trends
- Guido Appenzeller, a16z, "Welcome to LLMflation", 12 November 2024: https://www.a16z.com/llmflation-llm-inference-cost
- "The Price of Progress" (arXiv 2511.23455), not opened, via search summary: https://arxiv.org/abs/2511.23455
- Artificial Analysis, Coding Agents Index, accessed 2026-10-09: https://artificialanalysis.ai/agents/coding-agents
- Artificial Analysis, text-to-image leaderboard, accessed 2026-10-09: https://artificialanalysis.ai/image/leaderboard/text-to-image
- Decrypt, "Dumbest AI prompt: Claude beat careful game design", 28 July 2026: https://decrypt.co/374560/dumbest-ai-prompt-claude-beat-careful-game-design
- H100 rental price comparisons, 2026 (search summaries of comparison blogs): https://www.gmicloud.ai/en/blog/how-much-does-it-cost-to-rent-nvidia-h100-gpus-in-the-cloud and https://www.spheron.network/blog/runpod-h100-pricing-2026/
- RTX 5090 price reports, Aug-Sep 2026 (search summaries): https://spawningpoint.com/article/nvidia-rtx-5090-review-2026 and https://gpupoet.com/gpu/learn/price/september-2026/nvidia-geforce-rtx-5090
- DataCamp, "How to run Qwen3.8-27B locally" and community guides (search summary): https://www.datacamp.com/tutorial/how-to-run-qwen3-8-27b-locally
- LMArena WebDev leaderboard, snapshot dated 2026-10-08: https://arena.ai/leaderboard/code/webdev