How Fast Is AI Improving? The Evidence, and What It Means for Games
How fast AI is improving depends on which ruler you hold up, and the rulers themselves are wearing out. We went through the leaderboards, METR's time-horizon work, Epoch AI's compute and capability data and the open-versus-closed gap, then asked a narrower question that matters to people who make games: which capability thresholds have actually been crossed, and which are still ahead.
Key takeaways
- On Artificial Analysis's Intelligence Index (v4.3.2), the best model to date went from 24.9 in September 2025 to 57.6 in September 2026, a gain of 32.7 points against 13.5 in the previous twelve months. Epoch's separate index (ECI) also rose faster in the latest year (about 17 points against about 14), but much less dramatically. Two independent rulers agree on the direction and disagree on the size of the acceleration.
- METR's headline finding, that the length of tasks an AI agent can finish doubles every seven months, was revised toward roughly four months in 2026 by METR's own updated suite and independent re-analyses. METR also says measurements above 16 hours are unreliable with its current tasks, and its public page says it is no longer actively updated as of 8 September 2026. The best-known measure of agent progress has run out of road.
- Open-weights models trail the frontier by roughly three to four months on Epoch's index. Our own calculation on AA's index gives a median lag of about 4.3 months over the past two years, and about 3 months for the open-weights records set in 2026.
- Benchmarks die quickly. Epoch's staff describe most as saturating within months, a February 2026 preprint finds nearly half of 60 major benchmarks saturated, and the SWE-bench Verified runs Epoch published stop at about 83.5% in April.
- For game development, the thresholds that fell in the last 18 months are "write a playable small game from a prompt" and "finish a long turn-based game with a harness." The ones that have not fallen are real-time play from pixels, reliable engine-editor work (the best agent solves 53.8% of GameDevBench tasks) and a verified, consistent 100-hour story.
- Our labelled expectation for the next 12 to 24 months: more of the same on code and long-horizon agents, open weights within a few months of the frontier, and little progress on the memory and state problems that decide whether an AI-made game is any good after hour ten.
How fast is AI improving? Two scoreboards, one direction
Start with the two most-cited composite indexes, because they let you see the whole curve in one glance. Both combine many individual benchmarks into one number per model. They are built differently, scaled differently and updated on different schedules, so they get separate charts here. Put them on the same axis and you invite readers to compare numbers that mean different things.
Artificial Analysis (AA) re-scored every model it tracks on its Intelligence Index v4.3.2, which blends ten evaluations. That includes models released long before the index existed, and AA flags those older scores as back-filled estimates. Treat anything before about December 2025 as AA's retrospective estimate, not a contemporary measurement. Our chart takes the running maximum by release date: each point is a model that beat every earlier one.
Points before about December 2025 are AA back-filled estimates, not original-era measurements. Each step is a new record, so the line only moves up by construction. MiMo-V2.6-Pro and GLM-5.3 open-weights status should be checked against the license before you rely on it. Do not read this on the same axis as Epoch's ECI. Source: Artificial Analysis, Intelligence Index v4.3.2 (running maximum derived by CreateGame.ai), accessed 2026-10-09.
The shape is a staircase that gets steeper. GPT-4 in March 2023 sits at 6.7 on this scale. Claude 3 Opus in March 2024 is 8.7. By September 2024, OpenAI's o1-preview reaches 11.4, the first reasoning model to move the line noticeably. Twelve months later, in September 2025, the frontier is at 24.9 (GPT-5 Codex), and twelve months after that Claude Opus 5.5 stands at 57.6, released on 22 September 2026. Over the last year alone the line gained 32.7 points, compared with 13.5 the year before. That is an acceleration by AA's measure, and a large one.
We would not take the size of that acceleration at face value. The index is rescaled between versions (cached AA pages showed GPT-6 Astra at 55 to 60 while the live v4.3.2 data we pulled shows 52.7), so a number from last quarter and a number today may not be the same unit. The evaluations inside it get replaced as they saturate, which tends to keep the scale from flattening and can inflate late-period gains. And a running-maximum series is monotonic by construction, so it cannot show a slowdown unless releases stop altogether. What the chart does support is a plain statement: the best model available in October 2026 is much better by AA's tests than the best model of a year ago.
Epoch AI's Capabilities Index (ECI) is the independent check. Epoch built it by fitting a statistical model across roughly 40 benchmarks and about 200 models, giving each model one latent capability number and each benchmark a difficulty and a slope (Epoch, ECI methodology). The design matters: benchmarks near 0% or 100% carry almost no signal, so the method leans on whichever tests a model is currently in the middle of. It needs no new evaluations and no human votes, which is why it can be backdated from Epoch's own hub.
Epoch labels GLM-5.3 closed weights (AA labels it open), so it is absent from the open series. Gemini 4 Argon is not yet in ECI. Epoch itself warns that newer open models may not yet have enough evaluations to be scored. Source: Epoch AI benchmark data hub, ECI (CC BY 4.0), accessed 2026-10-09.
On ECI, the frontier moved from 135.8 (o1-mini, September 2024) to about 150 (GPT-5, August 2025) to 167.3 (Claude Opus 5.5, September 2026). That is roughly 14 points in the first year and 17 in the second. Still faster, but a modest acceleration of about 20% where AA shows 140%. We do not know which one is closer to the truth. Our guess is that both are right about different things: AA's mix of agentic and long-context tests has been picking up recent gains that older academic benchmarks no longer register, and ECI, which down-weights saturated tests by design, is smoothing the same gains. A reader who wants a single sentence should say "progress has not slowed and probably sped up in the last year, by an amount that depends on how you measure it."
What a single index hides
A composite number is an average of things that improve at different speeds. On Humanity's Last Exam, Epoch's collected results top out at 54.8% for GPT-6 Astra and AA's page shows 61.4% for Opus 5.5, scores that would have looked out of reach when the exam was built (Epoch HLE table via our snapshot). On AA's long-context test (AA-LCR), 25 of the 33 models in our snapshot score 0.79 or higher, so it no longer separates the field. On GDPval-AA, which scores work products, the spread is still wide: 0.683 for Opus 5.5 against 0.054 for the August 2025 open model gpt-oss-120b. Some capabilities are flat across vendors while others are still opening gaps.
METR and the doubling of task length
If the indexes tell you how good models are, METR's time-horizon work tries to tell you how long a job they can do alone. METR measures the length of task, in human-expert time, at which a model succeeds half the time. The original paper estimated that this horizon had been doubling about every seven months for six years, with Claude 3.7 Sonnet at roughly one hour (METR, March 2025). The same post noted that models then succeeded almost always on tasks under about four minutes and under 10% of the time on tasks over about four hours.
The 2026 updates changed the picture in three ways, and each deserves its own sentence.
First, the pace looked faster. METR's Time Horizon 1.1 added longer tasks and dropped flawed ones. Independent analyses of the updated data we read through summaries put the recent doubling time near four months, with one fit since 2023 giving about 129 days and others between 3.5 and 4 months (third-party analysis on GreaterWrong). Our own back-of-envelope agrees in spirit: if a one-hour horizon in early 2025 became roughly 12 hours for Claude Opus 4.6 a year later (METR's March 2026 note gives 11 hours 59 minutes), that is about 3.6 doublings in 12 months, or one every 3.4 months. Use a noise-corrected figure of 7 hours 38 minutes from the same note and you get about 2.9 doublings, or one every 4.1 months. These two measurements come from different versions of the task suite, so we treat the result as a rough range, not a rate.
Second, the measurement got shaky exactly where it got interesting. METR's note on modelling assumptions reports that reasonable alternative curve fits move the 50% horizon by about 1.5x and the 80% horizon by about 2x, and that confidence intervals often span a factor of two in each direction. Counting every sub-15-minute task as a success cut Opus 4.6's horizon from 12 hours to 9. Restricting to private tasks dropped it from 11h59m to 7h11m. The task suite has about 230 tasks from roughly 80 families, none longer than 30 hours (METR, 20 March 2026).
Third, the instrument is retiring. METR's time-horizons page says that measurements above 16 hours are unreliable with the current suite and that the page is no longer actively updated as of 8 September 2026 (METR time horizons). Search summaries report an early Claude Mythos Preview snapshot at 16 hours or more with a 95% interval of 8.5 to 55 hours; we have not read that figure on METR's own page and would not lean on it. Nothing we found gives a METR figure for the models released since mid-2026, including Opus 5.5, GPT-6 Astra or Fable 5.1.
What does the doubling work say for games? Taken at face value, agents have gone from tasks that take a human minutes to tasks that take most of a working day, which matches what builders report (implementing a feature, wiring a menu, verifying the result). It does not say a model can design a game: METR's tasks are software, ML and cybersecurity jobs with checkable outcomes, and fun is not a checkable outcome.
The compute behind it
Models get better for a few mundane reasons, and compute is the largest. Epoch's trends dashboard, updated 5 February 2026, puts the growth of training compute for frontier language models at about 5x per year since 2020, a doubling every 5.2 months (90% interval 4x to 6x). The total compute stock of AI chips is growing about 3.4x per year. Chip performance per dollar is up about 49% a year since 2023 (Epoch, Trends in AI). Epoch's decomposition work attributes the compute growth mainly to larger training clusters, with longer runs and better chips contributing less (Epoch, training compute decomposition, read via search summary).
Software matters as well. Epoch estimates that the same performance can be reached with roughly one third of the compute each year; its ECI paper puts the figure nearer six times when fitted differently. Those two numbers disagree, and we mention both because the uncertainty about how much progress comes from "more compute" versus "better recipes" is one of the biggest uncertainties in any forecast. If it is mostly compute, progress tracks capital and power, and Epoch notes that a gigawatt-scale data center takes about two years to build and roughly $38 billion up front. If it is mostly algorithms, progress can continue on cheaper hardware, and open-weights labs can keep up.
Release cadence by lab
Faster models only matter if you can use them, and the cadence of releases has shortened visibly. The table below lists the models from Anthropic and OpenAI that set records on AA or Epoch since mid-2025, with the gap in days to the previous model on the list. It is not a complete release history, and the gaps depend on which models we counted, so read it as a rough rhythm.
| Lab | Model | Release date | Days since previous |
|---|---|---|---|
| Anthropic | Claude Opus 4.5 | 2025-11-24 | - |
| Anthropic | Claude Opus 4.6 | 2026-02-05 | 73 |
| Anthropic | Claude Opus 4.7 | 2026-04-16 | 70 |
| Anthropic | Claude Opus 4.8 | 2026-05-28 | 42 |
| Anthropic | Claude Fable 5 | 2026-06-09 | 12 |
| Anthropic | Claude Opus 5 | 2026-07-24 | 45 |
| Anthropic | Claude Fable 5.1 | 2026-09-01 | 39 |
| Anthropic | Claude Opus 5.5 | 2026-09-22 | 21 |
| OpenAI | GPT-5 | 2025-08-07 | - |
| OpenAI | GPT-5.2 | 2025-12-11 | 126 |
| OpenAI | GPT-5.3 Codex | 2026-02-05 | 56 |
| OpenAI | GPT-5.4 | 2026-03-05 | 28 |
| OpenAI | GPT-5.5 | 2026-04-23 | 49 |
| OpenAI | GPT-5.6 (Sol/Luna/Terra family) | 2026-07-09 | 77 |
| OpenAI | GPT-6 Astra | 2026-09-03 | 56 |
| OpenAI | GPT-6.1 Sol | 2026-09-29 | 26 |
Not a complete release list: it leaves out Sonnet and Haiku models, Codex-only and Pro-only variants, and anything neither AA nor Epoch tracked. Median gap: Anthropic 42 days, OpenAI 56 days (our arithmetic over the gaps shown). Source: Artificial Analysis and Epoch AI release dates (via CreateGame.ai snapshot), accessed 2026-10-09.
Anthropic's median gap in this list is 42 days; OpenAI's is 56. The pattern is more telling than the median. Anthropic's first two intervals in 2026 were 73 and 70 days, then 42, then 12 (Opus 4.8 on 28 May and Fable 5 on 9 June), then 45, 39 and 21. OpenAI went from a 126-day gap (GPT-5 to GPT-5.2) to gaps of 28, 49 and 77 days and, most recently, 56 and 26 days (GPT-6 Astra on 3 September and GPT-6.1 Sol on 29 September). In the 15 days from 22 September to 7 October, Anthropic shipped Opus 5.5, Sonnet 5.5 and Haiku 5.5, and OpenAI, Google and Mistral each shipped something as well. Counting only AA record-setters, there were 4 in the first half of 2025, 7 in the second half, 6 in the first half of 2026 and 3 so far in the second half of 2026.
Shorter gaps do not mean bigger steps: the largest single jump on AA's index in this period is Claude Fable 5 on 9 June (41.8 to 49.6, a 7.8-point gain), and most are smaller. They do have a cost for builders, though: a feature tied to one model snapshot in spring has seen three newer ones since.
Benchmarks die fast
Everything above rests on benchmarks, and benchmarks are the weakest link. Epoch's staff have said on its podcast that most AI benchmarks saturate "in months," and that a very good one may last a year or two (Epoch, "Are AI benchmarks doomed?"). A February 2026 preprint analysing 60 benchmarks from major developers' technical reports found nearly half showing saturation, with the rate increasing as benchmarks age, and defined saturation as the loss of the ability to tell top models apart (arXiv 2602.16763, read via search summary). The older Papers With Code study of 3,765 benchmarks found much the same pattern in vision and language (Koch et al., arXiv 2203.04592).
You can see it in the coding data that games depend on. Epoch's own runs of SWE-bench Verified top out at 83.5% for Claude Opus 4.7 (16 April 2026), with GPT-5.5 at 80.6%, and the runs stop there. Newer models have not been run, partly because the test is near its ceiling. AA's index dropped the older Terminal-Bench variants and now uses Terminal-Bench 4.0, a harder version. The tbench.ai 2.0 leaderboard shows the best agent-plus-model submission at 84.7% on GPT-5.5. When your top scores are above 80%, the test no longer tells you about the margin that matters.
Benchmark names are also not stable units: "Terminal-Bench" has at least three meanings in our snapshot, AA's pre-2026 scores are retroactive estimates, and Epoch's SWE-bench Verified values do not exist for the latest models, so we never mix them in one chart. The most informative evaluations are the ones that do not exist yet: a score of 8% on a new agentic test says more than 97% on an old one. For game work, what remains is a handful of Godot tasks, a few Pokémon runs and a very low score on real-time play (see the coding post and the game-playing post).
The open-versus-closed gap
Open-weights models are the other half of the story, because they set the floor on what any game can use for free or run locally. They are also where the data is cleanest, so we can say more here than for most topics.
Epoch published two versions of the answer. An October 2025 data insight found that between January 2023 and October 2025 the best open-weights model trailed the best closed model by about three months on average (Epoch, via search summary). A May 29, 2026 follow-up found that from January 2026 the lag on ECI had risen to about four months, an average gap of about eight ECI points, which Epoch compares to the distance between GPT-5 and GPT-5.5 (Epoch, open-closed ECI gap). Epoch flags that the current gap may be overstated because many new open models lack enough evaluations to be scored. Its separate analysis of models small enough for a single consumer GPU finds they catch the frontier in 6 to 12 months (Epoch Brief, via search summary).
We computed our own version on AA's index, shown below. For each open-weights record, we looked up the first model of any kind that scored at least as high, and counted the months between the two release dates.
Our calculation from AA's v4.3.2 running-maximum series. It depends on AA's back-filled estimates for pre-2026 models and on AA's open-weights flag. Epoch's own ECI-based figure is about 4 months since January 2026 and about 3 months for 2023 to October 2025. Source: Artificial Analysis leaderboard, our computation, accessed 2026-10-09.
The chronological shape is more useful than the average. In 2025 the open records trailed by four to eight months, with a high of 8.5 months for DeepSeek V3.1. From early 2026 the lag falls to between 1.2 and 4.4 months: GLM-5 at 2.8, GLM-5.2 at 3.4, Kimi K3 at 1.2 and MiMo-V2.6-Pro at 3.4. The median across all 16 records is about 4.3 months; across the seven records of 2026 it is 3.4. The Kimi K3 figure is the extreme: released 16 July, it matched the level the closed frontier first reached with Claude Fable 5 on 9 June.
Two things temper that. First, AA and ECI disagree on how big the gap is in absolute terms. On AA the best open model (MiMo-V2.6-Pro at 46.3) sits 11.3 points below the frontier (57.6), a ratio of 0.80, up from 0.56 at the end of 2024. On ECI the best open model with a clear open-weights label (Kimi K3 at 157.5) is 9.9 points below Opus 5.5, essentially the same distance as at end-2024 (9.6) and end-2025 (9.1). A constant ECI gap with a rising AA ratio is what you get when two indexes weight different capabilities; AA's score gap grew in points because the whole scale stretched, even as the lag in months shrank. Time lag is the better measure and the one Epoch uses.
Second, open-weights status is not always clean. AA and LMArena label GLM-5.3 open (MIT license, per LMArena), while Epoch labels it closed. MiMo-V2.6-Pro is flagged open by AA. Kimi K3 uses a custom license and MiniMax-M3 a community license. If you plan to ship a model inside a game, read the license text, as we do in the cost post.
For game makers the takeaway is practical. Whatever the frontier can do in a given month, an open model you can host or fine-tune will probably do about as well within a season. That has consequences for games that embed models (the "which model do we depend on" decision becomes a price question), for studios worried about vendor risk, and for anyone building on the premise that a capability will stay exclusive.
What the thresholds are for game development
Index points are not game features. To connect the curves to things a game maker cares about, we picked seven capability thresholds and checked what the published record says about each: when it was crossed, by whom, and how solid the evidence is. This is a judgement table, not a benchmark, so the "status" column is ours.
| Threshold | Where it stands (Oct 2026) | Evidence and date | How solid |
|---|---|---|---|
| Write a small playable game from a short prompt | Crossed, repeatedly | Pieter Levels' fly.pieter.com, built with Cursor in early 2025 and reported at about $50,000 a month in revenue by March (404 Media); Cursor Vibe Jam 2026 entries, with about 1 in 7 multiplayer games using Colyseus (Colyseus) | Strong for small games, anecdotal for quality |
| Run a long, multi-agent code project for days | Crossed for non-game code | Anthropic's 16-agent C compiler: nearly 2,000 sessions, about $20,000, roughly 100,000 lines, Claude Opus 4.6, February 2026 (Anthropic) | One well-documented case, with criticism of completeness |
| Work inside a game engine editor | Partly crossed | GameDevBench (Godot): best agent and method solved 53.8% of 333 tasks, 2D graphics tasks about 33% (arXiv 2602.11103); Unity and Unreal both ship MCP routes (our engines post) | Single benchmark, three revisions with different numbers |
| Finish a 20 to 40 hour turn-based game with a harness | Crossed | Gemini 2.5 Pro beat Pokémon Blue (May 2025), Gemini 3 Pro beat Crystal (Dec 2025), Claude Opus 4.7 beat Red (May 2026) (our play post) | Single runs, harness-dependent |
| Play real-time games from raw pixels | Not crossed | VideoGameBench: best result 0.48% of the full set and 1.6% of the Lite set (arXiv 2505.18134) | Benchmark is from 2025; newer models unreported in what we found |
| Keep a 100-hour story consistent | Not demonstrated | Context windows have reached 1M tokens on most frontier models and AA-LCR scores exceed 0.8; no benchmark tests narrative consistency over a game's length | Our inference |
| Design a game that is fun without human direction | Unknown | No measure exists | None |
A few of those need explanation.
Writing a small game. The easy one. Models that produced a working Snake clone in 2023 can now produce a multiplayer 3D flight game or a roguelike, often from one prompt. The caveat is the one every builder knows: "playable" is not "good." We found no benchmark that scores how fun a generated game is, and the human-vote leaderboard for generated apps (LMArena WebDev, led by Claude Opus 5.5 at 1813 Elo) rates visual and functional quality of web apps. Our web games post covers where such games ship.
Running an engine. The public data is weakest here and practical progress probably largest. GameDevBench's authors found that agents struggle and that multimodal feedback helps: GPT-5.4 rose from 41.1% to 52.0% when given screenshots. The benchmark has three published revisions (v1 reported 54.5% over 132 tasks, the ICML version 53.8% over 333, and another listing 49.0%). Engine work lags plain code work, and graphics lags logic (33.0% against 51.4% for gameplay tasks). Our coding post goes through it.
Playing games. Pokémon runs are the nearest thing to a long-horizon game test, and they are good at one thing: showing how many hours of play a model can hold together. The progression is suggestive. A model beat Pokémon Blue with a custom harness in May 2025; a year later Claude Opus 4.7 beat Red, and Anthropic says Claude Fable 5 beat FireRed using only vision in June 2026. But real-time pixel play remained at roughly 1% on VideoGameBench, and the Pokémon harnesses differ run to run. Progress on the long turn-based side is real; progress on the fast, perceptual side is, in the published record, absent.
Holding 100 hours of story together. This is the threshold we most want to discuss because it is the one a game with generated stories needs most and the one nobody has measured. The ingredients exist: 1M-token context is standard on frontier models in our snapshot (32 of the 33 models have 250,000 tokens or more; only gpt-oss-120b has 131,072), AA-LCR scores above 0.8 are common, and the fiction-comprehension benchmark Fiction.liveBench showed near-perfect scores at 16,000 tokens for GPT-5 and o3 Pro in late 2025 (Epoch's mirror, via search summary). But recall of facts in a long passage is not the same as consistency in a story that the model itself is writing and the player is changing. As an estimate: if a player sees 2,000 words of text per hour for 100 hours, that is 200,000 words, or roughly 270,000 tokens at 1.35 tokens a word. That fits in a 1M window with room to spare, so storage is not the obstacle. The obstacle is that each added token raises the chance of a contradiction, and nobody publishes the rate. This is one of the few places where we think a new benchmark would be more useful than another model.
Our expectations for the next 12 to 24 months
These are our own forecasts, reasoned from the curves above and not taken from any source. Prices at a fixed capability should keep falling and the best model's price should not (see the cost post); we do not forecast specific prices. Real-time play from pixels will probably improve slowly, and a jump from 1.6% to 10% on VideoGameBench would not mean mastery. The meaningful signal would be a model that plays a new action game at novice level without a game-specific harness.
1. Code and agent capability keep rising at roughly the recent pace. If frontier compute keeps growing about 5x per year and efficiency about 3x, the next year's frontier should be well above today's. We expect the AA index to keep climbing at something like the rate of the last twelve months, with considerable uncertainty and a real possibility of slower measured gains as the index's evaluations saturate. We put 70% on "the best model on AA's index a year from now scores at least 70 on a version of the index comparable to v4.3.2" and say plainly that this is a judgement call, not a calculation. The arithmetic behind it is simple: 57.6 plus a gain between the year-before rate (+13.5) and the latest-year rate (+32.7) gives 71 to 90.
2. The engine-editor threshold is crossed for common tasks. GameDevBench's best result of 53.8% involves agents with feedback loops. We would be surprised if the best published result a year from now were below 70%, because the same mechanism (let the model see the game) is cheap to add and harness vendors are working on it. Unity's docs say its AI tools are not a way to prompt a finished game into existence, and Epic's MCP server is still experimental in Unreal Engine 5.8. We expect both statements to age, but not to be reversed, within a year.
3. Open weights stay within three to five months. The 2026 lag figures are between 1.2 and 4.4 months. A reasonable expectation is a median of about four months and individual records occasionally within weeks. That implies that any capability a frontier lab demonstrates in 2027 is available to a small studio on a rented GPU by roughly mid-year, subject to licenses and hardware.
4. Long-term consistency stays hard. We expect context windows to stay at or around 1M tokens and expect more progress from memory systems, summarisation and state tracking than from the model alone. This is a prediction about where the work happens, not about whether it will be done. A game that tracks facts in a database and asks the model to write within them will outperform one that relies on the model remembering.
What would change our minds
- A sustained slowdown on ECI. If Epoch's frontier gains drop below about 8 points per year for two consecutive quarters, our "same pace" forecast is wrong, and the AA acceleration was mostly a measurement artefact.
- An open-weights gap above six months. That would mean the closed labs had found something that does not spread through publication, hiring or distillation. We would watch whether Epoch's measured lag, now about four months, widens.
- A real-time play result without a custom harness. Even a modest one would make our fifth expectation too pessimistic.
- Training runs stalling on power or data. Epoch's own numbers show a gigawatt-scale site takes about two years to build. A visible stall in announced clusters would slow the frontier more than anything on the algorithm side.
- A metric for narrative consistency. If one appears and models do well, we will move the 100-hour story into the "crossed" column. If models do badly, we will rewrite the paragraph above.
What this means if you are making a game
Two practical suggestions fall out of the evidence. Design around the capability that has actually moved: code generation, long-horizon coding agents and long-context reading are good enough to change how a small team prototypes, while real-time perception, consistency over long play and judgement about fun are not. And plan for a model swap: with releases every few weeks, the favourite of March is not the favourite of October, so keep models behind an abstraction and re-run a small evaluation of your own tasks on each release. At CreateGame.ai, where we are building a one-sentence game creator that is not yet released, nothing we describe depends on one model staying best.
If you want to follow our work as these curves move, the waitlist is at creategame.ai.
How we reviewed this
We reviewed published sources only: leaderboards, papers, Epoch and METR pages, vendor pricing and documentation, and write-ups of builds. We ran no models, tests or builds ourselves and make no claim of hands-on testing.
- Cross-model data comes from the shared snapshot we assembled on 9 October 2026 from Artificial Analysis (Intelligence Index v4.3.2), LMArena and Epoch AI's benchmark hub (ECI, CC BY 4.0). AA's running-maximum series and the open-model lag are our own derivations from AA's data. AA's pre-2026 scores are AA's retroactive estimates.
- AA and Epoch are never shown on a shared axis. They are separately scaled and not numerically comparable.
- Epoch trend numbers (5x per year compute, 3.4x chip stock, the 4-month open-weights lag) were read from Epoch's pages on 9 October 2026. The October 2025 three-month lag figure, the consumer-GPU result, the training-compute decomposition and the saturation preprint were read through search summaries; we did not open every underlying page.
- METR: we read the March 2025 post, the March 2026 modelling-assumptions note and the time-horizons page. We did not open the Time Horizon 1.1 release post. The "about four months" doubling comes from third-party analyses read through summaries, and the Mythos Preview figure from a summary. We found no METR data for models released after about May 2026.
- Our arithmetic. The 3.4 to 4.1 month doubling range uses a one-hour horizon in early 2025 and METR's 11h59m and 7h38m figures from different suite versions; it is an illustration, not a METR result. The 100-hour story token estimate (200,000 words, about 270,000 tokens) and the 2027 forecasts are our estimates.
- Not verified: GLM-5.3's open-weights status conflicts between sources. Gemini 4 Argon is pre-release. GameDevBench has three published versions with different headline numbers. Pokémon and Vibe Jam results are self-reported. We did not read OpenAI's pricing page (HTTP 403).
FAQ
How fast is AI improving in 2026? By Artificial Analysis's Intelligence Index, the best model gained 32.7 points between September 2025 and September 2026, against 13.5 the year before. Epoch's index shows a smaller acceleration (about 17 points against 14). Both say progress did not slow.
Is AI progress slowing down? We found no evidence of a slowdown in the frontier data through early October 2026. Many individual benchmarks have flattened because they are saturated, which can look like a slowdown when it is not. METR's measure has hit its own ceiling, so one of the best long-task signals has gone quiet.
How far behind are open-weights models? Epoch measures about four months on its Capabilities Index since January 2026, and about three months from 2023 to October 2025. On AA's index we compute a median lag near 4.3 months over two years and about 3.4 months for 2026's records.
Can AI make a whole game today? It can make a small, playable one, often from a prompt, and several have shipped to the web. It does not reliably design a large game, and the best agent solved 53.8% of Godot tasks on GameDevBench. Fun and long-term consistency have no benchmark.
Sources
- Artificial Analysis, model leaderboard and Intelligence Index v4.3.2, accessed 2026-10-09: https://artificialanalysis.ai/leaderboards/models
- Artificial Analysis, Claude Opus 5.5 model page (index version, component evals), accessed 2026-10-09: https://artificialanalysis.ai/models/claude-opus-5-5
- Epoch AI, benchmark data hub (ECI, SWE-bench Verified runs, Terminal-Bench 2.0 mirror, HLE), accessed 2026-10-09: https://epoch.ai/benchmarks
- Epoch AI, Epoch Capabilities Index (methodology page): https://epoch.ai/blog/capabilities-index
- METR, "Measuring AI Ability to Complete Long Tasks" (19 March 2025): https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
- METR, time horizons page (accessed 2026-10-09; says not actively updated as of 8 September 2026): https://metr.org/time-horizons
- METR, "Impact of modelling assumptions on time horizon results" (20 March 2026): https://metr.org/notes/2026-03-20-impact-of-modelling-assumptions-on-time-horizon-results/
- "METR Time Horizons: Now 10x/Year" (GreaterWrong, third-party analysis, read via search summary): https://www.greaterwrong.com/posts/EYb2K9acKfyG2bome/metr-time-horizons-now-10x-year
- Epoch AI, Trends in AI dashboard (updated 5 February 2026): https://epoch.ai/data/trends
- Epoch AI, "Training compute growth is driven by larger clusters, longer training, and better hardware" (via search summary): https://epoch.ai/data-insights/training-compute-decomposition
- Epoch AI, open-versus-closed ECI gap (29 May 2026): https://epoch.ai/data-insights/open-closed-eci-gap
- Epoch AI, "Open-weight models lag state-of-the-art by around 3 months on average" (via search summary): https://epoch.ai/data-insights/open-weights-vs-closed-weights-models
- Epoch Brief, "What's the best model you can run on a single consumer GPU?" (via search summary): https://epochai.substack.com/p/whats-the-best-model-you-can-run
- Epoch AI, "Are AI benchmarks doomed?" (podcast page): https://epoch.ai/epoch-after-hours/are-ai-benchmarks-doomed
- "When AI benchmarks plateau: a systematic study of benchmark saturation," arXiv 2602.16763 (via search summary): https://arxiv.org/pdf/2602.16763v1
- "Mapping global dynamics of benchmark creation and saturation in artificial intelligence," arXiv 2203.04592: https://arxiv.org/pdf/2203.04592
- Epoch AI, "LLM inference prices have fallen rapidly but unequally across tasks" (12 March 2025): https://epoch.ai/data-insights/llm-inference-price-trends
- Chi et al., GameDevBench, arXiv 2602.11103: https://arxiv.org/abs/2602.11103
- Anthropic Engineering, "Building a C compiler with a team of Claudes": https://anthropic.com/engineering/building-c-compiler
- 404 Media, "This Game Created by AI Vibe Coding Makes $50,000 a Month" (5 March 2025): https://404media.co/this-game-created-by-ai-vibe-coding-makes-50-000-a-month-yours-probably-wont
- Colyseus, "Vibe Jam 2026: 1 in 7 multiplayer games shipped with Colyseus": https://colyseus.io/blog/vibe-jam-2026/
- VideoGameBench, arXiv 2505.18134: https://arxiv.org/abs/2505.18134
- Epoch AI, Fiction.liveBench mirror (via search summary): https://epoch.ai/benchmarks/fictionlivebench
- Anthropic, API pricing, accessed 2026-10-09: https://platform.claude.com/docs/en/about-claude/pricing
- Sibling post, Best AI for Coding Games: /blog/best-ai-for-coding-games/
- Sibling post, Can AI Play Video Games?: /blog/can-ai-play-video-games/
- Sibling post, How Much Does It Cost to Make a Game with AI?: /blog/cost-of-making-a-game-with-ai/
- Sibling post, Unity and Unreal's AI Tools: /blog/unity-unreal-ai-tools/
- Sibling post, Web Games in 2026: /blog/are-games-moving-to-the-web/