CreateGame.ai

Reports

World Models and the Future of Game Development: What Genie and Its Rivals Can Actually Do

A world model draws the next frame of a game from your button presses, with no engine, no level file and no rules written by a human. We went through the published specs, papers, demos and executive statements to ask what that means for the future of game development, and our answer is narrower, and stranger, than either the hype or the stock-market panic suggested.

Key takeaways

  • As of October 2026 the best-documented interactive world models run at 20 to 40 frames per second at up to 720p, keep a scene coherent for somewhere between one and a few minutes, and cost enough that Google caps its public Project Genie sessions at 60 seconds.
  • A game is a set of rules that must behave the same way twice. Today's world models keep state implicitly inside pixels, which is why they are good at looking like games and poor at being them.
  • The money is moving toward robotics and driving simulation. Decart, Odyssey and NVIDIA have all pointed their newest flagship models at physical AI, not at players.
  • Our thesis: over the next three to five years world models become an upstream layer (ideation, previs, asset and environment generation, texture and mood passes) feeding engines, not a replacement for them. Engines are absorbing the same models from the other side.
  • The genuinely new thing may not be a faster way to make Elden Ring. It may be a new kind of short, personalised, non-deterministic interactive media that today's genre labels do not fit.

What a world model is, and why games keep coming up

A world model, in the sense used by every lab in this article, predicts the next state of an environment given the previous states and an action. In practice the "state" is a video frame. You press forward, the model draws what the camera would see one step later, and then it feeds its own output back in as context for the next step. DeepMind's lead researcher Jack Parker-Holder described it to heise as closer to a language model than a classic video model, because it generates frame by frame and causally instead of producing a whole clip at once (heise, May 2026).

That autoregressive loop is the source of both the magic and the problems. Odyssey states the issue plainly: feeding generated output back in makes the system drift outside its training data, and a world model's state space is far larger than text's (Odyssey, May 2025). The short sessions, forgotten objects and cost below all follow from that one design choice.

Games became the public face of the technology for a mundane reason: gameplay footage is abundant and comes labelled with the player's actions. The original Genie in early 2024 was trained on platformer footage and produced roughly one frame per second (Wikipedia summary of Ars Technica's reporting). GameNGen, a paper from Google researchers in August 2024, trained a diffusion model on DOOM and ran it at 20 frames per second on a single TPU, with human raters only slightly better than chance at telling short clips from the real game (arXiv 2408.14837). Games are the easiest place to show that a network can hold a world together for a while, and the hardest place to show it can do so for good.

The state of play in October 2026

Google DeepMind: Genie 3 and Project Genie

The timeline is short enough to list. Genie 1 (early 2024) made 2D worlds at about one frame per second. Genie 2 (December 2024) added 3D but produced worlds of 10 to 20 seconds at 360p, according to the Ars Technica and PCMag reporting summarised on Wikipedia. Genie 3, announced on 5 August 2025, generates navigable worlds in real time at 24 frames per second and 720p, "retaining consistency for a few minutes," in DeepMind's own wording (DeepMind).

DeepMind's launch post is unusually frank about what Genie 3 cannot do. The list is worth reading as a game developer would: a limited action space for the agent itself, no reliable modelling of multiple independent agents interacting, no accurate real-world locations, text that is only legible when it was in the prompt, and interaction that lasts "a few minutes" instead of hours (DeepMind). Every one of those is a prerequisite for a shipping game.

On 29 January 2026 Google put Genie 3 behind a consumer wrapper called Project Genie, available to US adults on the Google AI Ultra plan. It combines Genie 3 with Nano Banana Pro for the preview image and Gemini, and it caps each generated world at 60 seconds. Google's own page lists the shortcomings: worlds may not look realistic or follow the prompt, characters can be harder to control with higher latency, and the "promptable events" shown in the August research preview are not in the prototype (Google). TechCrunch reported that the 60-second limit exists partly because an autoregressive model needs a lot of dedicated compute, which puts a ceiling on how much access DeepMind can hand out (TechCrunch, Jan 2026).

At Google I/O on 19 May 2026 the team added a Street View integration, so you can start from a real location instead of a prompt. US locations came first (TechCrunch, May 2026; heise). We did not find a Genie 4 or any new Genie model announced as of 9 October 2026. Several secondary sites claim September updates, but none of them linked a primary source, so we are not relying on them.

The most revealing detail in the heise interview is where the Genie team's attention lies. Parker-Holder and product manager Diego Rivas told the outlet that the researchers are thinking primarily of robotics and disaster simulation, not of games on demand, and Waymo has already built a driving simulator, the Waymo World Model, on top of Genie 3 (heise; Wikipedia on the Waymo model). The same interview contains the single most useful sentence in this whole debate. Asked about the competitive landscape, the team said that compared with language models, "we are in 2021" (heise). We will come back to what that implies, because 2021 for language models meant GPT-3 and a lot of impressive-but-unreliable text, and a lot of people were wrong about what came next in both directions.

Microsoft: Muse (WHAM) and the Quake II demo

Microsoft's contribution is the one most carefully built for game developers, and also the one most visibly constrained. Muse, published in Nature in February 2025 with Ninja Theory, is a "World and Human Action Model" that can generate game visuals, controller actions, or both. The 1.6-billion-parameter version was trained on Bleeding Edge, a multiplayer game, on over a billion image-and-action pairs representing roughly seven years of continuous human play, at 300 by 180 pixels (Microsoft Research; Nature). Microsoft open-sourced the weights and a demonstrator tool.

What makes the Nature paper interesting is its framing. The authors argue that for a model to support creative work it needs three properties: consistency, diversity and persistency, the last meaning that edits you make to a scene stay in the scene (Microsoft Research). That is a design-tool vocabulary, not a player-product vocabulary. Muse was pitched as a way to explore variations on a game that already exists. We find that the most credible near-term use of the whole category.

In April 2025 Microsoft followed with WHAMM, a browser demo of Quake II generated in real time. Coverage described 640 by 360 resolution, just over 10 frames per second, about a week of footage from one level, a 120-second play limit, laggy controls, and enemies and scenery that vanished when the player looked away (The Decoder; we have not retested it). Microsoft framed it as a demonstration of potential. We found no Muse or WHAMM successor in 2026.

Odyssey: the clearest published scaling story

Odyssey, a lab founded by self-driving veterans, has published the most legible progression. Odyssey-1 (May 2025) streamed a new frame every 40 milliseconds for five minutes or more but was described by its makers as "a glitchy dream" with limited utility (Odyssey). Odyssey-2 Pro (23 January 2026) added a developer API and streams 720p at 22 frames per second (Odyssey). Odyssey-2 Max (21 April 2026) is roughly three times larger and, by the company's own measurement, raised its VBench 2 physics score from 49.67 to 58.52 (Odyssey).

That last number deserves skepticism. The benchmark table sits in the company's own launch post, includes values "from our own measurements," and compares against three Cosmos-Predict variants and a gaming model, LingBot-World-Fast, which has no VBench 2 physics score listed. A physics score from a video-quality benchmark is also not a game developer's notion of "the physics works." Still, the direction is consistent: bigger model, better dynamics, same real-time constraint, which is Odyssey's argument that scale will help here as it did for language.

Physics score on VBench 2, as reported by OdysseyBigger Odyssey models score higher, but this is the lab's own table
Physics score on VBench 2, as reported by Odyssey0102030405060Odyssey-2 Max58.5Odyssey-2 Pro49.7Odyssey-248.6Cosmos-Predict2.5-14B44.9Cosmos-Predict2-14B39.2Cosmos-Predict2.5-2B35.6Model

Vendor-published table that mixes values from the VBench 2 paper and Odyssey's own measurements. LingBot-World-Fast is listed in the source but has no VBench 2 physics score. Not independently reproduced. Source: Odyssey, Introducing Odyssey-2 Max (21 Apr 2026), accessed 2026-10-09.

The strategic detail comes later. Odyssey raised $310 million at a $1.45 billion valuation in June 2026, led by Natural Capital with Amazon, AMD Ventures, GV and In-Q-Tel participating (Pulse 2), and in mid-September it unveiled Odyssey-3, which trade coverage describes as a foundation model for controlling robots, humanoids, vehicles and drones, with a public release promised "in the coming weeks" and no pricing or license published at announcement (The AI Insider). A lab whose first public product was a playable, glitchy dream world now leads with humanoid control policies.

Decart: from a Minecraft clone to driving data

Decart's first model, Oasis, was among the early real-time playable world models and went viral as a Minecraft-like demo (Maginative). The company's newest flagship, Oasis 3 (10 June 2026), is an API-first model for physical AI that starts with autonomous vehicles. Decart says it runs at 22 frames per second at 512 by 768 with under 200 milliseconds of latency, and it generates a synchronised three-camera driving view (Decart). TechCrunch reported a price of $0.02 per second and a recent $300 million raise at a valuation near $4 billion, and its reviewer noted that although you can interact for hours, the model "degrades significantly" the longer a world runs (TechCrunch, June 2026).

That price is the first real unit-economics number any of these companies has published, so let us do the arithmetic, and flag that this is our estimate and for a driving model, not a game model. At $0.02 per second, one minute costs $1.20 and one hour costs $72. A traditional game session costs the studio effectively nothing in marginal compute, because the player's own console or PC does the rendering. Even if we assume prices fall tenfold, a $7 hour of play would be a cost no free-to-play economy or $70 premium title can currently carry for every player. Cloud game streaming has fought that fight for fifteen years and is still a niche.

Open and Chinese models: Matrix-Game and others

The open-weights side is where we would watch for developer-facing tooling. Skywork's Matrix-Game series released 1.0 in May 2025, 2.0 in August 2025 and 3.0 on 27 March 2026, under the MIT license. The 3.0 repository claims 720p at 40 frames per second with a 5-billion-parameter model, a camera-aware memory meant to keep scenes consistent over minute-long sequences, and training data mixing Unreal Engine synthetic data, automated AAA game data and real-world video (Skywork on GitHub). Those are the authors' claims; we have not run the model and found no independent benchmarks. Note the pipeline: engines generate the data that teaches the world model, which then imitates the engine.

Dynamics Lab's Mirage 2 (August 2025) claimed more than ten minutes of continuous generation, around 200 milliseconds of latency and single-GPU operation, which would beat Genie 3 on session length. Those claims came through press coverage of a research preview, and one hands-on review found the controls did not feel smooth (Gigazine). Tencent's HunyuanWorld-Voyager (September 2025) is also open source and takes a camera-trajectory approach to 3D-consistent video (CG World).

World Labs, Runway and NVIDIA

World Labs takes a different bet. Marble, generally available since 12 November 2025, does not stream frames at all. It builds a persistent 3D world from text, images, video or coarse layouts and exports it as Gaussian splats, meshes (including low-fidelity collider meshes for physics) or video (World Labs). That output can be dropped into an engine, which is the point. The company raised $1 billion in February 2026, with Autodesk contributing $200 million (Reuters via Yahoo Finance). On 1 September 2026 it announced Atlas, an "omni" world model in early access with partners; the company published no paper, model card or code with the announcement, and the performance claims are its own (Implicator).

Runway's GWM Worlds 2, a research preview announced on 3 September 2026, streams playable worlds at 720p and 24 frames per second with generated audio and no preset session length. Access is gated and we found no pricing, with the only detail on limits coming from secondary coverage that mentions degraded texture and geometry under fast camera motion and imperfect long-term memory (Gradually). NVIDIA's Cosmos 3, released on 1 June 2026, is an open "omnimodel" for physical AI: robotics, vehicles, simulation. NVIDIA describes it as a model for robotics, vehicles and simulation, not for games (NVIDIA via GlobeNewswire).

Side by side

We gathered the published figures into one table. Every figure comes from the vendor or the cited coverage and none was independently measured. A dash means we found nothing published.

Model (date) Resolution Frame rate Coherence / session Access
Genie 2 (Dec 2024) 360p n/a 10 to 20 seconds Not public
Muse WHAM-1.6B (Feb 2025) 300 x 180 n/a Several minutes, in research examples Open weights
WHAMM, Quake II (Apr 2025) 640 x 360 above 10 120-second demo cap Browser demo
Odyssey-2 Pro (Jan 2026) 720p 22 Minutes Paid API
Genie 3 / Project Genie (Aug 2025 / Jan 2026) 720p 24 "A few minutes" in research; 60 s in product US, AI Ultra plan
Matrix-Game 3.0 (Mar 2026) 720p 40 (5B model) Minute-long consistency claimed Open weights, MIT
Oasis 3 (Jun 2026) 512 x 768 22 Hours, with degradation API, $0.02 per second
GWM Worlds 2 (Sep 2026) 720p 24 No preset length Gated preview
Advertised frame rate of interactive world modelsMost sit at 20 to 25 fps; Matrix-Game 3.0 claims 40 with a 5B model
Advertised frame rate of interactive world models010203040Matrix-Game 3.0 (5B)40Genie 324Runway GWM Worlds 224Odyssey-2 Pro22Decart Oasis 322GameNGen (DOOM)20WHAMM (Quake II)10Model

All values are vendor or paper claims, not independent measurements. WHAMM was reported as 'above 10' fps and is plotted at 10. Resolutions differ: 720p for Genie 3, Odyssey-2 Pro, Matrix-Game 3.0 and GWM Worlds 2; 512x768 for Oasis 3; 640x360 for WHAMM. Sources: DeepMind, Odyssey, Decart, Skywork, Runway (via Gradually), The Decoder, arXiv 2408.14837. Source: DeepMind Genie 3 announcement and vendor posts (see note), accessed 2026-10-09.

The frame-rate chart shows near parity: nearly everyone sits between 20 and 25 frames per second, with Matrix-Game's 40 the outlier. Frame rate is the metric the field has mostly solved. The bottleneck is elsewhere.

How long an interactive world lastsAdvertised coherence and enforced caps, in seconds. Orders of magnitude matter more than ranks
How long an interactive world lasts050100150200250300Genie 2 (upper bound of 10-20 s)20Project Genie session cap60WHAMM demo cap120Oasis 3 per-call timeout120Odyssey-2 family (120+ s reported)120Odyssey-1 (5+ minutes reported)300Model or product

The metrics are not equivalent: Genie 2 and Odyssey values are advertised coherence, Project Genie and WHAMM are enforced product caps, and Oasis 3's value is a per-call API timeout reported in a technical write-up (TechCrunch notes interaction can run for hours but quality degrades). Genie 3's research preview claimed 'a few minutes' of consistency, which has no exact figure and is not plotted. Source: Google Project Genie blog and vendor posts (see note), accessed 2026-10-09.

The second chart shows the real constraint. The numbers are not strictly comparable (some are advertised coherence, some enforced caps, and Oasis 3's is a per-call API timeout), so read it as orders of magnitude, not a ranking. Everything sits between about 20 seconds and a few minutes, and the product consumers can actually touch enforces 60 seconds. A game session is tens of minutes at the low end, and a save file lasts weeks.

What world models cannot do yet

They have no ground truth

The sharpest critique we found comes from the people who run engine companies and the investors who back them, and it is a technical argument as much as a commercial one. Tim Sweeney said on X that world models "are limited by a lack of memory, and by the fact that memory is in an inefficient and expensive format," adding that the ideal outcome is a merger in which the AI side supplies loosely organised audiovisual knowledge and the engine supplies "consistent reproducible data representation and simulation" (Mobilegamer.biz). Andreessen Horowitz partner Jonathan Lai made the same point from the other direction: games are deterministic, world models are probabilistic, and neither player nor developer knows what happens until it does, which makes them "poorly suited for traditional games" (Mobilegamer.biz).

Two 2026 survey papers say the same in academic language. One organises the field around the loop a conventional engine runs (action, explicit game state, observation) and finds that world models struggle with rule-following, consequences that persist beyond the current view, and effects at rule-defined moments, all of which "revolve around the game state, which most models keep implicit" (arXiv 2607.14076). A broader survey finds evaluation well established for bounded game-playing but "persistent state in learned worlds" less so (arXiv 2609.16679). GameNGen's authors said as much in 2024: the model saw a little over three seconds of history, and a bigger context window gave only marginal gains (arXiv 2408.14837).

Consider what this means for design. A shooter needs ammunition to be a number. An RPG needs a quest flag to stay set for forty hours. A competitive game needs two clients to agree on where a grenade landed. In an engine those are variables. In a world model they are whatever the next frame happens to imply.

The demos are good at the wrong half of a game

Naavik's Miikka Ahonen described Project Genie in February as an impressive research demo whose worlds break down after a minute, run at 720p and 24 frames per second with noticeable input lag, and show objects that "subtly change" while the model "forgets" what it established (Naavik, Feb 2026). His central point is the one we keep returning to: games are defined not by the worlds they depict but by the rules they enforce, and the appeal of a ranked ladder or an enemy pattern depends on repeatability, "which is exactly what world models don't do." He also notes that tools that do generate rule-bearing games from prompts, which are not world models, have produced hundreds of games without a breakout hit.

The cost problem has no engine-shaped answer

We covered the Oasis 3 arithmetic above. Genie 3's per-session cost is unpublished and we will not guess one. The structural point stands: an autoregressive model streaming frames to every player consumes inference for the whole session, while a conventional game's marginal cost per player-hour is close to zero after the sale. That is why we are skeptical of a world-model live-service game, and less skeptical of a short, premium experience such as an exhibit, an ad or a training scenario.

The reaction from studios and markets

When Project Genie launched, game stocks sold off. Mobilegamer.biz reported Take-Two and Roblox down 10 to 12 percent and Unity down about 20 percent in a day; other outlets published bigger Unity figures, so the exact number depends on the source and the window (Mobilegamer.biz; Naavik). Naavik called it a category error: the market behaved as if a replacement engine had been announced when the product was an experimental demo.

The executive responses were consistent. Unity CEO Matthew Bromberg called Genie-style output "unsuitable on their own for games that require consistent, repeatable player experiences" and said Unity converts world-model output into "structured, deterministic, and fully controllable simulations" (Mobilegamer.biz). Take-Two's Strauss Zelnick found the sell-off puzzling, said generative AI has no part in GTA 6, and argued that making a hit is "a completely different animal" from making an asset, as reported by TechSpot and The Game Business; VGC's headline called the idea of push-button hits "laughable".

Engine vendors and publishers have an obvious interest in telling investors their position is safe. But the technical arguments underneath are independently stated by researchers and in DeepMind's own limitations list, so we think the substance holds even if the tone is defensive. Unity and Epic are also building generative AI into their products, which we cover in what Unity and Unreal are shipping.

The developer mood is less sanguine. The 2026 GDC State of the Game Industry survey (more than 2,300 respondents, fielded November to December 2025) found 52 percent of professionals think generative AI is having a negative impact on the industry, up from 30 percent a year earlier and 18 percent before that, with 7 percent calling the impact positive (GDC). The question is about generative AI in general, not world models, but it describes the audience any world-model product must win over.

Where the money is going

Look at what the biggest rounds and newest model releases of 2026 were for.

Funding rounds announced by world-model companies in 2026The big checks are going to labs whose newest flagships target robotics, driving and 3D tooling
Funding rounds announced by world-model companies in 2026$0$200$400$600$800$1,000World Labs (Feb 2026)$1,000Odyssey Series B (Jun 2026)$310Decart (spring 2026)$300Company

Valuations: Odyssey $1.45B post-money (Pulse 2); Decart nearly $4B (TechCrunch); World Labs did not disclose (Reuters). Decart's round is described only as a few weeks before its 10 June 2026 launch. Source: Reuters via Yahoo Finance; Pulse 2; TechCrunch, accessed 2026-10-09.

World Labs raised $1 billion in February (Reuters via Yahoo). Decart raised $300 million in the spring (TechCrunch). Odyssey raised $310 million in June (Pulse 2). Against that, the newest flagship models from Decart (Oasis 3, driving), Odyssey (Odyssey-3, robots and vehicles) and NVIDIA (Cosmos 3, physical AI) are pitched at machines that learn to act, and the Genie team told heise that robotics is its main concern. Waymo uses a Genie derivative to rehearse tornadoes and elephants in the road.

The reasoning is not mysterious. A robotics team will pay for a simulator by the second because the alternative is hand-building environments, which Decart says takes "hundreds of expert hours" each (Decart), or crashing real hardware. The failure modes that hurt games (probabilistic output, drift, no hard rules) are tolerable when the goal is a diverse stream of plausible training scenarios. And those buyers have budgets. A player who gets bored after a minute does not.

This is what our headline claim rests on: the technology credited with the future of game development is, commercially, being built for something else, and games get what falls off the table. Language models were also built for other purposes and games benefited. But the roadmap will follow robotics and driving needs (physical accuracy, long-horizon stability, action conditioning, API access), while interactive entertainment needs hard rules, persistence, authorial control and near-zero per-player cost, and nobody with a billion dollars is solving those first.

Our thesis for the future of game development: an upstream layer, not a replacement

Here is the argued view. We label the forecasts as ours and give the reasoning.

Claim 1: World models will not replace engines for shipped, rules-driven games within five years. The reasoning is the three constraints above. State persistence is unsolved in a way that scale alone may not fix, because persistence is a data-structure problem and not a pixel-quality problem. The most promising 2026 research responds by adding explicit state: the survey we cite above organises the field around exactly that gap (arXiv 2607.14076), and we saw descriptions of hybrid proposals in which executable code holds the state while a video model only renders it (we did not read those preprints in full, so we do not rely on them). Once you add explicit state, you have rebuilt part of a game engine. Sweeney's "merging of world models and engines" is the same conclusion from the other side.

Claim 2: The real adoption path runs through ideation and content, where the engine stays in charge. Muse was framed that way from the start. Marble outputs meshes and splats that go into engines. Matrix-Game trains on engine-generated data. Unity's Bromberg describes world-model output being ingested into the engine. Naavik lists the use cases as visual exploration, mood setting and prototyping. This is the same shape as how image models entered game art: concept art and placeholders first, shipped assets later and unevenly. We would expect the sequence to repeat, with environment blockouts, skyboxes, background plates and look-development as first uses. For the image side of that story see our ranking of AI image tools for game art.

Claim 3: A new interactive format will appear, and it will not be called a game at first. What a world model does well is respond to a person plausibly and immediately in a place nobody built. That suits a few minutes of exploration, a personalised ad, a museum exhibit or a "what would my street look like in snow" toy, which is precisely the Street View pitch. Short, shareable generated interactive clips may be a bigger business than anyone expects. We hold this one with less confidence than the first two.

Claim 4: Timelines. These are our estimates, built on trends in the table above, not on any vendor roadmap.

Here is what would change our mind, so the claims can be tested:

What this means if you make games

For a solo developer or a small team, the practical conclusions are modest. Do not wait for a world model to make your game; nothing published suggests it will within your project's timeline. Do use them for reference environments, visual directions and walkable mood boards. If you use an engine, expect the vendor to bring the models to you. For the AI that is useful today, see our look at which models can code games and how fast AI is getting better. For AI as a player rather than a generator, see whether AI can play video games.

A deterministic engine plus a language model that writes the rules and content is a different bet from a pixel-generating world model, and it has working products today (AI game makers compared). CreateGame.ai, the pre-launch product behind this blog, is aimed at that second approach, so weigh our perspective accordingly.

How we reviewed this

We reviewed published material only: vendor blog posts and launch announcements from Google DeepMind, Microsoft Research, Odyssey, Decart, World Labs and NVIDIA; the Nature and arXiv papers cited; the GitHub repository for Matrix-Game 3.0; trade and general press coverage (TechCrunch, Mobilegamer.biz, The Game Business, Naavik, heise, Reuters via Yahoo Finance); and the GDC 2026 State of the Game Industry write-up. We did not run any of these models, we did not get Project Genie access, and we did not interview anyone. Every performance number is a vendor claim unless we say otherwise.

What we could not verify. The Runway GWM Worlds 2 and World Labs Atlas details come from secondary coverage and company statements without papers. We did not open Nature's full text, the HunyuanWorld-Voyager license or Mirage 2's primary materials, and we read only summaries of the Take-Two executive's remarks. The Odyssey VBench table is the company's own. Stock reactions are reported differently by different outlets. We found no Genie 4 announcement and treat reports of September 2026 Genie updates as unconfirmed. The $72-per-hour figure is our arithmetic from a driving model's per-second price and shows scale, not the price of a game.

FAQ

What is Google's Genie world model? Genie is DeepMind's family of interactive world models. Genie 3 generates navigable 3D-style environments from a text or image prompt at 24 frames per second and 720p, with consistency lasting a few minutes. Since 29 January 2026 Genie 3 has been available to Google AI Ultra subscribers through Project Genie, with each session capped at 60 seconds.

Can I make a game with Project Genie? Not in the sense of shipping one. You can create, explore and remix worlds and download videos, but there is no engine export, and worlds are neither persistent nor rule-bound. Google calls it a prototype.

Will world models replace game engines? We think not within five years. Engines hold explicit, deterministic state, and world models hold it implicitly in generated pixels, which causes drift and inconsistency. Epic's Tim Sweeney and Unity's Matthew Bromberg have both said they expect the two to converge, with the engine supplying reliable state.

What are Genie's main competitors? Odyssey (Odyssey-2 Max, Odyssey-3), Decart (Oasis 3), World Labs (Marble and Atlas), Runway (GWM Worlds 2), Skywork's open-source Matrix-Game 3.0, Microsoft's Muse, and NVIDIA's Cosmos 3. Most of the newest flagship releases target robotics and driving simulation instead of entertainment.

How much does it cost to run a world model? Only one price is public. Decart charges $0.02 per second for Oasis 3 via API, which works out to $72 per hour. Google has not published Genie's cost per session, but its 60-second cap was explained by compute constraints.

What is the future of game development with AI? In our reading: engines remain the backbone, language models write more of the code and content, and world models become an upstream tool for exploring and generating environments. A new short-form interactive medium is plausible, but we would not expect it to replace the games we know.

If you want to see where we are taking AI-built games, you can join the CreateGame.ai waitlist at https://creategame.ai.

Sources

  1. Google DeepMind, "Genie 3: A new frontier for world models" (5 Aug 2025): https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/
  2. Google, "Project Genie" (29 Jan 2026): https://blog.google/innovation-and-ai/models-and-research/google-deepmind/project-genie/
  3. TechCrunch, "I built marshmallow castles in Google's new AI-world generator Project Genie" (29 Jan 2026): https://techcrunch.com/2026/01/29/i-built-marshmallow-castles-in-googles-new-ai-world-generator-project-genie/
  4. TechCrunch, "Google's Genie world model can now simulate real streets with Street View" (19 May 2026): https://techcrunch.com/2026/05/19/googles-genie-world-model-can-now-simulate-real-streets-with-street-view/
  5. heise online, "Like 2021 with LLMs: Google researchers on the future of world models" (25 May 2026): https://www.heise.de/en/background/Like-2021-with-LLMs-Google-researchers-on-the-future-of-world-models-11304842.html
  6. Wikipedia, "Genie (world model)" (used for Genie 1/2 specs and the Waymo World Model, which cite Ars Technica and PCMag): https://en.wikipedia.org/wiki/Genie_(world_model)
  7. Naavik, "Project Genie and the Stock Market's Category Error" (8 Feb 2026): https://naavik.co/digest/project-genie-and-the-stock-markets-category-error/
  8. Mobilegamer.biz, "Unity boss Bromberg moves to calm markets as Google's Genie causes gaming stock slide": https://mobilegamer.biz/unity-boss-bromberg-moves-to-calm-markets-as-google-genie-causes-gaming-stock-slide/
  9. Microsoft Research, "Muse: Our first generative AI model designed for gameplay ideation": https://www.microsoft.com/en-us/research/blog/introducing-muse-our-first-generative-ai-model-designed-for-gameplay-ideation/
  10. Nature, "World and Human Action Models towards gameplay ideation" (2025): https://doi.org/10.1038/s41586-025-08600-3
  11. The Decoder, "Microsoft releases real-time AI-generated playable demo of Quake II" (Apr 2025): https://the-decoder.com/microsoft-releases-real-time-ai-generated-playable-demo-of-quake-ii/
  12. Valevski et al., "Diffusion Models Are Real-Time Game Engines" (GameNGen), arXiv 2408.14837: https://arxiv.org/abs/2408.14837
  13. Odyssey, "Introducing Odyssey-1" (28 May 2025): https://odyssey.ml/introducing-odyssey-1
  14. Odyssey, "The GPT-2 Moment for World Models Is Here" (23 Jan 2026): https://odyssey.ml/the-gpt-2-moment-for-world-models
  15. Odyssey, "Introducing Odyssey-2 Max" (21 Apr 2026): https://odyssey.ml/introducing-odyssey-2-max
  16. The AI Insider, "Odyssey Unveils Odyssey-3" (17 Sep 2026): https://theaiinsider.tech/2026/09/17/odyssey-unveils-odyssey-3-world-model-for-robots-humanoids-vehicles-and-drones/
  17. Pulse 2, "Odyssey Raises $310 Million" (17 Jun 2026): https://pulse2.com/odyssey-raises-310-million-to-accelerate-world-simulation/amp/
  18. Decart, "Introducing Oasis 3" (10 Jun 2026): https://decart.ai/publications/introducing-oasis-3-first-interactive-world-model-for-physical-ai
  19. TechCrunch, "Decart's new world model can simulate hours of photorealistic driving, with some caveats" (10 Jun 2026): https://techcrunch.com/2026/06/10/decarts-new-world-model-can-simulate-hours-of-photorealistic-driving-with-some-caveats/
  20. Maginative, "Decart Unveils Real-Time AI Game Engine": https://www.maginative.com/article/decart-unveils-real-time-ai-game-engine-raises-21m-to-make-ai-generated-worlds-more-affordable/
  21. World Labs, "Marble: A Multimodal World Model" (12 Nov 2025): https://www.worldlabs.ai/blog/marble-world-model
  22. Reuters via Yahoo Finance, "World Labs raises $1 billion" (18 Feb 2026): https://finance.yahoo.com/news/ai-pioneer-fei-fei-lis-202957884.html
  23. Implicator, on World Labs Atlas (Sep 2026): https://www.implicator.ai/world-labs-atlas-withholds-paper-price-partners/
  24. Skywork AI, Matrix-Game 3.0 README: https://github.com/SkyworkAI/Matrix-Game/blob/main/Matrix-Game-3/README.md
  25. NVIDIA via GlobeNewswire, "NVIDIA Launches Cosmos 3" (1 Jun 2026): https://www.globenewswire.com/news-release/2026/06/01/3303987/0/en/nvidia-launches-cosmos-3-the-open-frontier-foundation-model-for-physical-ai.html
  26. Gradually, "GWM Worlds 2" (Runway, Sep 2026, secondary): https://www.gradually.ai/en/ai-models/gwm-worlds-2/
  27. "From Pixels to States: Rethinking Interactive World Models as Game Engines," arXiv 2607.14076: https://arxiv.org/abs/2607.14076
  28. "AI for Games in the Foundation Model Era," arXiv 2609.16679: https://arxiv.org/abs/2609.16679
  29. The Game Business, "Take-Two CEO: Google's Project Genie AI is a great sign of things to come": https://www.thegamebusiness.com/p/take-two-ceo-googles-project-genie
  30. Video Games Chronicle, Take-Two CEO on Project Genie: https://www.videogameschronicle.com/news/take-two-ceo-says-its-laughable-to-say-ai-like-google-project-genie-can-create-hit-games-at-the-press-of-a-button
  31. Gigazine, "Mirage 2" (Aug 2025): https://www.gigazine.net/gsc_news/en/20250825-ai-generative-world-mirage-2
  32. CG World, HunyuanWorld-Voyager (Sep 2025): https://cgworld.jp/flashnews/01-202509-HunyuanWorld-Voyager.html
  33. GDC, "2026 State of the Game Industry reveals impact of layoffs, generative AI and more": https://gdconf.com/article/gdc-2026-state-of-the-game-industry-reveals-impact-of-layoffs-generative-ai-and-more/
  34. TechSpot, Take-Two CEO on generative AI: https://www.techspot.com/news/111193-take-two-ceo-company-embracing-generative-ai-but.html