Can AI Play Video Games? From Pokémon to Minecraft, What the Results Show
A year ago the honest answer to "can AI play video games" was "badly, and slowly, with a lot of help." In 2026 the answer has changed in some places and stayed stubbornly the same in others. This piece goes through the public record: the Pokémon livestreams, the academic benchmarks, the Minecraft work, DeepMind's SIMA line and the esports-era systems that came before.
Key takeaways
- Language models can now finish a long, turn-based game. Claude Opus 4.7 beat Pokémon Red in May 2026, Gemini 2.5 Pro beat Pokémon Blue in May 2025, and Gemini 3 Pro beat Pokémon Crystal (including the hidden boss Red) in December 2025. Anthropic says Claude Fable 5 beat Pokémon FireRed with a vision-only harness in June 2026.
- The same models are still near zero on real-time games with raw pixels. In the VideoGameBench paper the best result was 0.48% of the full benchmark and 1.6% of the Lite version.
- On the BALROG leaderboard, the top language model (GPT-6-Astra-Max) scores 68.3% progress, but the best NetHack score on the board is still 13.2%. The best vision-model entry on the same site is 35.7%, and it is from April 2025.
- Pokémon is a decent progress meter and a weak benchmark. Harnesses differ run to run, the games are heavily documented online, and nearly every result is a single run.
- The older esports systems (OpenAI Five, AlphaStar) were stronger at their one game than any LLM agent is at any game, but they were trained for years on that game. The new work is about generality.
- For game developers, the practical use today is testing, long-horizon simulation and NPC research, not "AI replaces the player."
Why games became the favourite agent test
Games are an unusually convenient exam. They have clear state and a score or a goal, and they punish mistakes in ways a grader doesn't have to interpret. They also test things that static question-answering benchmarks barely touch: remembering what happened four hours ago, noticing that a plan has silently stopped working, and acting under uncertainty where the next screen depends on your last button press.
The authors of VideoGameBench make the argument directly. Models are strong at coding and math, they write, but perception, spatial navigation and memory management remain understudied, and video games are designed to be intuitive for humans. A game is a test that was built to be learnable by a person with no manual. That is a different bar from a bar exam.
There is also a cultural reason. A livestream of a model walking into the same wall for four days is both a research artifact and entertainment, and it spreads. Claude, Gemini and GPT all have Twitch channels. That attention has pushed the labs to publish progress in game terms, which is how Pokémon became the most-watched informal agent benchmark of the last two years.
A caution before the numbers. Games differ in what they measure, and the harness (the code that feeds screenshots or game state to the model and turns its replies into button presses) often matters more than the model. We will flag this repeatedly, because it is the single most common way these results get misread.
The Pokémon saga
Claude: from three badges to a finished game
Anthropic's side project began as an internal experiment with a text harness around an emulator, and became a public stream with Claude 3.7 Sonnet. The early picture, as recounted in Julian Bradshaw's shortform on LessWrong, was grim: roughly 35,000 actions to get 3 badges with 3.7 Sonnet, and no further progress from Claude Opus 4, whose later run sat at about 380 hours and 54,000 actions without a fourth badge.
Bradshaw's December 2025 post, "Insights into Claude Opus 4.5 from Pokémon", states that Claude had made no real progress on Red since 3.7 Sonnet. Sonnet 4, Opus 4, Opus 4.1 and Sonnet 4.5 reached the same blocking points faster but did not get past them. Opus 4.5 was the first to break through Team Rocket Hideout and Erika's Gym, and at the time of writing it had used 48,854 reasoning steps over 300-plus hours. The post also lists failure modes that anyone who has watched the stream will recognise: ignoring visible objects outside its current focus (cut-able trees, spinner tiles), hallucinating a sought-after object onto nearby walls, and one wrong note in memory stalling progress for days. At one point Claude spent about four days and roughly 8,000 reasoning steps circling the roof of a gym.
A follow-up dated January 29, 2026 (by Josh Snider, on LessWrong) reports that Opus 4.5 had collected all eight badges after about 230,000 steps, then got stuck on the Victory Road boulder puzzles. In Bradshaw's May 2026 write-up, Opus 4.6 reached Victory Road in about 30,000 steps in its run but stalled for months on those same puzzles, and Opus 4.7 finally beat the game on 16 May 2026, "a year late" relative to the challenge Bradshaw had set. The post credits harness changes, including a zoom tool and a way to store screenshots as visual references, though it also notes the Elite Four took a couple of tries and the model remembered to buy healing items and revives.
Later in the same comment thread Bradshaw gives the headline figure for Opus 4.7: about 325 hours, with a lightly harnessed setup, against roughly 26 hours for an average human on Red (per HowLongToBeat, as he cites it).
Then Anthropic changed the game. On June 9, 2026 it released Claude Fable 5 and showed it beating Pokémon FireRed (the 2004 remake, not the original) with a minimal, vision-only harness: raw screenshots, no maps, no navigation aids, no extra game-state information. Earlier Claude models, the announcement says, struggled even with harnesses that gave them extra tools. Bradshaw's shortform puts the run at just over 50 hours, which is more than six times faster than the Opus 4.7 figure, though he adds that no full stream was provided, only two short edited videos, one of which was quickly taken down. Press coverage (NerdZap) describes a Charmander line that ended as a level 78 Charizard, and gives no hours, steps or costs. We therefore treat "50 hours" as an outside estimate of an unpublished run, not a measurement.
What it costs to try this yourself
A developer who tried to replicate the Fable launch demo published a useful reality check on DEV Community. Using about 800 lines of Python, two tools (press buttons, update notes) and a 2,000-turn cap, the run took about 6 hours 45 minutes of wall clock time, earned the first badge at turn 1,785 and cost $73.50 in total, around 3.7 cents a turn. The author's own extrapolation, which is an estimate and not a measurement, was that a full playthrough would take tens of thousands of turns and low four-figure dollars. Our arithmetic agrees with the shape of that estimate: at 3.7 cents a turn, 30,000 turns would be about $1,100 and 60,000 turns about $2,200. The author also lists what Anthropic did not publish (harness code, turn counts, costs, number of attempts) and notes the prior-knowledge confound: the model may know FireRed from training data. That is one developer's replication with a different ROM and a different harness, and the replication did not reach full completion, so it tells us about cost scale rather than capability.
Gemini: first to finish, and a cleaner comparison
Google's run got to the credits first. Gemini 2.5 Pro beat Pokémon Blue in May 2025, which Sundar Pichai announced. The streamer, Joel Zhang, has since written up the setup. By the figures in the LessWrong posts above, the first Blue run took 106,505 reasoning steps over 813 hours. (The shortform gives 816 hours for the same run, a small disagreement we leave as is.) A second run, with a more advanced scaffold, beat Blue in about 36,801 actions over 406 hours, and the Elite Four and Champion took seven tries.
The most informative single comparison in public is Zhang's December 2025 post, Gemini 3 Pro vs 2.5 Pro in Pokémon Crystal. Both models ran on the same harness: six tools including a mental map, notepad, map markers, code execution and custom agents. Gemini 3 Pro won the race. By the post's numbers it beat Red, the hidden final boss, at turn 24,178, using 1.88 billion tokens. Gemini 2.5 Pro was at its fourth badge at turn 38,204 when the race ended, and the author's projected finish for it was about 157,000 turns and over 15 billion tokens (a projection from its efficiency ratio, not an observed result). The Red fight lasted about 7 hours of real time; Gemini 3 Pro's party was one level 75 Typhlosion plus teammates at levels 8 to 19, facing Red's level 70 to 80 team, and the model devised a multi-stage plan around Smokescreen, Leftovers, PP exhaustion and a revive loop.
The author is careful about the limits: it is a single head-to-head run rather than a controlled benchmark, the harness was built mainly for 2.5 Pro, and turn and token comparisons are more reliable than time comparisons because of API downtime. We trust those caveats more than we trust the headline.
GPT: fewest steps, with an asterisk
OpenAI's GPT-5 finished Pokémon Red in 6,470 steps against 18,184 for o3, according to TechRadar's coverage of the stream, working out to roughly seven days versus over fifteen. It did so by leveling mostly one Pokémon. Bradshaw's Opus 4.5 post lists GPT-5.1 beating Pokémon Crystal in 9,454 reasoning steps over 108 hours using a minimal harness with a minimap, while the shortform mentions o3 beating Red in about 18,000 actions over roughly 388 hours with a custom harness.
The step counts across these runs are not comparable. A "step" in one harness can be a single button press and in another a multi-press plan; some harnesses give the model a minimap or pathfinding, others do not. Counting hours is a little better, but hours include API latency, which depends on the day and the provider.
One table, loudly caveated
| Run | Game | Result | Effort reported |
|---|---|---|---|
| Gemini 2.5 Pro, run 1 | Blue | Beat game, May 2025 | 106,505 steps, 813 h |
| Gemini 2.5 Pro, run 2 | Blue | Beat game | 36,801 actions, 406 h |
| o3 | Red | Beat game | about 18,000 actions, about 388 h |
| GPT-5 | Red | Beat game | 6,470 steps, about 7 days |
| GPT-5.1 | Crystal | Beat game | 9,454 steps, 108 h |
| Gemini 3 Pro | Crystal | Beat Red (boss), Dec 2025 | 24,178 turns, 1.88B tokens |
| Claude Opus 4.7 | Red | Beat game, May 2026 | about 325 h |
| Claude Fable 5 | FireRed | Beat game, June 2026 | unpublished; about 50 h per outside estimate |
Sources are linked in the text above. Different games (Red and Blue are nearly identical; Crystal is longer; FireRed is a remake), different harnesses, different step definitions.
Different games (Red/Blue, Crystal, FireRed) and different harnesses. Gemini run 1 is 813 h in LessWrong's Opus 4.5 post and 816 h in Bradshaw's shortform. The Fable 5 figure ('just over 50 hours') is Bradshaw's reading of two short edited videos; Anthropic published no hours. The human figure is as cited by Bradshaw. Source: Julian Bradshaw on LessWrong (Opus 4.5 post and shortform), accessed 2026-10-09.
The chart shows hours to finish where a source reports them. The pattern is clear even with the caveats: the time to finish has dropped by an order of magnitude in about a year, and a typical human is still faster than all of them, at around 26 hours for Red.
The run that did not use an LLM to play
One more result belongs here because it shows what the best games on this list may be telling us. According to Tom's Hardware, a developer says an AI decision model called Jev, which is not an LLM, beat Pokémon Red in under a week while Claude Opus 5 acted as a coach, rewriting option lists and diagnosing dead ends. A secondary source (remio.ai) gives 37 hours 40 minutes and 16,150 model decisions, and the secondary sources disagree on who built it and even which version was played. We could not read the full article, so treat this as a developer's claim with unclear details. Still, it matches a theme we return to below: the strongest game results increasingly come from a system, not from a model on its own.
Benchmarks that try to be fair
BALROG
BALROG (ICLR 2025) bundles existing reinforcement-learning environments of very different difficulty, from BabyAI gridworlds that take seconds, to NetHack, which takes experts years. The headline metric on the leaderboard is "% Progress," the average completion percentage across environments. The site says it is updated every Monday; we read it on 2026-10-09.
The language-model table now looks like this at the top:
| Model | % Progress | Entry date |
|---|---|---|
| GPT-6-Astra-Max | 68.3 | 2026-09-18 |
| Claude-Opus-5-Max | 63.4 | 2026-09-20 |
| GPT-5.6-Sol-Max | 60.0 | 2026-09-18 |
| Gemini-3-Pro | 58.1 | 2026-02-03 |
| Gemini-3.1-Pro-Thinking | 57.0 | 2026-02-25 |
| Gemini-3.1-Pro | 56.9 | 2026-02-21 |
| GPT-5.6-Terra-Max | 53.2 | 2026-09-20 |
| Gemini-3-Flash | 48.1 | 2026-02-13 |
Older entries give the trend. The leaderboard lists Claude 3.5 Sonnet at 32.6 and GPT-4o at 32.3 in November 2024, Gemini 2.5 Pro at 43.3 in April 2025, and Claude Opus 4.5 at 43.5 in February 2026. So the best score on the board roughly doubled in under two years, from the low 30s to 68.3.
Scores as shown on the leaderboard; the page also shows error margins we omit. Entries are not all run at the same time or harness version. The best NetHack score on the board is 13.2%. Source: BALROG leaderboard, accessed 2026-10-09.
Per-environment results are more revealing than the average. The top entries have saturated the easy ends: BabyAI is at 100% (three models tie), BabaIsAI at 100% (GPT-6-Astra-Max), Crafter at 76.8 and TextWorld at 75.7. MiniHack tops out at 65.0. NetHack, the hardest, tops out at 13.2. The leaderboard page does not define the NetHack score units, so we can't say how far 13.2 is from meaningful play, only that it is the lowest ceiling on the board by a wide margin. The average is therefore carried by games that are, for current models, mostly solved.
There is a second table for vision-language models, where the model sees an image of the environment instead of text. It has barely moved: the best entry is Gemini 2.5 Pro (March 2025 experimental) at 35.7, with Claude 3.5 Sonnet at 35.5. Nothing from 2026 is on it. We take that as a gap in the leaderboard, not as evidence about 2026 models, but it matches the pattern in the next section: pixels are the problem.
One more caveat about BALROG: we cannot independently confirm the newest entries, and the site gives no run counts or cost per entry. We would want those before leaning on small gaps like 57.0 versus 56.9.
VideoGameBench
VideoGameBench (Zhang, Griffiths, Narasimhan and Press; submitted May 2025, revised May 2026) asks models to play ten 1990s games in real time from raw screenshots plus a short description of controls and objectives. Seven games are public (Doom II, Kirby's Dream Land, Zelda: Link's Awakening, Civilization, The Need for Speed, The Incredible Machine, Pokémon Crystal) and three are secret, to reduce overfitting. Because inference latency wrecks real-time play, the authors added a Lite mode where the game pauses while the model thinks.
The results table in the paper (HTML version) is stark:
| Model | VideoGameBench | Lite |
|---|---|---|
| GPT-4o | 0.09% | 1.6% |
| Claude Sonnet 3.7 | 0.48% | 1.6% |
| Gemini 2.5 Pro | 0.48% | 1.6% |
| Llama 4 Maverick | 0% | 0% |
| Gemini 2.0 Flash | 0% | 0% |
| Qwen2.5-VL 7B | 0% | 0% |
| Qwen2.5-VL 32B | 0% | 0% |
Models are from spring 2025. The paper was revised in May 2026; we could not confirm whether newer models were added. Qwen2.5-VL 7B also scored 0 and is omitted. Source: Zhang et al., VideoGameBench (arXiv 2505.18134), accessed 2026-10-09.
On Lite, the only game with any progress was Kirby's Dream Land, at 4.8% for the three top models. Doom II and Link's Awakening stayed at 0% for everyone. The failure modes the authors describe are the ones the Pokémon streams already showed: a "knowing-doing gap" where the model recognises that the exit is at the bottom of the screen and keeps pressing down regardless of where the character is, misreadings of what is on screen, and scratchpad mistakes where overwriting notes sends the model back and forth.
The caveat is on our side. The paper's table contains models from spring 2025, and the arXiv page notes a revision in May 2026 without a visible changelog; we could not confirm whether newer models have been added. Given how fast the Pokémon results moved over the same period, we would not assume these zeros still hold for the newest frontier models. We do think the basic message survives: real-time play from pixels with no help was very hard for the 2025 generation.
Other benchmarks worth knowing
We looked at a few others without finding the kind of leaderboard that would justify a chart. Orak is a multi-game benchmark; in its Minecraft crafting task it reported o3-mini, Gemini 2.5 Pro and Claude 3.7 Sonnet tied at the top with a score of 75.0, while most open models scored zero. And the community keeps building "plays X" streams for Civilization V, Slay the Spire, Old School RuneScape and Factorio; Anthropic says in the Fable 5 announcement that giving Fable 5 file-based memory in Slay the Spire improved its performance about three times more than for Opus 4.8, and that it reached the final act about three times as often. That is Anthropic's claim from a vendor page and we have not seen the underlying data.
Minecraft: the open-ended test
Minecraft is the game most research groups pick when they want open-endedness rather than a fixed ending. The reference result for LLM agents is still Voyager (2023), which wrapped GPT-4 in an automatic curriculum, a growing library of executable skills and iterative self-verification. By its authors' numbers it obtained 3.3 times as many unique items, traveled 2.3 times farther and unlocked tech-tree milestones up to 15.3 times faster than the prior best. (An earlier version of the paper says 3.1 times for the first figure; the later version says 3.3.) The detail worth remembering is the method: Voyager writes code. It does not press keys; it calls an API.
A parallel line, Ghost in the Minecraft, reported a big jump on the ObtainDiamond task (+47.5% over prior methods, from a baseline success rate near 20%) using text-based knowledge and memory. And in 2025 and 2026 new benchmarks appeared: MineNPC-Task found about a 33% subtask failure rate for GPT-4o over 44 tasks, with errors clustered around preconditions and memory reuse; PillagerBench tests teams of LLM agents in competitive scenarios.
A different kind of Minecraft evaluation is MC-Bench, started by a high-school student: models are asked to build a structure from a prompt and humans vote blind on which build is better. It is a creativity and spatial-reasoning test, not a play test, and we did not find current standings.
What this body of work says about "can AI play Minecraft" depends on what you mean by play. Agents that call a code API can do well on resource-gathering and crafting chains, as in Orak. Agents that must look at pixels and operate a keyboard and mouse are a harder problem, which is where DeepMind's work sits.
DeepMind's SIMA line
SIMA 2, announced November 13, 2025, is the most direct attempt at the general problem: one agent, many commercial 3D games, instructions in natural language, keyboard and mouse like a human. The original SIMA, the blog notes, learned over 600 language-following skills. SIMA 2 is built on Gemini; per TechCrunch it runs on Gemini 2.5 Flash-Lite, and its training mixes human demonstrations with labels that Gemini generated.
The numbers are spread across secondary sources, and they do not agree on a single headline. The Decoder reports that on held-out environments (MineDojo and ASKA), SIMA 2 completed 45 to 75% of tasks versus 15 to 30% for SIMA 1. MarkTechPost reports about 62% on the main evaluation suite against roughly 70% for humans, and the original SIMA at around 31% (we could not open MarkTechPost's page ourselves, so those figures come from search summaries and should be checked). DeepMind's own post gives charts without values in its text.
Ranges as reported by The Decoder; DeepMind's blog shows the comparison as an unlabelled chart. Other outlets report about 62% for SIMA 2 on a different main suite (human reference about 70%), which is not comparable. Source: The Decoder, citing DeepMind, accessed 2026-10-09.
What DeepMind says about limits is more useful than the percentages: SIMA 2 has a "relatively short memory of its interactions," long-horizon multi-step tasks are open, and precise low-level keyboard and mouse actions remain a challenge. The more interesting claim is self-improvement. SIMA 2 can learn through trial and error in new environments, including worlds generated by Genie 3, with Gemini generating tasks and rewards. DeepMind says this works without further human-labeled data, but it gives no metric for it in the text. It is research-only; access is limited to a small group of academic and game-industry partners. If you want the world-model side of this story, our piece on world models and the future of game development covers Genie in more detail.
The old guard: OpenAI Five and AlphaStar
It is fashionable to say AI has only now started playing games, which would surprise the people who watched 2019. OpenAI Five beat the Dota 2 world champions, Team OG, on April 13, 2019, after about 10 months of training with a distributed system; according to Wikipedia's summary, the public online event that followed saw the bots play 42,729 games and win 99.4%. Training was self-play reinforcement learning with no human game data, at roughly 180 years of play per day on 256 GPUs and 128,000 CPU cores in 2018. The bots read about 20,000 numbers per frame from the game's developer API; they did not look at pixels.
AlphaStar reached Grandmaster rank in StarCraft II in 2019, placing in the top 0.2% of ranked players, under human-like constraints including a camera view and network latency, as reported at the time of the Nature paper.
These systems matter for context because they set a different kind of bar. They were superhuman or near it, but narrow: one game, a game-specific interface, years of compute. Nobody has asked OpenAI Five to play Pokémon. The new wave trades peak performance for breadth: one model that can be pointed at an unfamiliar game, with a prompt rather than a training run. That trade is why the 2026 results look so much worse in absolute terms and so much more interesting in kind. It is also why, to us, comparing them directly is a category error.
What the pattern says about agents
Perception is the bottleneck, and it is getting better
Bradshaw's Opus 4.5 post says the model "no longer has any trouble finding doors" and recognises gyms, centers and marts the moment they appear. The Opus 4.7 work added a zoom tool, and the Fable 5 demo reportedly works from raw screenshots alone. VideoGameBench, with its 2025 models, says almost the opposite: models cannot find the exit. Both can be true. Those are different model generations, and the Pokémon harnesses, being turn-based, give the model time to look.
Memory and note-keeping decide long runs
Every long run turns on the same thing: what the model writes down. The Claude posts describe one wrong note stalling progress for days. Gemini's harness has a notepad and map markers. The Fable 5 replication author used a free-form scratchpad and a summary of history every 20 turns, and Anthropic says memory made a larger difference for Fable 5 than for Opus 4.8 in Slay the Spire. As game benchmarks mature, they are turning into memory benchmarks.
The harness is part of the result
The harness is not a footnote. Bradshaw's account says Gemini 2.5 Pro beat Blue with a stronger harness, and the Gemini team later beat the game with progressively weaker harnesses. The Claude harness has had hints and a "Critic Claude" in the past, and later removed them, along with limits on memory files and button presses. Opus 4.5's reports of getting stuck for 8,000 steps are in part reports about what the harness lets the model see and do. When someone says "Model X beat the game in N hours," ask what was in the harness.
Training data is a confound we cannot rule out
All of these games have been walked through, in text, thousands of times. A model that "knows" how to get past the Silph Co. lock is not necessarily reasoning its way through. Bradshaw notes he has some reason for suspicion that earlier runs may have entered training data, but not enough information to be sure; a commenter argues the models already knew the general routes anyway. The Fable 5 replication author raises the same confound. Both observations are fair. The way around it is secret games, as in VideoGameBench's three hidden titles, or procedurally generated worlds, as in NetHack and Crafter. Those are exactly where the numbers stay low.
Cost and variance
We have almost no cost data from the labs. The only per-turn number we found is the independent replication's 3.7 cents a turn. And almost every headline result is one run, often with a human operator restarting the stream. That is not a criticism of the streamers; it is expensive to do more. But it means error bars are missing from nearly every table on this page, and gaps of a few percent mean little.
What this means for game developers
The literal question, "can an AI beat my game," has a growing yes for turn-based, text-and-tile games and a mostly-no for real-time ones that must be played from pixels. Three practical implications seem solid to us.
First, playtesting. An agent that can walk a 2D RPG from title screen to credits, writing notes as it goes, is a plausible bug-finder for softlocks and unreachable content. The Pokémon runs are a good source of examples of what such an agent notices and misses: it ignores objects outside its focus, and treats trainer encounters as random. That is useful information about where it will be blind. We have not run such a test ourselves and we are not aware of a public study comparing agent playtesting to human QA; treat this as an informed guess.
Second, NPC and companion research. SIMA 2 is framed by DeepMind as an agent that "plays, reasons and learns with you." Games that want AI companions may get ideas from there, though the research is closed to most developers. For the more immediate question of whether language models are fast and cheap enough for dialogue in a shipped game, see our piece on AI NPCs and latency.
Third, harness design is becoming a product category. If the system matters more than the model, as Jev and the Gemini scaffolds suggest, then the people building memory, mapping and tool layers around models have as much influence as the model labs. We cover how fast the underlying models are moving, and how to read those scores, in how fast AI is improving and in our pillar review of AI models for game development.
How we reviewed this
We read the public record: the LessWrong posts by Julian Bradshaw and Josh Snider on Claude Plays Pokémon, Joel Zhang's blog posts on Gemini Plays Pokémon, the BALROG paper and the live leaderboard (read 2026-10-09), the VideoGameBench paper and its HTML version, DeepMind's SIMA 2 post, the Voyager and Ghost in the Minecraft papers, Anthropic's Fable 5 announcement, press coverage of the GPT-5 and Jev runs, and an independent developer's replication write-up. We did not run any agents or games ourselves.
Things we could not verify. We could not read the full text of the Tom's Hardware article on Jev, the MarkTechPost SIMA 2 page, or Zhang's original per-run Twitch logs, so those figures come from search summaries or secondary sources. The Fable 5 FireRed time ("just over 50 hours") comes from a comment by Bradshaw, who says no full stream exists, and Anthropic's page does not state it. The model names on the BALROG board (GPT-6-Astra-Max, Claude-Opus-5-Max and others) are as the site lists them; we did not independently check the runs. Pokémon step and hour counts use different definitions across harnesses. We also did not find a current 2026 Minecraft leaderboard for LLM agents, and we did not find a 2026 update to the VideoGameBench results.
FAQ
Can AI beat Pokémon? Yes, several have. Gemini 2.5 Pro beat Pokémon Blue in May 2025, GPT-5 finished Red in 6,470 steps, Gemini 3 Pro beat Crystal in December 2025, and Claude Opus 4.7 beat Red in May 2026. Anthropic says Claude Fable 5 beat FireRed with a vision-only harness in June 2026.
Which AI is best at video games? It depends on the game and the harness. On BALROG's language-model table GPT-6-Astra-Max leads at 68.3%, ahead of Claude-Opus-5-Max at 63.4%. In Pokémon, the fastest reported finishes come from GPT-5 (fewest steps) and, by an outside estimate, Claude Fable 5 (fewest hours). These are not the same test.
Can AI play games in real time? Mostly not from raw pixels, according to VideoGameBench, where the best 2025 models completed under 1% of the real-time benchmark and 1.6% of the Lite version. Specialized systems such as OpenAI Five did play real-time games at champion level, but they read game state through an API and trained on one game.
What is the BALROG benchmark? BALROG is a benchmark of existing RL environments, from easy gridworlds to NetHack, that measures how far language and vision-language models progress, reported as an average "% Progress."
Is "Claude Plays Pokémon" a real benchmark? It is a public demonstration more than a benchmark. Harnesses change between runs, there is usually one run per model, and the games are well documented online, so it works better as a progress meter than as a controlled comparison.
Can AI play Minecraft? Agents that call code APIs, like Voyager, can gather resources and climb the tech tree far faster than older methods. DeepMind's SIMA 2 plays from pixels with keyboard and mouse and reports 45 to 75% task completion on held-out environments including MineDojo, per The Decoder.
CreateGame.ai is a pre-launch tool for building browser-playable games from a sentence; if that interests you, join the waitlist at https://creategame.ai.
Sources
- VideoGameBench paper, arXiv 2505.18134 (abstract): https://arxiv.org/abs/2505.18134
- VideoGameBench paper, HTML v3 (results table, games list): https://arxiv.org/html/2505.18134v3
- BALROG paper: https://arxiv.org/abs/2411.13543
- BALROG leaderboard (read 2026-10-09): https://balrogai.com
- Julian Bradshaw, "A Year Late, Claude Finally Beats Pokémon" (LessWrong, 16 May 2026): https://www.lesswrong.com/posts/sehJYg5Yny9fvpbpt/a-year-late-claude-finally-beats-pokemon
- Julian Bradshaw, "Insights into Claude Opus 4.5 from Pokémon" (LessWrong, Dec 2025): https://www.lesswrong.com/posts/u6Lacc7wx4yYkBQ3r/insights-into-claude-opus-4-5-from-pokemon
- Josh Snider, "Claude Plays Pokemon: Opus 4.5 Follow-up" (29 Jan 2026): https://www.lesswrong.com/posts/gogZyeistdaDFuhbG/claude-plays-pokemon-opus-4-5-follow-up
- Julian Bradshaw's shortform (Fable FireRed, Gemini and Claude run figures): https://www.lesswrong.com/posts/ekF2EDwKyZJNuxBTb/julian-bradshaw-s-shortform
- Anthropic, "Claude Fable 5 and Claude Mythos 5" (9 June 2026): https://www.anthropic.com/news/claude-fable-5-mythos-5
- NerdZap on the Fable 5 FireRed run (10 June 2026): https://nerdzap.com/news/claude-fable-five-pokemon-firered-vision/
- DEV Community, independent replication of the vision-only run: https://dev.to/qingze_hu_c4c251c1b353ede/i-replicated-the-vision-only-pokemon-run-anthropic-showcased-on-the-fable-5-launch-page-then-i-35ah
- Joel Zhang, "Gemini 3 Pro vs 2.5 Pro in Pokemon Crystal" (12 Dec 2025): https://blog.jcz.dev/gemini-3-pro-vs-25-pro-in-pokemon-crystal
- Joel Zhang's blog index: https://blog.jcz.dev
- DeepakNess on Gemini 2.5 Pro playing Pokémon: https://deepakness.com/raw/gemini-2-5-pro-pokemon/
- TechRadar on GPT-5 finishing Pokémon Red: https://www.techradar.com/ai-platforms-assistants/chatgpt/gpt-5-just-completed-pokemon-red-in-a-new-world-record-time-claude-gemini-and-chatgpt-o3-arent-even-close
- Tom's Hardware on Jev: https://www.tomshardware.com/tech-industry/artificial-intelligence/developer-says-jev-decision-model-beat-pokemon-red-in-under-a-week-non-llm-engine-succeeds-where-traditional-chatbots-stalled-for-months-but-claude-opus-5-coached-the-model-through-its-dead-ends
- remio.ai on the Jev run: https://www.remio.ai/post/jev-pokemon-red-run-finished-in-37-hours-but-claude-helped-build-the-winning-sys
- Orak benchmark, arXiv 2506.03610: https://arxiv.org/abs/2506.03610
- Voyager project page: https://voyager.minedojo.org/
- Ghost in the Minecraft, arXiv 2305.17144: https://arxiv.org/abs/2305.17144
- MineNPC-Task, arXiv 2601.05215: https://arxiv.org/pdf/2601.05215
- PillagerBench, arXiv 2509.06235: https://arxiv.org/pdf/2509.06235
- TechCrunch on MC-Bench (20 March 2025): https://techcrunch.com/2025/03/20/a-high-schooler-built-a-website-that-lets-you-challenge-ai-models-to-a-minecraft-build-off
- DeepMind, "SIMA 2" (13 Nov 2025): https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds/
- The Decoder on SIMA 2: https://the-decoder.com/deepminds-latest-ai-agent-learns-by-exploring-unfamiliar-games-and-ai-built-worlds/
- MarkTechPost on SIMA 2: https://www.marktechpost.com/2025/11/16/google-deepmind-introduces-sima-2-a-gemini-powered-generalist-agent-for-complex-3d-virtual-worlds/
- OpenAI Five paper, "Dota 2 with Large Scale Deep Reinforcement Learning": https://arxiv.org/abs/1912.06680
- Wikipedia, OpenAI Five: https://en.wikipedia.org/wiki/OpenAI_Five
- ScienceAlert on AlphaStar Grandmaster (30 Oct 2019): https://www.sciencealert.com/starcraft-ii-has-a-new-grandmaster-and-it-s-not-human