CreateGame.ai

Reports

AI RPG: Which Models Make the Best Text RPGs and Game Masters

An AI RPG is only as good as its worst hour: the point at which the narrator forgets your character's name, hands the dragon a sword it already lost, or quietly refuses to continue the story. We reviewed the published benchmarks, product pages, pricing, papers and forum reports to work out which models are worth running a text RPG on, which products wrap them well, and what an hour of play really costs.

Key takeaways

  • No benchmark measures what an AI game master does. Creative-writing boards, long-context tests and roleplay proxies each cover a slice, and the two main creative-writing boards we found rank the same models almost in reverse (EQ-Bench has Claude Opus 5 first and Claude Fable 5 sixth; arena.ai's creative category has them the other way round).
  • Memory is the real bottleneck, not intelligence. Frontier models now advertise 1M-token windows, but consumer RPG products still run on small slices: NovelAI's top tier is documented at up to 28,672 tokens, and AI Dungeon's tiers reportedly span about 4,000 to 16,000.
  • The most credible fix is not a bigger window. Voyage, Friends & Fables and the forum GMs who report success keep state (HP, gold, inventory, who knows what) outside the model.
  • Our arithmetic puts an hour of play at roughly $0.05 to $0.15 on cheap open-weights models, about $0.76 on Claude Sonnet 5.5 or GPT-6.1 Sol and $1.52 on Claude Opus 5.5, assuming 40 turns with an 8,000-token context. Flat-rate products work out to about $0.17 to $0.83 per hour at 60 hours a month (the top of that range is AI Dungeon Mythic at $49.99).
  • Content policy differs more than model quality does, and the 2021 AI Dungeon moderation revolt shows what happens when a product changes its mind.
  • For most people the practical pick is a purpose-built product (AI Dungeon, NovelAI, Friends & Fables or Voyage). For control, SillyTavern plus a mid-priced open-weights model is the cost-effective setup, and a frontier model is worth paying for when prose quality matters more than price.

What an AI game master actually has to do

When people search for an AI RPG they usually mean one of three things, and the models rank differently for each.

The first is the text adventure: you type what you do, the model narrates what happens. AI Dungeon popularised this in 2019, and it remains the reference product. The model's job is mostly prose and plausibility. Rules are loose.

The second is the AI game master in a rules-heavy game: a D&D-style system with character sheets, hit points, initiative and skill checks. This asks a language model to be a referee, which is a very different job from being a novelist. The model must apply rules consistently, track numbers precisely and refuse to let the story override the dice.

The third is character roleplay: a persistent persona (a companion, a rival, a shopkeeper) you talk to across weeks. Here the key properties are voice stability and memory of the relationship, not plot.

We keep these separate because the evidence for each differs. For the first, creative-writing leaderboards are a reasonable proxy. For the second, the only dedicated benchmark we found suggests current models struggle. For the third, there is no public benchmark we would trust, though the products now ship their own memory features. If you are building real-time characters rather than playing a story, our post on AI NPCs and real-time dialogue latency covers the speed and cost side, which is a different problem.

Underneath all three sits the same technical constraint: the model sees only what is in its context window on each turn. Everything else, from the quest you accepted in session one to the name of the tavern, must be re-supplied by the product or it is gone. Much of what follows is really a comparison of how well each product manages that re-supply.

The products people actually use

The market is a mix of a decade-old text adventure, a subscription writing tool, a persona-chat giant, a group-play game with rules, and a do-it-yourself frontend. We looked at the vendor pages, press coverage and pricing for each. Some pages did not load for us, so we say where a number comes from secondary reporting.

The AI RPG products people usePrice, memory and model details as published; blanks mean we could not verify them.
Products
AI Dungeon (Latitude)Context about 4k to 16k tokens by tier (via search summary); credits buy more; Latitude fine-tunes such as Wayfarer
NovelAIUp to 28,672 tokens on Opus; home page says GLM 4.6, docs name older models; lorebook, no rules engine
Friends & Fables5e-style rules engine, tracks HP and inventory, 3 to 6 players; model not stated
Voyage (Latitude)World Engine keeps state outside the model; Gemma for text; vendor-reported results
Character.AIPersona chat; open-ended chat removed for under-18s Nov 2025; Stories format
Hidden DoorNarrator-driven worlds, modifiers, cards; partner program
Infinite WorldsBrowser text adventure launched Feb 2024; GPT-4-class text and image generation at launch
SillyTavernFrontend only; World Info lorebooks; any local or cloud model

Plan tables for AI Dungeon and Character.AI come from search summaries; the live pages did not render for us. Source: Vendor pages and press (see post sources 11-25), accessed 2026-10-09.

AI Dungeon

Latitude's AI Dungeon is the original, and its plan structure has become a ladder of context and model access. According to Latitude's membership overview as reported in search results, the paid tiers are Adventurer ($9.99 a month, up to 4,000 tokens of context), Champion ($14.99, 8,000), Legend ($29.99, 16,000 plus access to larger "Ultra" models) and Mythic ($49.99) [14]. We could not load that help page ourselves, so treat the context figures as provisional. The live pricing page, which also failed to render the plan table, does confirm how context works: every tier gets a base amount, some models let you buy extra context with Credits each turn, and credits do not expire [15].

On the model side, Latitude has leaned on its own fine-tunes for years. Wayfarer, a Mistral NeMo 12B fine-tune, is described by the maker as built to deliver "genuine conflict and tension", and Wayfarer 2 12B, published on Hugging Face in August 2025, advertises higher stakes with more frequent character failure and death [16]. That design choice is the clearest example of a product-level fix for a model-level problem: base chat models are trained to be agreeable, and a narrator who never lets you lose is a bad narrator.

Voyage, Latitude's second attempt

The most interesting 2026 development is Voyage, which Latitude launched on 21 April 2026 as an AI-native RPG platform [17]. TechCrunch's write-up describes a "World Engine" that coordinates several AI systems to narrate actions, track characters and objects, and remember relationships, and says it took about five years to build. Latitude's own account, repeated by Runtime Wire, is that the engine separates world state from narration: it tracks what exists and what changed (health, currency, inventory, locations, relationships) and the language model turns that into prose [18].

This is the architecture we would bet on, and we will return to it. What we cannot say is how well it works. The details come from Latitude's press release and statements; the TechCrunch reporter's hands-on observations were that NPCs were unscripted (a captor troll began talking about his marriage) and that betraying a character made them avoid or rival the player [17]. Runtime Wire notes that open beta began in late August 2026 and that whether the engine stays coherent over real multiplayer sessions "is still unproven" [18]. Models named in the coverage are Latitude's proprietary ones plus Google's Gemma for text; subscription tiers of $15, $30 and $50 were announced for after the beta [17].

NovelAI

NovelAI is the writer's tool of the group, and it has the clearest privacy positioning. Its home page promises "complete anonymity", says prompts are encrypted, and says story generation is "powered by GLM 4.6" [12]. The documentation lists Tablet at $10 a month, Scroll at $15 and Opus at $25, all with unlimited text generation, and states that Opus gets up to 28,672 tokens of context [11]. The same docs page still names older models (Xialong, Krake, Llama 3 Erato) as what Opus subscribers can use, which conflicts with the home page's GLM 4.6 claim. The docs and the front page are out of step, and we could not tell which is current for each tier. The docs page does not state a content policy; the home page says "No restrictions. No compromises." in its image section, which is marketing language rather than policy [12].

NovelAI is not a game master. It is a continuation engine with a lorebook and a memory field. That suits writers who want to steer, and it suits players who like to co-author a story, but it does not referee anything.

Friends & Fables

Friends & Fables is the closest thing in this list to a tabletop game. Its comparison page (late-2025 pricing) describes a 5e-inspired rules engine, an AI game master that tracks hit points, inventory and conditions, character sheets, turn-based combat, battlemaps, and a long-term lore system the AI reads during play [19]. The free plan allows 5 to 25 AI turns per day and three players; Starter is $19.95 a month for four players, Pro $29.95 for five, and Legend $39.95 for six, with monthly credits spent on premium narration models and image generation [19]. The page does not say which language models it uses. That is a pattern across the category: the product is the product, and the model underneath changes without notice.

Character.AI and persona chat

Character.AI is by volume the biggest roleplay destination, but it is a chat product first and a game second. Its recent history is mostly about policy. On 29 October 2025 the company announced it would remove open-ended chat for under-18 users, effective 25 November, and introduced "Stories", a structured choose-your-path format where users pick two or three characters and a genre and then make choices [20]. In January 2026 Character.AI and Google agreed to settle five lawsuits over teen deaths and mental-health harms, with terms undisclosed [21]. When the company needed a safer way to let people play with AI characters, it chose a guided structure with discrete choices over free text.

Hidden Door, Infinite Worlds and others

Hidden Door launched in 2022 with a $2M pre-seed round and described its first product as "Roblox meets D&D" with an AI narrator [22]. Its current site describes playable stories where you pick a world, add modifiers and create a character, with a partner program for authors, but lists no prices and no dated announcements [23].

Infinite Worlds, a solo-developer browser game that launched on 14 February 2024, uses GPT-4-class text and Stable Diffusion images [24]. We could not confirm its current model or pricing. We also looked for a product called "AI Realm" and found nothing verifiable under that name; the closest hit was a generic iOS app called AI-RPG with too few ratings to judge [25].

SillyTavern and the open-model route

SillyTavern is a free, open-source (AGPL-3.0), locally installed frontend, not a model. It connects to local engines or to cloud APIs, and its World Info feature lets you keep lore outside the character card to save tokens [13]. This is where most enthusiasts who want control end up, and it is why open-weights models matter in this category more than in most. The community consensus from a mid-2026 guide we could only partially access points to DeepSeek V4 variants for value, Claude for prose at a premium, and Kimi, GLM and MiMo as rotation picks [26]. The guide judges prose by anecdote, not benchmark, and we could not open it in full, so read it as a snapshot of forum opinion.

Which models write the best RPG prose

There is no single leaderboard for this, so we looked at four instruments and what each can and cannot say.

EQ-Bench and arena.ai disagree

EQ-Bench Creative Writing v3 is an LLM-judged benchmark: models write responses to prompts and judge models score them with a rubric and pairwise Elo [4]. Its longform variant asks a model to plan a story and then write eight chapters of about 1,000 words, and reports a "degradation" figure for how much quality falls between the first and last chapter [5]. That is the closest published analogue to a long campaign, though eight chapters is a short campaign.

A 31 August 2026 analysis by Digital Applied compared EQ-Bench with the arena.ai Creative Writing category (snapshot 27 August 2026, 1,214,472 votes across 393 models). On EQ-Bench the top five were claude-opus-5 (2116.1), kimi-k3 (2070.8), GLM-5.3 (2062.4), gpt-5.6-sol (1964.1) and ox-alpha (1959.7). On arena.ai the top was claude-fable-5 (1505), then claude-opus-4-6-high (1500), claude-opus-4-7-high (1489) and gemini-3-pro (1483) [3]. Fable 5 is first on arena.ai and sixth on EQ-Bench; Kimi K3 is second on EQ-Bench and twenty-third on arena.ai [3]. The analysis' own explanation is that arena.ai uses blind human pairwise votes while EQ-Bench uses Claude judges and deliberately adversarial prompts, and that EQ-Bench does not control for judges favouring their own outputs [3].

EQ-Bench Creative Writing v3, top fiveLLM-judged Elo; arena.ai's human-vote board ranks these models very differently.
EQ-Bench Creative Writing v3, top five05001,0001,5002,0002,500AnthropicMoonshotZ.aiOpenAIclaude-opus-52,116Mkimi-k32,071ZGLM-5.32,062gpt-5.6-sol1,964ox-alpha1,960Model

Secondhand: reported by Digital Applied on 31 Aug 2026; we could not load the live table. On arena.ai's Creative Writing category (27 Aug snapshot) claude-fable-5 is first at 1505 and kimi-k3 is 23rd. Do not compare the two scales. Source: Digital Applied analysis of EQ-Bench and arena.ai, accessed 2026-10-09.

Model names are as the source writes them and may not match our snapshot's version strings (it lists Claude Opus 5.5 and Fable 5.1). We could not load the live EQ-Bench table, so this is secondhand from one analysis [4].

A third opinion from professional writers

Surge AI's Hemingway-bench had professional writers (screenwriters, poets, speechwriters, editors) run more than 5,000 blind pairwise comparisons on real writing prompts. Surge reports that EQ-Bench's autograder agreed with its expert writers as little as 43% of the time in some categories, and that EQ-Bench appeared to reward heavy use of literary devices [6]. At the top of Surge's board were Gemini 3 Flash, Gemini 3 Pro and Claude Opus 4.5. The page shows two dates (4 February and 29 September 2026), so these may be older models than today's leaders [6].

Put together: Claude models are strong on human-preference measures, Kimi K3 and GLM-5.3 are strong on the LLM-judged one, and the people best qualified to judge prose distrust the LLM-judged board. Prose that wins by sounding literary is not necessarily what you want from a narrator who must stay out of your way.

What a roleplay benchmark would measure, and why none exists

BenchLM states plainly that "there is no dedicated roleplay benchmark" and ranks models instead on proxies such as instruction-following, WildBench and MuSR (multi-step narrative tracking). Its top five on the 9 October 2026 update were mostly open-weights models, led by Qwen3.5-27B (95), and Claude Opus 4.5 sat at 43rd with 74.5 [8]. We read it as a reminder that "good at instruction-following" and "pleasant narrator" are different properties, not as a ranking to follow.

There are academic efforts. RPGBench (Yu et al., arXiv 2502.00595) tests models as game engines in two tasks: creating a playable world in a structured event-state format, and simulating multi-round gameplay while tracking state and enforcing rules. Its abstract reports that top LLMs can produce engaging stories but often struggle to keep game mechanics consistent and verifiable, especially in long or complex scenarios [7]. The abstract names no models and the benchmark predates this year's releases, so it shows the shape of the problem, not today's leaders. A community RP-Bench leaderboard reportedly scores character consistency and lorebook use, but we could not open it.

CALYPSO, an AIIDE 2023 paper, framed the model as an assistant to a human Dungeon Master, and its evaluation was formative and user-centred rather than an automated consistency measure [27]. The design lesson is that the human holds the canon.

Context windows and memory

What the headline numbers say

Almost every frontier model in our snapshot advertises a 1,000,000-token window, including all the Claude 5.x models, GPT-6 and 6.1, Gemini 4 Argon, Kimi K3, GLM-5.3, MiMo-V2.6-Pro and DeepSeek V4. The exceptions are Grok 4.7 (500,000), Mistral Large 4 (524,288), Qwen3.8 27B (256,000) and the older gpt-oss-120b (131,072) [1].

Here is a back-of-envelope sizing, our estimate at 0.75 words per token. A 300-token reply is about 225 words. A 40-turn hour produces 12,000 tokens of narration (40 x 300); a 3-hour session is 36,000; a 20-session campaign is 720,000. That fits a 1M window on paper with no room for rules or lore, and resending it every turn would cost a great deal.

What the long-context tests say

Artificial Analysis's Long Context Reasoning benchmark (AA-LCR) asks 100 multi-step questions over documents of about 10,000 to 100,000 tokens, graded pass/fail by another model [2]. On v1.1 the top three were Kimi K3 (88.7%), Step 5 Preview (88.3%) and MiMo-V2.6-Pro (86.3%) according to AA's page; our snapshot, accessed the same day, has Claude Fable 5.1 at 85.3% and Opus 5.5 at 84.7% [1][2]. The spread across current frontier models is narrow, roughly 80% to 89%. Two things follow.

First, this is a ceiling, not a floor. AA-LCR documents stop at 100,000 tokens, a tenth of the advertised window, and the questions are retrieval-and-reasoning, not "does the narrator remember that the innkeeper was lying".

Second, the independent evidence on long context is less flattering than vendor claims. Chroma's July 2025 "Context Rot" report tested 18 models, including Claude Opus 4 and Sonnet 4, GPT-4.1, Gemini 2.5 Pro and Qwen3, and found performance became less reliable as input grew even on simple tasks; on a conversational memory test (LongMemEval) models did far better on a focused ~300-token prompt than on the full ~113,000-token one [10]. Distractors, meaning text that is topically related but does not answer the question, hurt more as length grew [10]. A fantasy campaign is made of distractors: dozens of similar-sounding characters, towns and factions.

Fiction.liveBench, which tests comprehension of long fiction, shows scores falling as stories get longer, but we could only reach mirrors of the board, and its leaders (GPT-5 and o3-pro at 97.2% at 16,000 tokens) are 2025-era models [9].

Long-context reasoning (AA-LCR v1.1)Documents of 10k to 100k tokens; current models cluster at 80-89%, older open models fall to about 50%.
Long-context reasoning (AA-LCR v1.1)020406080100MoonshotXiaomiAnthropicDeepSeekOpenAIGoogleZ.aiMetaMKimi K388.7MiMo-V2.6-Pro86.3Claude Fable 5.185.3Claude Opus 5.584.7DDeepSeek V4.1 Flash84GPT-6 Luna83.3GPT-6.1 Sol83Claude Sonnet 5.582.7Claude Haiku 5.582.7Gemini 3.8 Flash81.3ZGLM-5.379.7GPT-6 Astra80.7Gemini 4 Argon79.7gpt-oss-120b52Llama 4 Maverick50Model

Test documents stop at 100k tokens, a tenth of the advertised 1M windows. It measures retrieval and reasoning, not narrative consistency. Source: Artificial Analysis (via CreateGame.ai models snapshot), accessed 2026-10-09.

How products handle memory

AI Dungeon sells context directly: higher tiers get bigger windows, and credits buy more per turn on some models [15]. That is honest, and it is why a 16,000-token window on the top tier looks small beside the 1M of the underlying frontier models: the product pays per token. NovelAI uses a lorebook and memory field up to 28,672 tokens on Opus [11], so you curate what the model sees. SillyTavern's World Info injects lore entries when keywords appear, a retrieval trick that saves tokens but can miss an entry the player refers to indirectly [13]. Friends & Fables says its AI reads a "long-term lore system" during play [19]; how it chooses what to read is not documented on the page we saw.

Voyage's design is the most ambitious: the World Engine keeps health, currency, inventory, locations and relationships as explicit state across what Latitude says are thousands of turns, and the language model only narrates [17][18]. The TechCrunch article reported players averaging nearly 3,000 choices in beta [17], which is the right scale to test this on, but the reliability claim itself is vendor-reported.

What GMs report

We looked for public experience reports and found a recent Paizo forum thread (30 September to 2 October 2026) where Pathfinder GMs compared notes. One poster ran Curse of the Crimson Throne with ChatGPT as GM: character creation and the first two sessions went well, session 3 broke down (NPCs knew things they should not, gold and gear bonuses were miscalculated), and by session 4 the AI was ignoring recent events [28]. Another said AI characters tend to know everything the AI knows. A poster who reported good results with Claude as GM described a workflow of Markdown rules files, CSV spells and YAML stat blocks, and an instruction to ask rather than improvise, which they said kept the system consistent about nine times in ten [28]. Another reported success with ChatGPT using a running session-summary file [28].

This is a handful of anecdotes from one forum, not a study. But the pattern lines up with RPGBench and with Voyage's design: the failures are in state and knowledge boundaries, and the fixes are external files and tools.

Content policy and censorship

Players choose AI RPG tools for what they will and will not narrate as much as for quality. We reviewed what vendors say publicly; we did not test any product with sensitive prompts or read every usage policy line by line.

AI Dungeon learned this the hard way. In 2021, after OpenAI pressed Latitude to act on sexual content involving minors that the system was generating, Latitude introduced a blocklist and human review of flagged stories; users revolted over false positives (a laptop described as eight years old was flagged) and over the privacy implications of moderators reading private stories [29][30]. The episode is the origin of the "uncensored" segment: NovelAI emerged as an alternative that used open models outside OpenAI's control [29]. Today Voyage says some content is mature, comparable to Steam titles, with safety measures and parental controls [17].

NovelAI markets privacy and a lack of restrictions, though its documentation page has no explicit content policy and the "No restrictions" line sits in the image-generation section of the home page [11][12].

Character.AI has moved the other way. Under-18 users lost open-ended chat in November 2025 after litigation, and an age-assurance system now sorts accounts into brackets [20][21]. One secondary source reports the content filter remains in place for adults; we could not confirm this from the company.

Hosted frontier models apply their providers' usage policies, generally stricter than roleplay-specific products, and these change. OpenAI announced an opt-in "adult mode" for ChatGPT for early 2026, delayed it in March, and the Financial Times reported it indefinitely postponed; we found no later reporting of a launch [31]. If you build on an API, read the provider's rules before you design around a genre.

Open-weights models shift the question to whoever hosts them. A local model through SillyTavern follows only your rules; a hosted service built on open weights takes on the moderation burden itself, and the Latitude story is a warning about doing that under pressure.

The practical upshot: pick the product whose policy matches your use before you compare prose. A better model that refuses your genre is worse than a weaker one that does not.

What an hour of play costs

We built the arithmetic below ourselves. These are estimates, not measurements, and all prices are list prices from the Artificial Analysis snapshot on 9 October 2026 [1].

Assumptions. One hour is 40 turns (a turn every 90 seconds). Each turn resends a context of 8,000 tokens (system prompt, lorebook entries, recent history) and generates 300 tokens of narration. That makes 40 x 8,000 = 320,000 input tokens and 40 x 300 = 12,000 output tokens per hour. We assume a non-reasoning mode; reasoning models bill thinking tokens as output and would cost more. The second scenario uses a 32,000-token context: 1,280,000 input tokens and the same 12,000 output.

Worked example, Claude Sonnet 5.5 ($2 input, $10 output per million tokens). Input: 0.32 x $2 = $0.64. Output: 0.012 x $10 = $0.12. Total: $0.76 per hour at 8,000 tokens. At 32,000: 1.28 x $2 = $2.56, plus $0.12, is $2.68.

Estimated API cost per hour of AI RPG play40 turns an hour, 300 tokens of narration per turn, full context resent each turn, no prompt caching.
Estimated API cost per hour of AI RPG play$0$2$4$6$8$10$12$148,000-token context32,000-token contextClaude Fable 5.1$3.8$13.4Claude Opus 5.5$1.52$5.36Claude Sonnet 5.5$0.76$2.68GPT-6.1 Sol$0.76$2.68Kimi K3$1.14$4.02GLM-5.3$0.5$1.84Gemini 3.8 Flash$0.29$1.01MiMo-V2.6-Pro$0.15$0.57DeepSeek V4.1 FlashDeepSeek V4.1 Fla…$0.11$0.4Claude Haiku 5.5$0.04$0.13Model

Our estimate: hourly cost = turns x context tokens x input price + turns x 300 x output price, using list prices of 9 Oct 2026. Non-reasoning mode assumed; reasoning tokens would add cost. Flat-rate products come to about $0.17-$0.83 per hour at 60 hours a month. Source: Artificial Analysis list prices (via CreateGame.ai models snapshot); arithmetic by CreateGame.ai, accessed 2026-10-09.

Model $/1M in / out 8k context, per hour 32k context, per hour 8k, 80% cached input
Claude Fable 5.1 10 / 50 $3.80 $13.40 $1.30
Claude Opus 5.5 4 / 20 $1.52 $5.36 $0.55
Claude Sonnet 5.5 2 / 10 $0.76 $2.68 $0.27
GPT-6.1 Sol 2 / 10 $0.76 $2.68 $0.27
Kimi K3 3 / 15 $1.14 $4.02 $0.45
GLM-5.3 1.4 / 4.4 $0.50 $1.85 $0.21
Gemini 3.8 Flash 0.75 / 3.75 $0.29 $1.01 $0.11
MiMo-V2.6-Pro 0.435 / 0.87 $0.15 $0.57 $0.04
DeepSeek V4.1 Flash 0.3 / 1.2 $0.11 $0.40 $0.04
Claude Haiku 5.5 0.1 / 0.5 $0.04 $0.13 $0.015

The cached column assumes 80% of each turn's input is a repeated prefix billed at the vendor's cache-hit price (for Sonnet 5.5: 0.2 x 0.32 x $2 + 0.8 x 0.32 x $0.10 + 0.012 x $10 = $0.128 + $0.026 + $0.12 = $0.27). Prompt caching works well for RPGs because the system prompt and lore are the same each turn, but it needs careful prompt layout and cache lifetimes vary by provider, so this column is a best case [1].

Three things stand out. A cheap, fast model is two orders of magnitude cheaper than the premium tier ($0.04 against $3.80). The input side dominates the bill, so the cost of "better memory" (a bigger context) multiplies it by four. And the absolute figures are small: even the Opus 5.5 scenario is lower per hour than a cinema ticket.

Compare with flat-rate products. At 60 hours a month (two hours a day), NovelAI Tablet at $10 is $0.17 an hour, Opus at $25 is $0.42, AI Dungeon Adventurer at $9.99 is $0.17 and Legend at $29.99 is $0.50, and Friends & Fables Starter at $19.95 is $0.33 [11][14][19]. Those products are cheaper than a frontier API for heavy players, because they run smaller context on smaller or fine-tuned models. If you play two hours a month, the API is cheaper.

The API figures exclude bundled image generation, and SillyTavern is free software that costs you an evening to wire up.

Which model for which kind of RPG

We have not played these models as game masters ourselves, so these recommendations come from the published evidence above, not from sessions we ran.

Prose-first solo adventure with a lot of freedom. A purpose-built product (AI Dungeon, NovelAI, Voyage) is the sensible default: someone else has solved prompt layout, memory and pacing. If you want to bring your own model, the creative-writing boards point to Claude Opus 5.5 or Sonnet 5.5 for human-preference quality and to Kimi K3 or GLM-5.3 if you want the best open-weights model on the EQ-Bench board. Expect to pay for it in context-resend cost.

Rules-heavy game master (D&D, Pathfinder, a custom system). Choose a tool or product that holds state outside the model: Friends & Fables, Voyage, or your own setup with files. On the model side, long-context reasoning scores are tightly bunched, so the cheapest model that follows instructions reliably is usually the right one. Haiku 5.5, GPT-6 Luna and MiMo-V2.6-Pro all have AA-LCR scores of 83% or above in our snapshot at a small fraction of the flagship price [1]. Use a stronger model only for set-piece scenes.

Companion or persona chat. The binding constraints are policy and memory, not intelligence. Character.AI's Stories format, SillyTavern's character cards and lorebooks, and NovelAI's memory field all address this differently. Whatever you choose, check the policy first.

Local and private play. SillyTavern with KoboldCpp and a fine-tune that fits your GPU; one 2026 guide suggests a 24B Mistral-based fine-tune for a 24 GB card [26]. The trade-off is that the largest open models in our snapshot are far too large for consumer hardware, so local players use something much smaller than the models benchmarked above. Older open models score poorly on long context: gpt-oss-120b and Llama 4 Maverick get 52% and 50% on AA-LCR against 80% or above for current models [1].

Building your own. If you are a developer, the architecture that keeps showing up is the same: a small deterministic engine for state and rules, retrieval over a lore store, and a language model that narrates. Pricing for that model side is in our cost breakdown, engine code is covered in our coding post, and the wider picture is in our guide to the best AI models for game development. CreateGame.ai, which is pre-launch, is aimed at generating whole playable games from a sentence, not at hosting open-ended narration like the tools above.

Where AI RPGs still break

The failures we would rank first, from the evidence above, are knowledge leakage (NPCs that know what the narrator knows, seen in the Paizo thread [28]), number drift (RPGBench's core finding [7]), positivity bias (the reason Wayfarer was tuned for failure [16]), and policy whiplash (Character.AI in 2025, AI Dungeon in 2021 [20][29]). None is fixed by a bigger model alone. The most credible path is products that treat the language model as a narrator, not the whole game, and the early evidence from Voyage, Friends & Fables and the forum GMs points the same way. Whether Voyage's engine delivers is, as Runtime Wire put it, the beta's test [18]. We will revisit when there is independent data.

How we reviewed this

We reviewed published sources only; we did not play or test any of these products or models. The sources are Artificial Analysis leaderboard data and the AA-LCR page (index v4.3.2, accessed 9 October 2026) via our shared models snapshot; a secondary analysis of EQ-Bench and arena.ai (Digital Applied, 31 August 2026); Surge AI's Hemingway-bench; the RPGBench and CALYPSO papers; Chroma's Context Rot report; and vendor pages and press for each product.

We could not verify the live EQ-Bench table (client-side rendered), Fiction.liveBench's own site, the RP-Bench community leaderboard, AI Dungeon's current plan table (figures come from a search summary of Latitude's help page), Latitude's 2026 model notes, OpenRouter's roleplay rankings, any product called "AI Realm", Friends & Fables' underlying models, or which of NovelAI's docs and home page is right about each tier's model. The cost-per-hour numbers are our estimates from list prices and will vary with turn length, caching and reasoning mode. Scores from different leaderboards use different scales and should not be compared.

FAQ

What is the best AI for an RPG? There is no single winner. Claude models lead human-preference creative-writing votes on arena.ai; on LLM-judged EQ-Bench, Claude Opus 5 leads with Kimi K3 and GLM-5.3 close behind [3]. For a ready-made game, the main options are AI Dungeon, NovelAI, Friends & Fables and Voyage.

Can ChatGPT or Claude be an AI dungeon master? Yes, with limits. Forum GMs report good early sessions and breakdowns by session 3 or 4 as context fills, and better results with external rules files and a running summary [28]. RPGBench found models struggle to keep mechanics consistent in long games [7].

How much does it cost to play an AI RPG? Flat-rate products run roughly $10 to $50 a month [11][14][19]. By API, our estimate is about $0.05 to $0.15 an hour on cheap open-weights models and $0.76 to $1.52 on Sonnet 5.5, GPT-6.1 Sol or Opus 5.5.

Which AI RPG tools are uncensored? NovelAI advertises no restrictions on its home page but has no explicit content policy in its docs [11][12]. Local models through SillyTavern follow only your rules. Character.AI, hosted frontier APIs and Voyage apply their own policies.

How long can an AI remember my story? NovelAI documents up to 28,672 tokens on its top tier [11]. Models advertise 1M tokens, but independent tests show accuracy falling on long inputs [10]. State-tracking designs like Voyage's aim to sidestep the problem [17].


If you want to follow what we are building at CreateGame.ai, you can join the waitlist at https://creategame.ai.

Sources

  1. Artificial Analysis, model leaderboard, accessed 9 Oct 2026 (via the CreateGame.ai models snapshot): https://artificialanalysis.ai/leaderboards/models
  2. Artificial Analysis, Long Context Reasoning (AA-LCR) v1.1: https://artificialanalysis.ai/evaluations/artificial-analysis-long-context-reasoning
  3. Digital Applied, "Which AI model writes best? Leaderboards disagree," 31 Aug 2026: https://www.digitalapplied.com/blog/which-ai-model-writes-best-leaderboards-disagree
  4. EQ-Bench Creative Writing v3: https://eqbench.com/creative_writing.html
  5. EQ-Bench Longform Creative Writing: https://eqbench.com/creative_writing_longform.html
  6. Surge AI, Hemingway-bench: https://surgehq.ai/blog/hemingway-bench-ai-writing-leaderboard
  7. Yu et al., "RPGBench: Evaluating Large Language Models as Role-Playing Game Engines," arXiv 2502.00595: https://arxiv.org/abs/2502.00595
  8. BenchLM, best models for roleplay (updated 9 Oct 2026): https://benchlm.ai/best/roleplay
  9. Fiction.liveBench scores via aggregator: https://www.datalearner.com/en/benchmarks/fiction-livebench
  10. Chroma, "Context Rot," 14 Jul 2025: https://trychroma.com/research/context-rot
  11. NovelAI documentation, subscription: https://docs.novelai.net/en/subscription
  12. NovelAI home page: https://novelai.net/
  13. SillyTavern documentation: https://docs.sillytavern.app/
  14. Latitude, AI Dungeon membership plan overview (figures via search summary; page did not render for us): https://latitudegames.notion.site/Membership-Plan-Overview-9085c85b381a4e29b6caf644b66c6382
  15. AI Dungeon pricing and context FAQ: https://play.aidungeon.com/pricing
  16. LatitudeGames, Wayfarer 2 12B model listing: https://featherless.ai/models/LatitudeGames/Wayfarer-2-12B
  17. TechCrunch, "Voyage is an AI RPG platform...," 21 Apr 2026: https://techcrunch.com/2026/04/21/voyage-is-an-ai-rpg-platform-for-creating-custom-gaming-worlds-with-ai-generated-npc-interactions
  18. Runtime Wire, "Latitude opens Voyage...": https://runtimewire.com/article/latitude-opens-voyage-ai-roleplay-world-engine
  19. Friends & Fables, plan comparison: https://fables.gg/blog/old-gregs-tavern-vs-friends--fables-plan--feature-comparison
  20. TechRadar, Character.AI launches Stories: https://www.techradar.com/ai-platforms-assistants/character-ai-launches-stories-to-keep-teens-engaged-as-it-scales-back-open-ended-chat-for-under-18s
  21. Bloomberg Tax, Character.AI and Google settle teen chatbot lawsuits: https://news.bloombergtax.com/legal-exchange-insights-and-commentary/character-ai-google-agree-to-settle-teen-chatbot-harm-lawsuits
  22. GamesBeat, Hidden Door launch (2022): https://gamesbeat.com/hidden-door-reveals-its-ai-powered-narrative-game-building-platform/
  23. Hidden Door home page: https://www.hiddendoor.co
  24. Game Developer, Infinite Worlds launch press release: https://www.gamedeveloper.com/press-release/play-a-unique-adventure-created-just-for-you-ai-text-based-adventure-game-infinite-worlds-launches-globally-today
  25. App Store, AI-RPG: https://apps.apple.com/qa/app/ai-rpg/id6449506940
  26. APIsRouter, best models for roleplay in 2026 (partially accessed): https://apisrouter.com/best-models-for-roleplay
  27. Zhu et al., "CALYPSO: LLMs as Dungeon Masters' Assistants," arXiv 2308.07540: https://arxiv.org/abs/2308.07540
  28. Paizo forums, "Anyone used AI as a GM with success?" (30 Sep to 2 Oct 2026): https://paizo.com/threads/rzs8tv4g
  29. The Register, AI Dungeon moderation (2021): https://www.theregister.com/2021/10/08/ai_game_abuse/
  30. TechSpot, AI Dungeon now censored: https://www.techspot.com/news/89571-machine-learning-text-adventure-ai-dungeon-now-censored.html
  31. TechCrunch, OpenAI delays ChatGPT's adult mode again, 7 Mar 2026: https://techcrunch.com/2026/03/07/openai-delays-chatgpts-adult-mode-again/