CreateGame.ai

Reports

AI World Building for RPGs: Models, Workflows and Where It Breaks

World building AI is good at the part of a setting nobody remembers writing (the fourth tavern, the minor saint, the regional dish) and poor at the part players will remember (the premise, the contradiction that was meant to be there, the fact that the baron is dead). We reviewed the published benchmarks, tool pages, surveys, legal reports and public posts from game masters and writers to see which models and tools suit RPG world building, where they fail, and a workflow that keeps the canon yours.

Key takeaways

  • The dedicated world-building tools split on AI. LegendKeeper states on its site that it "does not have generative AI features"; Kanka and Campfire say nothing about AI on the pages we could read; for World Anvil we found only a community thread and a competitor's blog post, not an official statement. The most-used workflow is therefore a general LLM beside a notes tool, not an integrated feature.
  • Context size is no longer the limit. Most frontier models advertise 1M tokens, enough for roughly 750,000 words of lore by our arithmetic. The limit is consistency: independent tests show accuracy falling as inputs grow, and a campaign bible is mostly near-duplicates, which is the worst case.
  • Cost is trivial. Resending a 100,000-word world bible to Claude Sonnet 5.5 costs about $0.27 per request uncached, by our estimate, and a few cents on small models.
  • Sentiment among the people who make games is hostile and getting more so: 52% of respondents to GDC's 2026 survey said generative AI is harming the industry, up from 30% a year earlier, and the figure is 63% among game design and narrative staff.
  • The legal position is clearer for outputs than for inputs. The US Copyright Office says prompts alone do not make you the author; selection, arrangement and your own edits can be protected. If you publish, major TTRPG outlets (Paizo, the ENNIEs, Wizards for D&D) bar or restrict AI content.
  • Image models are good enough for moodboards and rough maps, and an open-weights option exists, but there is no benchmark for maps. Treat generated maps as concept art, not as the map players navigate by.

What "world building" asks of an AI

Three jobs hide inside the phrase, and the models are not equally good at them.

The first is generation: inventing names, factions, histories, creatures, magic systems, rumours. This is a brainstorming task with a human filter, and it is where language models are most useful. A weak idea costs seconds to discard.

The second is consistency: remembering that the river runs north, that the cult worships a drowned god, that the queen's sister was exiled twelve years ago. This is a retrieval-and-reasoning task over a growing document, and it is where models fail quietly. The error does not look like an error. It looks like a plausible sentence that contradicts page forty.

The third is judgment: deciding what the setting is about. This is the premise, the tone, the central tension, the thing that makes a world feel authored. Almost every practitioner we found says this belongs to the human, and we agree, for a reason that is as much practical as philosophical: a premise is a constraint, and constraints are what make the generated filler coherent.

The rest of this post follows that split. If you want the interactive version of the problem, where a model runs a world rather than writes it, see our post on AI RPGs and game masters. This one is about preparing the world.

The tools: dedicated world-building apps and what they say about AI

The mainstream tabletop world-building tools are wiki-like databases with maps, timelines and relationship graphs. They were built before the current wave of models, and their communities are protective of that. We looked at the vendors' own pages.

World-building tools and what they say about AIFrom each vendor's own pages where we could read them; blanks mean unverified.
Tools
LegendKeeperFAQ: "No, LegendKeeper does not have generative AI features."
Campfire18 writing and worldbuilding modules; no AI statement on the page we read
KankaNo AI features or policy found in the official pages we reached
World AnvilAI status unconfirmed: 2024 community thread says AI out of scope; a competitor claims an AI feature called the Sage
Friends & FablesCharacters, monsters, items, locations feed an AI game master
General LLMs (ChatGPT, Claude, Gemini)No built-in lore management; consistency depends on what you paste or retrieve

This is a review of published pages, not a hands-on test. Source: Vendor pages (LegendKeeper, Campfire, Kanka, Friends & Fables); World Anvil community thread, accessed 2026-10-09.

LegendKeeper is explicit. Its FAQ answers the question directly: "No, LegendKeeper does not have generative AI features." [7] It sells a browser-based wiki, maps, timelines and whiteboards at $9 a month ($7.50 on the annual plan), with a free tier for viewing, exporting and collaborating, and a 14-day trial [7]. That is a position, not a gap.

Campfire describes itself as writing and world-building software with "18 writing and worldbuilding modules" (character creator, interactive maps, timeline maker, calendar creator, conlang software among them); accounts are free with limited module use, and paid plans start "as low as $0.50/month" with lifetime purchases available [8]. The page we read says nothing about AI features or policy. Directories list it under AI writing tools, but we could not find an official description of any AI function, so we do not claim one.

Kanka is a community-driven campaign and world manager; its features page says its core features are free [9]. We found no AI features or AI policy in the official documentation we could reach, and no changelog entry that says otherwise. Because we could not read the full changelog, this is an absence of evidence in the pages we saw, not proof of a policy.

World Anvil is the largest by reputation and the hardest to pin down, because its own site returned an access error to our fetch. What we have is secondhand. A 2024 community suggestion asking for an "AI Novel Writer" drew a team response, as quoted in search results, that AI features were "not within the scope" of the platform, along with strong objections from members, one of whom wrote that they and many others would leave "in droves" if it arrived [10]. A competitor, Sudowrite, wrote that an AI called "the Sage" is built into World Anvil and can expand a city stub into a history and districts, while noting that the prose can read as generic [11]. We cannot reconcile the two. A search summary also suggested "Sage" is the name of a World Anvil membership tier, which would make the Sudowrite claim a possible mix-up; we could not verify that either. If you care, ask World Anvil or check its current feature pages before relying on any of this.

The pattern is more informative than any single product. The people who use these tools are often the people who object most strongly to generative AI, and the vendors know it. The AI-native alternatives are generic assistants (ChatGPT, Claude, Gemini) or writing tools such as Sudowrite and NovelAI, and, for playable worlds, products such as Friends & Fables, whose free world-building suite for characters, monsters, items and locations feeds an AI game master [28]. Hidden Door and Voyage also treat world creation as a first-class activity, with Voyage's creators able to describe settings, quests and villains in natural language [25].

One caution about the Sudowrite comparison we cite. It was written by a vendor with an interest, compares ChatGPT on GPT-4 and Claude 3 Opus (both several generations old), and publishes no method or data [11]. It is useful for its list of tools and for the observation that general chatbots have no built-in lore management, so consistency depends on how you paste context. It is not a current ranking.

Which models for which part of the job

There is no world-building benchmark. What exists is a set of adjacent measures, and we use them as proxies with that caveat in plain view.

Generation: creative-writing boards

For fast, varied invention, the closest public instruments are the creative-writing boards. A 31 August 2026 analysis found that EQ-Bench Creative Writing v3 and arena.ai's Creative Writing category rank the same models nearly in reverse: Claude Fable 5 is first on arena.ai and sixth on EQ-Bench, Kimi K3 is second on EQ-Bench and twenty-third on arena.ai [29]. Professional writers who ran Surge AI's blind comparisons reported that EQ-Bench's autograder agreed with them as little as 43% of the time in some categories [30]. We unpack that in the AI RPG post. For this post the conclusion is narrower: do not choose a lore generator by leaderboard rank. Try two or three on your own setting and keep the one whose first drafts you edit least.

Consistency: long-context tests

For holding a large bible in mind, the relevant measure is long-context reasoning. Artificial Analysis's AA-LCR asks multi-step questions over documents of about 10,000 to 100,000 tokens [2]. Current frontier models bunch tightly between roughly 80% and 89% in our 9 October 2026 snapshot; older open models fall to about 50% (gpt-oss-120b at 52%, Llama 4 Maverick at 50%) [1].

Long-context reasoning vs input pricePrice barely predicts AA-LCR score: cheap models sit near the flagships.
Long-context reasoning vs input price50607080900246810AA-LCR v1.1 score (%)Input price ($ per 1M tokens)Kimi K3MiMo-V2.6-ProClaude Fable 5.1Claude Opus 5.5GPT-6.1 SolClaude Sonnet 5.5GPT-6 AstraGemini 4 ArgonGLM-5.3MiMo-V2.6-Flashgpt-oss-120bLlama 4 MaverickAnthropicOpenAIGoogleMoonshotZ.aiXiaomiDeepSeekMeta

AA-LCR documents run 10k to 100k tokens; it does not test the full advertised 1M windows or narrative consistency. List prices on 9 Oct 2026; some are introductory. Source: Artificial Analysis AA-LCR and model pricing (via CreateGame.ai models snapshot), accessed 2026-10-09.

Two readings of the chart matter. First, price barely tracks score. Claude Haiku 5.5, at $0.10 per million input tokens, scores 82.7%, within a few points of Claude Opus 5.5 at $4 (84.7%), and Kimi K3 leads the snapshot at 88.7% at $3 [1]. For retrieving and cross-referencing facts in a lore document, cheap models are not meaningfully worse on this test. Second, the test caps at 100,000 tokens, so it says nothing about the 1M windows the same models advertise.

Chroma's "Context Rot" report is the independent evidence we trust more here. Across 18 models, including Claude Opus 4 and Sonnet 4, GPT-4.1, Gemini 2.5 Pro and Qwen3, performance became less reliable as input length grew even on simple tasks; distractors (text that is related but does not answer the question) hurt more at greater lengths; and on a conversational-memory test, accuracy was far higher on a focused ~300-token prompt than on the full ~113,000-token one [3]. A campaign bible is a distractor factory: fifty towns with similar naming conventions, a dozen factions with overlapping goals. The practical consequence is that sending the whole bible every time is the worst way to use a long window, and sending the three relevant pages is the best.

Judgment: not a model question

No leaderboard measures taste, and none should. The strongest argument we found against outsourcing it comes from the writers themselves, and we cover it below.

What the bible costs to send

Because context is cheap relative to a person's time, it is worth knowing how cheap. These are our estimates, using the rule of thumb of 0.75 words per token (tokens = words / 0.75) and Claude Sonnet 5.5's list price of $2 per million input tokens and Claude Haiku 5.5's $0.10, from the Artificial Analysis snapshot of 9 October 2026 [1].

World bible size Tokens Sonnet 5.5, per request Haiku 5.5, per request
10,000 words (a one-shot's notes) about 13,300 $0.027 $0.0013
50,000 words (a small campaign) about 66,700 $0.13 $0.0067
100,000 words (a long-running wiki) about 133,000 $0.27 $0.013
300,000 words about 400,000 $0.80 $0.040
750,000 words (fills a 1M window) about 1,000,000 $2.00 $0.10

Worked example for the 100,000-word row: 100,000 / 0.75 = 133,333 tokens; 133,333 / 1,000,000 x $2 = $0.27. Output is billed separately and is small by comparison for lookup questions. Cache-hit input is priced far lower in the snapshot ($0.10 against $2 per million for Sonnet 5.5; $0.01 against $0.10 for Haiku 5.5), so repeated questions over an unchanged bible cost a fraction of these figures if the provider's caching applies to your setup [1].

None of this means you should paste 750,000 words. It means the cost objection to "load the whole wiki" is gone and the accuracy objection remains. Retrieval of the relevant entries beats a full dump on every test we reviewed that bears on the question [3].

What creators and designers say works and fails

We read a mix of surveys, studio statements and individual accounts. Most are recent; some are older and we date them.

The industry mood

GDC's 2026 State of the Game Industry survey, covering more than 2,300 professionals, found 52% of respondents say generative AI is harming the industry, up from 30% last year and 18% the year before; only 7% say it is helping, down from 13%. Sentiment is worst among game design and narrative staff (63%), visual and technical art (64%) and programmers (59%). At the same time, 36% use AI tools at work, with research and brainstorming the most common use (81%) [12].

Share of game professionals who say generative AI is harming the industryGDC State of the Game Industry surveys; 2026 sample is more than 2,300 respondents.
Share of game professionals who say generative AI is harming the industry0%10%20%30%40%50%60%202420252026Share of respondentsSurvey yearSay generative AI is harming the industry

Reported by GDC's 2026 write-up; some secondary sources give slightly different prior-year baselines. In 2026, 63% of game design and narrative respondents held this view, and 36% of all respondents used AI tools at work. Source: GDC 2026 State of the Game Industry, accessed 2026-10-09.

That combination (hostile and still using it) matches what we see in individual accounts. Brainstorming is the use people admit to; shipping generated content is the use they reject.

A studio's pull-back

Larian Studios gave the most public example. After its Divinity reveal, CEO Swen Vincke said the studio was using AI for concept art exploration, presentations and placeholder text, with no AI content in the final game. After fan backlash he went further in a Reddit AMA, saying there would be no generative AI art in Divinity and that the studio would not use it in the concept-art process or in writing; writing director Adam Smith said the team tested generative AI for writing and rated the results "3/10 at best" [13]. Outlets differ on whether this was a reversal or a clarification, and we have not seen the primary posts, so we report the position as the secondary coverage gives it. As evidence about quality it is a single studio's judgment of a single test. As evidence about the reputational cost of admitting to using the tools, it is strong.

The writers' objection

Brandon Sanderson's December 2025 keynote at Dragonsteel Nexus, published on his blog in January 2026, argues that the value of art is in what making it does to the maker. His line: "Art is the means by which we become what we want to be." [14] He objects to AI art even when well made, and does not treat AI as a neutral tool [14]. That is a moral claim, not an empirical one, but it has an empirical corollary useful to world builders: if the pleasure of your campaign is that you made the world, outsourcing the making removes the pleasure, and no model improvement changes that.

Game masters on what works

The accounts we found are consistent with each other. A DM writing in April 2023 reported that ChatGPT produced a quick mix of usable and weak ideas, that some made it into the game, and that suggestions were often generic and needed heavy rework to fit the world and rules; a Midjourney map worked well for play in a virtual tabletop [23]. That post is three years old and the models have changed, so we read it for the shape of the experience rather than the quality level.

More recent is a Paizo forum thread from 30 September to 2 October 2026 about using AI as a GM. Posters who use AI for prep describe it as a feedback tool for descriptions and set pieces, where short outputs limit the damage from hallucinations; one poster who pre-loaded published adventure text reported improvised content and leaked secrets, and another objected on IP grounds to uploading a published book into a model [21]. One poster who reported good results with Claude described a file-based setup (Markdown rules, CSV spells, YAML stat blocks) and an instruction to ask rather than improvise [21]. These are forum posts, not a study; we cite them as practice, not proof.

Marketing-adjacent sources add little. A June 2025 Wayline blog post lists benefits (terrain, biome and city generators, lore generators) and pitfalls (ecological inconsistency, generic architecture, over-reliance), but it is promotional content from an asset marketplace with no data [24]. We mention it only because it recites the same failure list: generic output, no coherence without a human steering.

IP and copyright

Four questions come up, and the answers differ in how settled they are.

Can you copyright AI-generated lore or art? In the US, the Copyright Office's January 2025 report on copyrightability concluded that, with current technology, a prompt does not give the user enough control to be the author of the output, and that re-prompting is "re-rolling the dice"; wholly AI-generated material is not protected. Human contributions can be: original selection, coordination or arrangement of AI output, original modifications to it, and your own work that remains perceptible in the output [15]. For a world builder this means a wiki you wrote, edited and organised is protectable in those human elements; a stack of unedited generated entries is a weaker claim. Other jurisdictions differ, and this is not legal advice.

Can you publish it? The tabletop market has its own rules, regardless of the law. Paizo has barred AI-generated creative work in its products and its community marketplaces since March 2023 [19]. The ENNIE Awards stopped accepting products containing generative AI or created with the assistance of large language models beginning with the 2025-2026 cycle [17]. DriveThruRPG and its sister marketplaces have required labelling of AI-generated art and in 2023 stopped accepting commercial content primarily written by AI language generators, later adding a customer-facing filter with "Handcrafted" and "Contains AI-Generated Content" labels that rely on publisher self-reporting [18][32]. Wizards of the Coast says it requires D&D contributors to refrain from using generative AI to create final D&D products [20]. If you plan to sell, check the target platform's rules before you decide how much to generate.

What about the models' training data? The headline legal development is Bartz v. Anthropic. In June 2025 the court held that training on books was fair use but that keeping over 7 million pirated books in a central library was not; the parties settled for $1.5 billion, and the court gave final approval on 20 July 2026, with payments of roughly $3,000 per book across nearly 465,000 works [16]. The ruling is about acquisition and storage, and is not a blanket answer on whether outputs can infringe. It also leaves open the practical question that came up in the forum thread: whether you should upload a published book or adventure into a hosted model. That is a terms-of-service and licensing question for the material and the provider, and the objections are real [21].

What about your own text? Check your provider's data terms. A privacy-first writing product such as NovelAI advertises encrypted prompts and anonymity [31]; a general chat product may use conversations to improve models unless you opt out or use the API. We did not audit these terms and do not summarise them here.

Image models for maps and concept art

World building ends up in pictures: a regional map, a city skyline, a portrait of the baron. For this we have better data than for text, because image models have large public arenas.

The leaderboard picture

On Artificial Analysis's text-to-image arena (accessed 9 October 2026), OpenAI's GPT Image 2.5 Sunburst leads at 1198 Elo, followed by GPT Image 2.5 Flare (1191), GPT Image 2 (1172), Google's Nano Banana 2.1 (1160), Grok Imagine Image 2.0 (1156), Microsoft's MAI-Image-2.6 (1151), Nano Banana 2 (1126), FLUX 3 Image (1117) and Meta's Muse Image (1115) [4]. LMArena's text-to-image board agrees on the top three; positions four to ten differ, for instance Nano Banana 2.1 is fourth on AA and sixth on LMArena, and is flagged pre-release there [5]. The two Elo scales are not comparable and we do not put them on one axis.

Text-to-image Elo vs price per imageThe top model costs about six times the fourth for 38 Elo points more.
Text-to-image Elo vs price per image1,0001,0501,1001,1501,20000.050.10.150.20.25Artificial Analysis arena EloPrice per image ($)GPT Image 2.5 Sunburst (max)GPT Image 2.5 Flare (max)GPT Image 2 (high)Nano Banana 2.1Grok Imagine Image 2.0MAI-Image-2.6Nano Banana 2 (Gemini 3.1 Flash Image)FLUX 3 ImageMuse ImageMAI-Image-2.6-FlashQwen-Image-3.0-ProQwen-Image-3.0Ideogram 4.0 (Quality)OpenAIGooglexAIBlack Forest LabsAlibaba

AA's Elo scale is not comparable with LMArena's. The arena measures general preference, not map legibility or geographic plausibility. Qwen-Image-2.1 (1035 Elo, open weights, no API price) is omitted from the plot. Source: Artificial Analysis text-to-image arena, accessed 2026-10-09.

The price axis changes the picture. GPT Image 2.5 costs about $0.21 an image on AA's listing; Nano Banana 2.1 is $0.034 (matching Google's 1K price on its pricing page), Muse Image $0.01 and MAI-Image-2.6 $0.039 [4][6]. The top of the board costs six times as much as the fourth place for a drop of 38 Elo points. A rough map set is a good example: 100 generations (20 maps at five attempts each) is 100 x $0.211 = $21.10 at the top model and 100 x $0.034 = $3.40 on Nano Banana 2.1. Those are our arithmetic from list prices, and OpenAI's own pricing page was unreachable for us, so the GPT Image figure is AA's [4].

Open weights

If you want to run images locally or avoid per-image fees, AA lists open-weights entries: Qwen-Image-2.1 (1035 Elo, released 20 September 2026), Ideogram 4.0 Quality (1012), FLUX.2 [dev] (1000) and Qwen Image Max 2512 (999) [4]. They trail the best closed model by about 160 to 200 Elo points, which is large on this scale, but they are free of per-image charges and of vendor content filters. For a campaign you are not publishing commercially, that trade can be right.

What the arena does not measure

The arenas ask voters which image they prefer for a prompt. That tells you about aesthetics and prompt-following in general, not about whether a map is legible, whether labels are readable, whether coastlines and rivers obey geography, or whether the same character looks the same across ten images. We found no benchmark for any of those, and no map-specific leaderboard. For maps, this matters: a generated "world map" is a picture of a map, not a map. The rivers may not flow downhill and the borders may not match the political geography in your notes.

The tool market reflects this. Dungeon Alchemist has marketed AI-assisted map making since its 2021 Kickstarter, where you draw rooms and pick a theme and the software places walls, objects and lighting; text-to-map appears only as a user feature request in the material we found [26]. A 2026 iOS app, BattleMap Maker, advertises generation of battle maps from text descriptions [26]. We could not verify how well either performs, and the list of "AI battlemap generators" on the web is dominated by marketing pages.

Our reading is that image models earn their place for mood: a moodboard for the capital, a portrait for the exiled prince, a rough sketch to hand an artist. They are weaker for anything players must navigate. The map the table plays on should be drawn or edited by a person, with generated art as a reference. For more on the art side, see our post on the best AI for game art.

A practical workflow

This is our recommendation, built from the evidence above. We have not run it as a controlled experiment, and your mileage will depend on your setting and your tolerance for editing. The principle is one sentence: the human owns canon, the model proposes, and a small amount of structure stops contradictions.

1. Write the premise by hand. Two paragraphs: what the world is about, the central tension, the tone, three things that are never true here. Models fill gaps with the average of their training data, and a specific premise is what keeps the average out.

2. Keep one canonical file per domain. Geography, factions, timeline, cast, magic rules, each as plain text or Markdown in a folder or a wiki (LegendKeeper, Kanka, Campfire, World Anvil, or a folder of notes). A short "facts that must not change" file lists dates, deaths, borders and rules. Keep this in your own tool, not in a chat history, which is lost or truncated.

3. Use the model for volume, not for canon. Ask for twenty names for a minor port, ten rumours, five competing origin stories for the ruins. Paste in the premise and the relevant canon file, not the whole bible. Pick, edit, and write the result into the canon yourself. Nothing generated becomes canon until a person moves it.

4. Retrieve, don't dump. When you need consistency, paste the three relevant entries and ask the question. Chroma's results suggest this is better than sending everything [3]. If you use a notes app such as Obsidian with an agent such as Claude Code, community guides recommend a root instruction file that describes the vault and, for safety, read-only source notes with agent output going to a separate folder until you have reviewed it [27]. Those guides are about personal knowledge management in general, not RPG canon, so adapt the idea rather than the details.

5. Run a contradiction pass. After a big prep session, give the model the facts file and the new entries and ask it to list conflicts and not resolve them. We have not seen a published benchmark of how good models are at this task, so treat the output as a prompt for your own check, not a verdict. A cheap model is fine for the first pass; the AA-LCR results suggest you do not need a flagship for retrieval-style checks [1].

6. Track mutable state outside the model. If the world will be played, with live campaign state like HP, gold, faction standing and who knows what, keep it in a sheet or tool. The reports we reviewed point the same way, and RPGBench's finding is that models struggle to keep game mechanics consistent and verifiable in long games [22]. See the AI RPG post for what products like Voyage do about this.

7. Use images for mood and write maps by hand. Generate a handful of references with a mid-priced model (Nano Banana 2.1 at about $0.034 an image in our arithmetic above), then draw or edit the playable map. Keep a note of what was generated if you plan to publish, because platforms ask.

8. Decide in advance what you will disclose. If you might submit to a marketplace or award, read its rules first. The ENNIEs and Paizo's programs, for example, bar AI content outright [17][19].

If you are building tools rather than a setting, the same architecture shows up: a structured store of facts, retrieval, and a model that writes against the retrieved pages. Costs for the model side are in our cost breakdown, and a map of where each model fits is in our guide to the best AI models for game development. CreateGame.ai, which is pre-launch, is aimed at generating a playable game from a sentence, and these consistency limits are the design constraints we think about most.

Where it breaks

A short list, so you can recognise failure early.

Drift. The model contradicts an earlier fact without noticing. It is the dominant failure in the accounts we read and it grows with context length [3][21].

Genericness. Output regresses to the mean: the grim fantasy kingdom with a cursed forest. Every source that comments on quality, from the 2023 DM post to the 2025 marketplace blog to Larian's "3/10", describes this [13][23][24].

False authority. A confident paragraph reads like canon and gets pasted in. The workflow above exists to stop this.

Social cost. In a community that is hostile to generative AI, admitting to it has a price. GDC's numbers and the Larian episode show how fast sentiment is moving [12][13].

None of these is a reason not to use the tools. They are reasons to decide what you will let them touch.

How we reviewed this

We reviewed published sources only. We did not test any of the tools or models named here, and nothing in this post is a hands-on result. We used: Artificial Analysis model and image leaderboard data and the AA-LCR page (accessed 9 October 2026, via our shared models snapshot); LMArena's image board; Chroma's Context Rot report; the GDC 2026 State of the Game Industry write-up; vendor pages for LegendKeeper, Campfire, Kanka and Friends & Fables; the US Copyright Office report via the Copyright Alliance summary; publisher and awards policies from Paizo, the ENNIEs, DriveThruRPG and Wizards of the Coast; settlement coverage for Bartz v. Anthropic; the Paizo forum thread; and individual accounts by Brandon Sanderson, Larian and a DM blogger.

What we could not verify: World Anvil's current AI features and policy (its site returned an error, and we have only a 2024 community thread and a competitor's claim); Kanka's and Campfire's complete feature and changelog history; Larian's primary statements (we relied on secondary coverage); the Copyright Office report itself (we read a summary); the current text of DriveThruRPG's policy; Wizards of the Coast's current wording (our source is secondary and undated); the quality of Dungeon Alchemist and BattleMap Maker; and OpenAI's image pricing page, which was unreachable. Image arena Elo scales on AA and LMArena differ and are not compared. The cost tables are our estimates from list prices and a 0.75 words-per-token rule of thumb.

FAQ

What is the best AI for world building? There is no benchmark for it. Use a strong general model such as Claude or GPT for brainstorming and test two or three on your own setting; for lookups across your notes, long-context scores are tightly bunched, and cheap models score close to flagships on AA-LCR [1].

Does World Anvil have AI? We could not confirm. A 2024 community thread reports the team calling AI out of scope, and a Sudowrite article claims an AI feature called the Sage exists [10][11]. Check World Anvil's current feature pages.

Is it okay to publish AI-generated lore or art? Legally, wholly AI-generated material is not copyrightable in the US, though human selection and edits can be [15]. Commercially, Paizo and the ENNIEs bar AI content, DriveThruRPG requires labels, and Wizards restricts it for D&D products [17][18][19][20].

Can AI make a usable fantasy map? For mood and reference, yes. For navigation, we would not rely on it: no benchmark tests legibility or geographic plausibility, and the arenas only measure preference [4][5]. Draw the playable map yourself.

How much does it cost to give an AI my whole campaign setting? By our estimate, a 100,000-word bible is about 133,000 tokens, or $0.27 per request on Claude Sonnet 5.5 and about $0.013 on Claude Haiku 5.5 at list prices [1].

Which image model is best for RPG concept art? On AA's arena, GPT Image 2.5 Sunburst leads (1198 Elo) but costs about $0.21 an image; Nano Banana 2.1 (1160) is about $0.034 [4]. Qwen-Image-2.1 is the top open-weights entry we found [4].


If you want to follow what we are building at CreateGame.ai, you can join the waitlist at https://creategame.ai.

Sources

  1. Artificial Analysis, model leaderboard, accessed 9 Oct 2026 (via the CreateGame.ai models snapshot): https://artificialanalysis.ai/leaderboards/models
  2. Artificial Analysis, Long Context Reasoning (AA-LCR) v1.1: https://artificialanalysis.ai/evaluations/artificial-analysis-long-context-reasoning
  3. Chroma, "Context Rot: How Increasing Input Tokens Impacts LLM Performance," 14 Jul 2025: https://trychroma.com/research/context-rot
  4. Artificial Analysis, text-to-image arena, accessed 9 Oct 2026: https://artificialanalysis.ai/image/leaderboard/text-to-image
  5. LMArena, text-to-image leaderboard, accessed 9 Oct 2026: https://arena.ai/leaderboard/image/text-to-image
  6. Google, Gemini API pricing, accessed 9 Oct 2026: https://ai.google.dev/gemini-api/docs/pricing
  7. LegendKeeper, home page and FAQ: https://www.legendkeeper.com/
  8. Campfire, home page: https://www.campfirewriting.com/
  9. Kanka, features: https://kanka.io/features
  10. World Anvil community suggestion thread, "AI Novel Writer": https://www.worldanvil.com/community/voting/suggestion/182ee13a-ceb8-4bb8-b53d-904bfb2a326e/view
  11. Sudowrite, "What is the best AI for worldbuilding? We tested the top tools" (vendor-written, models dated): https://sudowrite.com/blog/what-is-the-best-ai-for-worldbuilding-we-tested-the-top-tools/
  12. GDC, "2026 State of the Game Industry": https://gdconf.com/article/gdc-2026-state-of-the-game-industry-reveals-impact-of-layoffs-generative-ai-and-more/
  13. GosuGamers, Larian CEO responds to Divinity generative AI backlash: https://www.gosugamers.net/entertainment/news/77735-larian-ceo-responds-to-divinity-gen-ai-backlash-says-game-won-t-include-ai-content
  14. Brandon Sanderson, "The Hidden Cost of AI Art" keynote transcript: https://www.brandonsanderson.com/blogs/blog/ai-art-brandon-sanderson-keynote
  15. Copyright Alliance, summary of the US Copyright Office AI report Part 2: https://copyrightalliance.org/ai-report-part-2-copyrightability/
  16. Authors Guild, final approval of the Anthropic copyright settlement: https://authorsguild.org/news/court-grants-final-approval-anthropic-copyright-settlement/
  17. ENNIE Awards, revised policy on generative AI usage: https://www.ennies.org/?p=2899
  18. ENWorld, DMs Guild and DriveThruRPG AI policy thread: https://www.enworld.org/threads/dms-guild-and-drivethrurpg-ban-ai-written-works-requires-labels-for-ai-generated-art.698936/
  19. Paizo, "Paizo and Artificial Intelligence": https://paizo.com/blog/paizo-and-artificial-intelligence
  20. Bell of Lost Souls, WotC updated statement on AI: https://belloflostsouls.net/?p=519090
  21. Paizo forums, "Anyone used AI as a GM with success?" (30 Sep to 2 Oct 2026): https://paizo.com/threads/rzs8tv4g
  22. Yu et al., "RPGBench," arXiv 2502.00595: https://arxiv.org/abs/2502.00595
  23. Jason Preston, "ChatGPT as a DM copilot," 19 Apr 2023: https://jasonp.substack.com/p/chatgpt-as-a-dm-copilot
  24. Wayline, "AI game worldbuilding," 13 Jun 2025 (marketing content): https://www.wayline.io/blog/ai-game-worldbuilding
  25. TechCrunch, Voyage launch, 21 Apr 2026: https://techcrunch.com/2026/04/21/voyage-is-an-ai-rpg-platform-for-creating-custom-gaming-worlds-with-ai-generated-npc-interactions
  26. Dungeon Alchemist feature request page and BattleMap Maker listing: https://dungeonalchemist.featurebase.app/p/ai-text-to-map-generation and https://apps.apple.com/app/id6756099903
  27. Claude Code for Marketers, "Obsidian as your second brain" (general workflow guide): https://claudecodeformarketers.com/blog/obsidian-as-your-second-brain/
  28. Friends & Fables, plan comparison: https://fables.gg/blog/old-gregs-tavern-vs-friends--fables-plan--feature-comparison
  29. Digital Applied, "Which AI model writes best? Leaderboards disagree," 31 Aug 2026: https://www.digitalapplied.com/blog/which-ai-model-writes-best-leaderboards-disagree
  30. Surge AI, Hemingway-bench: https://surgehq.ai/blog/hemingway-bench-ai-writing-leaderboard
  31. NovelAI home page: https://novelai.net/
  32. Hyperfictive, "The Label is Broken" (critique of DriveThruRPG AI labels): https://hyperfictive.substack.com/p/the-label-is-broken