AI NPC Dialogue in 2026: Which Models Are Fast and Cheap Enough for Real Time
An AI NPC lives or dies on a number that most model leaderboards bury: how long the player waits before the character starts to answer. We went through Artificial Analysis' latency data, vendor documentation, developer write-ups and the Steam and press reception of the games that have shipped, and the picture is more specific than "use a small model." Speed, cost, memory and safety each rule out different options, and the games that work have mostly decided that the language model should do less than the marketing suggests.
Key takeaways
- Time to first token (TTFT) is the filter. In Artificial Analysis' 9 October 2026 data, GPT-6 Luna's non-reasoning variant starts answering in 0.83 seconds; the same model at maximum reasoning effort takes 98.8 seconds. Using the wrong variant is the commonest way to get NPC latency wrong.
- Eleven hosted entries in AA's table start answering in under a second when reasoning is off or minimal, and most finish an 80-token line in roughly 1.2 to 2 seconds by our arithmetic. That is acceptable for text chat and borderline for voice.
- Cost is no longer the obstacle for cheap models: our estimate for a 1,200-token prompt and an 80-token reply is $0.16 to $0.46 per 1,000 lines on the cheapest options, against about $3.20 on Claude Sonnet 5.5. It becomes the obstacle when a game succeeds.
- Voice raises the bar. A published budget for a conversational voice agent allocates 800 ms across speech recognition, the model and speech synthesis, and leaves the language model about 400 ms.
- Shipped examples are mostly indie or optional features. Where Winds Meet limits the chatbot to side characters, PUBG Ally runs a 2-billion-parameter model on the player's GPU, and Ubisoft's Teammates is still a closed playtest. Steam reception for the games built around AI conversation runs from 48% to 63% positive.
- Safety failures are public and recurring: players talked Where Winds Meet's NPCs into completing quests with invented claims, and Fortnite's AI Darth Vader was coaxed into profanity within days.
- We did not run any of these systems. Every figure is from a published source, with its date, and our own arithmetic is labelled.
What an AI NPC has to do in a few hundred milliseconds
People in conversation leave very short gaps. The best-known cross-language study, Stivers and colleagues in PNAS in 2009, found that all ten languages studied minimise silence between turns, and that average gaps differed between languages only within a range of 250 ms around the cross-language mean [3]. The figure usually quoted from that line of work is around 200 ms (we did not re-check the full text). No game will hit that with a cloud model. What matters is what the player forgives, and that depends on the channel.
Text chat. A player typing to an NPC expects a pause, and streaming the reply token by token hides most of it. The constraint is reading speed. Brysbaert's 2019 meta-analysis of 190 studies puts adult silent reading of English non-fiction at 238 words per minute [4], about four words a second, or roughly five tokens a second at an assumed 1.3 tokens per word. Every model in this post generates more than ten times faster than that. For text, throughput is a non-issue and TTFT is everything: the player sees nothing until the first token arrives.
Voice. Speaking runs at roughly 150 words a minute (a commonly cited rate), so a model needs only 3 or 4 tokens a second to keep a text-to-speech engine fed. Again, throughput is cheap. The problem is the chain: speech recognition, then the model, then speech synthesis, with the first audio not starting until all three have started. A vendor guide from Smallest.ai, dated 28 September 2026, allocates an 800 ms time to first audio as 50 ms for audio capture, 150 ms for transcription, 400 ms for the model's first token, 150 ms for the first chunk of speech and 50 ms of network [5]. It adds that one 300 ms tool call pushes the total to roughly 1,100 ms. That guide is written by a voice-AI vendor and describes phone assistants, not games, so we use it as a reference point and not a standard.
Action games. Anything where the NPC must react in combat cannot wait on a language model at all. PUBG Ally's designers split behaviour into a behaviour tree running at game tick rate for movement, aiming and immediate combat, and a language model handling intent, coordination and speech [10]. We come back to that split below because it is the most transferable idea in this post.
Cinevva's NPC guide, published by a game-tool vendor, says responses under about 800 ms feel conversational and that cloud round trips often add one to two seconds [9]. Both are claims, not measurements.
Reading the latency numbers without being fooled
Artificial Analysis (AA) publishes a "Latency First Chunk" column and a "Total Response" column on its model leaderboard alongside speed in tokens per second [1]. Our shared model snapshot flags a trap in them. For reasoning models, the first chunk can include the model's thinking time, which at maximum effort can run to hundreds of seconds (Claude Opus 5.5 shows 662 seconds in the leaderboard table we read on 9 October; the snapshot's median rounds to 683). Those numbers are real, and useless for a game. A character who thinks for ten minutes before saying hello has not failed a latency test, she has failed a design test.
The effect is easy to see inside a single model family. GPT-6 Luna appears on the leaderboard at six effort settings.
AA's first-chunk latency for reasoning variants includes thinking time. AA lists no non-reasoning Haiku 5.5 entry, and no medium entry for GPT-6 Luna, so those points are absent. Do not use the high-effort numbers for NPC claims; they show why variant choice matters. Source: Artificial Analysis leaderboard (Latency First Chunk, speed), accessed 2026-10-09.
At non-reasoning, Luna's first chunk arrives in 0.83 seconds. At low effort it is 2.83 seconds, at high 16.6, at extra-high 23.6 and at maximum 98.8. Claude Haiku 5.5 shows the same slope with a worse floor: 9.7 seconds at low effort, 13.4 at medium, 22.7 at high and 289 at maximum [1]. AA's table lists no non-reasoning Haiku 5.5 entry, so we cannot say from its data whether Anthropic's smallest model can respond quickly with thinking disabled. We flag that rather than guess; our cost post priced NPC chat on Haiku-class models, and anyone following that estimate should confirm the effort setting they will actually run.
The rule we draw: for NPC dialogue, pick the non-reasoning or lowest-effort entry, and check that your provider lets you turn thinking off. An NPC does not need to solve olympiad problems. Hidden reasoning also costs money, which we covered in the cost post: 400 thinking tokens before an 80-token reply multiplies the output bill by six.
Two more caveats. AA's numbers use its own standard test prompt, whose length we did not confirm, and a real NPC prompt with persona, memory and recent dialogue will usually be longer, which raises TTFT. And the speed column is for hosted inference, which says nothing about running a model on a player's machine.
Which models clear the bar
We pulled every leaderboard row that is either labelled non-reasoning or is the lowest-effort setting of a family, then kept those with TTFT near or under about one second, plus a few instructive failures. To estimate how long a typical line takes, we added the time to generate 80 tokens (about two short sentences) at each model's median speed. That is our arithmetic: TTFT + 80 / tokens per second.
Estimate = TTFT + 80 / median tokens per second, from AA's hosted-endpoint medians. Open-weights models run locally will differ. AA's test prompt length was not confirmed, and a real NPC prompt with memory will be longer. Source: Artificial Analysis leaderboard (Latency First Chunk, speed), accessed 2026-10-09.
| Model (AA variant) | TTFT (s) | Speed (tokens/s) | Our estimate: 80-token line (s) | List price in/out ($ per 1M) |
|---|---|---|---|---|
| Qwen3.5 9B (non-reasoning) | 0.70 | 81 | 1.69 | not in our snapshot |
| Mistral Small 4 (non-reasoning) | 0.72 | 151 | 1.25 | not in our snapshot |
| GPT-5.6 Terra (non-reasoning) | 0.77 | 91 | 1.65 | $2 / $12 |
| Gemma 4 E4B (non-reasoning) | 0.82 | 25 | 4.02 | open weights |
| GPT-6 Luna (non-reasoning) | 0.83 | 126 | 1.46 | $0.10 / $0.50 |
| Claude Sonnet 5.5 (low) | 0.88 | 100 | 1.68 | $2 / $10 |
| GPT-5.5 Instant | 0.95 | 139 | 1.53 | not in our snapshot |
| Nemotron 3 Nano (non-reasoning) | 0.97 | 227 | 1.32 | open weights |
| Gemma 4 26B A4B (non-reasoning) | 1.02 | 88 | 1.93 | open weights |
| DeepSeek V4.1 Flash (non-reasoning) | 1.12 | 227 | 1.47 | $0.30 / $1.20 |
| Qwen3.6 35B A3B (non-reasoning) | 2.04 | 146 | 2.59 | open weights |
| GPT-6.1 Sol (low) | 2.80 | 47 | 4.50 | $2 / $10 |
| Qwen3.8 27B (non-reasoning) | 3.75 | 52 | 5.29 | $0.50 / $3.00 |
Source: Artificial Analysis leaderboard, read 9 October 2026 [1]; prices from the AA data in our shared snapshot for models it covers. The Gemma and Nemotron speeds are for AA's hosted endpoints; the open-weights models can be run locally at different speeds.
Four readings of the table.
First, there is a real cluster of options that start in under a second with reasoning off: Qwen3.5 9B, Mistral Small 4, GPT-5.6 Terra, GPT-6 Luna, Claude Sonnet 5.5 at low effort, GPT-5.5 Instant and Nemotron 3 Nano. If you stream tokens to a text box, these all work. For voice, a 0.7 to 0.9 second TTFT alone exceeds the 400 ms the Smallest.ai budget allows the model, so a voice game either accepts about a second of lag, hides it with an animation or filler line, or hosts the model close to the player.
Second, intelligence and speed run in opposite directions. AA's Intelligence Index v4.3.2 gives GPT-6 Luna's non-reasoning variant 18, against 56 for Claude Sonnet 5.5 at maximum effort and 57.6 for Opus 5.5 [1]. Sonnet 5.5 at low effort scores 36. The cheap fast models are not stupid at conversation, which is mostly an instruction-following and style task, but they are poor at remembering a detailed lore bible or resisting manipulation. The index measures work tasks such as terminal use and long-document reasoning, not character voice, and we know of no published benchmark for NPC believability. That gap deserves more attention than it gets.
Third, speed figures for small open models mislead. Gemma 4 E4B shows 25 tokens a second on AA's hosted endpoint, which reflects the provider's setup, not what a desktop GPU would do. None of the sources we found published on-device speed for the games discussed below.
Fourth, "low effort" does not guarantee fast. GPT-6.1 Sol at low effort takes 2.8 seconds to first token and runs at 47 tokens a second, while Claude Sonnet 5.5 at low effort starts in 0.88 seconds. Check the specific variant, not the family.
What a line of dialogue costs
The cost post worked out monthly bills for a game with 10,000 daily players; we will not repeat that arithmetic. A simpler unit for choosing a model is the cost of 1,000 lines, using the same assumptions: 1,200 input tokens per exchange (persona, memory summary, recent dialogue, player message) and 80 output tokens.
Our estimate: (1,200 x input price + 80 x output price) / 1,000,000 x 1,000. List prices per 1M tokens (input/output): Luna $0.10/$0.50, GLM-5.3-Flash $0.15/$0.50, DeepSeek V4.1 Flash $0.30/$1.20, Qwen3.8 27B $0.50/$3.00, Sonnet 5.5 $2/$10, Terra $2/$12. Excludes speech synthesis, which can cost more than the model. Source: Artificial Analysis list prices via CreateGame.ai model snapshot, accessed 2026-10-09.
The formula is (1,200 x input price + 80 x output price) / 1,000,000 per line. For GPT-6 Luna at $0.10 and $0.50 per million that is $0.00016 a line, or $0.16 per 1,000. DeepSeek V4.1 Flash is $0.46. Qwen3.8 27B is $0.84. Claude Sonnet 5.5 and GPT-5.6 Terra are $3.20 and $3.36. These are list prices from the AA data on 9 October 2026, before caching, and ignore speech, which we price below.
Two consequences. Under a dollar per thousand lines, a player who has twenty exchanges a day costs well under a cent. The expensive part of a successful game is not the language model; it is the voice. And the gap between a $0.16 and a $3.20 line is a factor of twenty, which means model choice is a design decision about how much of your game's dialogue deserves a frontier model. Several of the shipped games below answer "very little."
Voice pipelines
Most players say they want to talk to NPCs, and voice is where the costs and the latency pile up.
Speech synthesis. Inworld's documentation lists two current text-to-speech models. Realtime TTS-2 has a server-side P90 time to first byte of 100 ms, and TTS-2 Flash 20 ms, described as five times faster [6]. These are the vendor's own figures, for synthesis on its servers, not end-to-end audio at the player's ear. Inworld retired its older TTS models in June 2026 [6]. Its pricing, as recorded by an independent site on 27 September 2026, was $25 per million characters for TTS-2 and $15 for Flash on the free on-demand plan, falling to $12.50 and $7 on a $1,500 monthly Growth plan, with speech-to-text at $0.15 an hour [7]. Arcanum, which wrote that piece, runs a paid service for developers and its own benchmark, so weigh that.
We would not use any other vendor's latency table. The independent TTS benchmark we looked at (Codesota) said on its page of 7 October that it had withdrawn its earlier April figures, so the "75 ms" and "90 ms" figures that circulate in voice-AI blog posts now rest on vendors' own claims.
Cost of voice. An 80-token line is about 320 characters (assuming roughly four characters per token). At $25 per million characters that is $0.008, or $8 per 1,000 lines, fifty times the language-model cost on Luna; at the Flash rate of $15 it is $4.80 per 1,000. At the Growth plan's $7 it is $2.24. Our arithmetic, and a good reason that text chat is the economical default.
Streaming is the whole trick. A pipeline that waits for the full transcript, then the full reply, then the full audio will never reach 800 ms. Real systems stream partial transcripts to the model before the player has finished speaking, stream tokens to the synthesiser sentence by sentence, and start playback on the first chunk. This is why vendors' stage budgets sum to more than the perceived delay. It also constrains design: a model that writes a long preamble before the useful sentence adds audible delay, so NPC prompts should demand short replies.
Shipped voice games. Suck Up! (Proxima, 1 October 2025) is built around talking a vampire's way into homes and discloses a third-party connection to OpenAI's ChatGPT on its Steam page [19]. Whispers from the Star (Anuttacon, August 2025) offers voice conversation with a single astronaut character [8]. Ubisoft's Teammates lets players give squadmates voice commands, and its helper Jaspar can highlight enemies, give lore and change settings [13]. Ubisoft gives no latency figures or model names in its own post; later press reports a link to Google Gemini that we could not confirm from a primary source [15]. Whispers from the Star's Steam page, as snapshotted on aggregators, shows Very Positive overall (1,611 reviews) with recent reviews lower, and players complained of repetitive and argumentative AI dialogue and crashes [22]. A published review summary also said the AI responses are fresh at first and then lose pace without game structure.
Small models and on-device inference
The case for running the model on the player's machine is cost (zero marginal cost per line), latency (no network), and offline play. The cost is VRAM, quality and engineering.
The clearest production account is Krafton's PUBG Ally, a "co-playable character" in PUBG: Battlegrounds, described in an NVIDIA developer post on a Krafton Q&A of 25 June 2026. It chains NVIDIA speech recognition, a quantised Mistral-NeMo-Minitron-2B small language model (2 billion parameters) and a custom in-house text-to-speech model, runs on the player's GPU in whatever VRAM PUBG leaves free, and works on cards with as little as 8 GB. It handles English, Korean and Chinese on one map (Sanhok) in one mode (AI Duo), and the public beta ran in Arcade Mode from 17 to 30 June 2026. Evaluation included large-scale playtests with feedback from more than 1,000 players. Krafton gives no latency numbers: the post says only that removing the network round trip made response times more predictable [10].
The design choices matter more than the model. Krafton uses the behaviour tree as "System 1" and the language model as "System 2," keeps prompts KV-cache-friendly so the model does not re-read the world every turn, and grounds the model through tool calls that return plain-text descriptions of the live match, not by stuffing the state into memory [10]. Scope is narrow on purpose.
Krafton's inZOI, in early access since March 2025, uses a 0.5-billion-parameter Mistral NeMo Minitron model through NVIDIA ACE for its "Smart Zoi" characters, which revise their daily schedules and show their inner thoughts [11]. NVIDIA has also added Qwen3-8B to ACE's in-game inferencing plugin [12]. A newer trade article describes Unreal Engine plugins with a 4-billion-parameter Qwen model, but we could not verify its date or source and do not rely on it.
The AA data adds a hosted reference point: Qwen3.5 9B responds in 0.70 seconds and Gemma 4 E4B in 0.82 seconds on AA's endpoints, with Qwen3.5's 0.8B and 2B and Gemma 4 E2B variants listed without data [1]. What we lack is anything that measures how a 2B to 9B model performs as a character against a frontier model, in a blind player test. The 2B model in PUBG Ally works because it is asked to do so little: take a command, pick a behaviour, say something short.
Shipped games and what we know about them
Most of the "AI NPC" announcements of 2023 and 2024 never became products, and Inworld, the best-known platform, has retired its Character Studio and repositioned as developer infrastructure for voice and inference [7]. Convai, which still sells game-focused tools, listed an Indie plan at $29 a month for 3,000 interactions in Cinevva's July 2026 guide [9]; we could not read Convai's own pricing page, so check it. This is what we could verify about games that are out.
| Game | Studio | Status | What the AI does | Model | Reception |
|---|---|---|---|---|---|
| Where Winds Meet | Everstone / NetEase | Released Nov 2025 | Text chat with a subset of side NPCs; friendship tiers | "A large language model" [16] | Split; 80% positive Steam reviews reported at launch [17] |
| inZOI | Krafton | Early access since Mar 2025 | "Smart Zoi" characters plan and reflect | 0.5B Minitron on-device [11] | Not measured by us |
| PUBG Ally | Krafton | Public beta 17-30 Jun 2026 | Voice teammate | 2B Minitron on-device [10] | 1,000+ playtesters; no published scores |
| Suck Up! | Proxima | 1 Oct 2025, $8.99 | Voice persuasion of residents | ChatGPT connection disclosed [19] | 63% of 211 positive, Mixed [19] |
| Vaudeville | Bumblebee Studios | Early access Jun 2023; 1.0 28 Nov 2025 | Detective questions suspects by text or voice | Not stated | 48% of 283 positive, Mixed [20] |
| Whispers from the Star | Anuttacon | 14 Aug 2025 | Voice talk with one astronaut | Not stated | Very Positive overall (1,611 reviews) [22] |
| World Apart | Nuverse (ByteDance) | Early access 23 Sep 2026 | Store page says NPCs take typed chat | Not stated | Too new |
| RyzaChat:AI | SpiralAI under Koei Tecmo licence | Aug 2026, several regions | Typed commands drive exploration and battles | Separate "Game Master AI" | Not measured |
| Teammates | Ubisoft | Closed playtest, a few hundred players | Voice-commanded squadmates | Not stated | Not published [13] |
Sources are cited in the table; the list of games and dates comes mostly from the Arcanum RPGs roundup of 26 September 2026 [8], which says it checked primary sources and did not play the games, and which also sells developer services. Two lines in the table are weakly sourced: the 80% for Where Winds Meet is a secondary report of a launch-week figure, and RyzaChat:AI and World Apart rest on Arcanum alone.
Three patterns. The games where the language model is the whole game (Suck Up!, Vaudeville) are the ones with mixed reception. The games where it is an optional layer on a large game (Where Winds Meet, inZOI) get the louder discussion and the bigger audiences. And the biggest studios are still at prototype: Ubisoft's own post describes a closed playtest of a few hundred players [13], while a Ukrainian summary of Ubisoft's annual report says none of the experiments has reached a commercial release [15].
Vaudeville and Suck Up! were read from their Steam pages on 2026-10-09. The Where Winds Meet figure is a secondary report of launch-period reviews, and the game uses chat only for some side characters, so it is not a measure of the AI feature. Positive share is not a verdict on the AI. Source: Steam store pages; Notebookcheck (Where Winds Meet), accessed 2026-10-09.
Memory and consistency
An NPC who forgets you after two minutes is a chatbot. The foundational design is still the Stanford "generative agents" paper of Park and colleagues (arXiv 2304.03442, April 2023), which stored each agent's experiences as a natural-language memory stream, synthesised them into higher-level reflections, and retrieved them to plan; its ablations showed that removing observation, planning or reflection each reduced believability, in a sandbox of 25 agents [23]. Almost every NPC memory system since is a variant of that loop.
In practice the shipped systems are more modest. PUBG Ally keeps a short-term memory for the current match and a long-term memory of the player's profile and match history across matches, with storage and retrieval details unpublished [10]. Where Winds Meet tracks an affinity score and restricts the model to a fixed set of game-state flags, and its main story stays hand-written [9]. The Cinevva guide's summary of what works is short replies, limited knowledge and a fixed action list [9], which is the same advice every developer talk gives.
Three problems recur. Consistency: unconstrained models invent facts, drift from lore and spoil plots, and the more you put in the prompt the more each line costs and the slower TTFT gets. Cache design: a stable persona prefix can be cached, but re-ordering memory into the front of the prompt every turn loses the saving. And evaluation: PUBG's team describes layered automated checks and A/B playtests because outputs are non-deterministic [10]. A writer at Frisson Labs argues, as an interested builder of AI companions, that models are often too knowledgeable and too helpful to be believable characters, falling back into "know-it-all" or "therapist" modes [26]. We have not seen a published measurement of that, but it matches the Whispers from the Star complaint about argumentative, repetitive dialogue.
Safety, jailbreaks and rules
The best public evidence that players will attack your NPC comes from Where Winds Meet. Within days of launch, players found that repeating the NPC's own lines back at it eventually convinced it a quest was done, and that typing invented actions in parentheses made the model believe them (one player "gifted" an imaginary cat, and it accepted). Another coverage piece described a player persuading an NPC he had fathered a child, then that the child had died, with the NPC grieving; another found that generic encouragement raised affection regardless of the character's goal, and that the model rarely held the line when conversations turned sexual, even though the protagonist can be a teenager [16]. The reporting says the developer made no official statement, though returning players found the NPCs more constrained, which the outlet reads as tightened guardrails; that is community observation, not confirmed [16]. Rock Paper Shotgun quoted the writer Dan Griliopoulos calling the chatbots "in a bad state" for being too general and breaking the fiction [18].
Fortnite's AI Darth Vader in May 2025 is the other canonical case. Players got a voice-synthesised Vader, reportedly built with ElevenLabs from archived recordings of James Earl Jones, to swear and use slurs within days; Epic said its "filters did not catch a specific variation of an expletive" and hotfixed it [25]. SAG-AFTRA separately filed an unfair labour practice charge over the use of the voice [25].
The structural lesson from both is that a model with a game-state tool and no validation will believe whatever the player types. The design response is a fixed list of allowed actions, server-side validation of any state change, and treating the model's output as dialogue only. Gravitee, an API-gateway vendor, publishes a threat model for LLM NPCs covering injection through chat, item names and in-world text, and leaked system prompts [27]; it is marketing, but the attack surface it lists is the right one. We found no published measurement of how often NPC guardrails fail in shipped games.
Storefront rules apply too. Valve requires developers who ship live-generated AI content, defined as content created by AI while the game runs, to say what guardrails prevent illegal content, displays that disclosure on the store page, offers players a way to report illegal live content, and does not currently permit adult-only sexual content generated live [24]. A game that ships a chat NPC without a moderation layer risks removal. SAG-AFTRA's consent rules for AI voice work, covered in our Unity and Unreal post, matter once you voice a character.
Player reception
Player sentiment is the least favourable number in this space, and it is mostly about taste, not technology.
When Ubisoft showed NEO NPC at GDC 2024, fans reacted on social media with lines like "robotic soulless nonsense," and some pointed to recent layoffs. Ubisoft's narrative director Virginie Mosser said at the time that the characters "do not have free will" and exist to play roles in a story, and that human voice actors would still be needed [14]. In its November 2025 Teammates post, Mosser says she still writes the story and personalities and that NPCs improvise within set boundaries [13]. That framing, human authors with AI improvisation inside a fence, is where the industry has settled.
The Steam numbers are mixed. Suck Up! stands at 63% positive of 211 reviews and Vaudeville at 48% of 283 (all-time, read 9 October 2026) [19][20]. Vaudeville's developers wrote in January 2024 that reviews were "evenly split between enthusiasts and haters" and that the average could deter buyers, at about 7,500 sales [21]. Where Winds Meet's launch was far larger and the chatbots were polarising: Wccftech called them one of its best features, KitGuru called them highly divisive, and a Spanish outlet argued the game did not need a chatbot [17].
Frisson Labs' May 2026 essay gives the sharpest account of why these features have not spread: unit economics, the difficulty of measuring fun, and uncanniness [26]. Its author has a commercial stake and no data, so we treat it as a hypothesis. Our reading is that the novelty of talking to a character wears off, and the cases players praise are ones where the conversation does something in the game.
What we would build with
This is our judgment from the documents, not from builds.
A text-chat NPC in a browser or indie game. Use a non-reasoning small model through an API: GPT-6 Luna or DeepSeek V4.1 Flash for price, Mistral Small 4 or Nemotron 3 Nano if you want an open-weights model, Claude Sonnet 5.5 at low effort if the lines must be better and you can pay twenty times more. Stream the reply. Cap output at about 80 tokens. Keep a fixed action list.
A voice NPC. Budget about a second and design for it: a thinking animation, a verbal filler, a first sentence that is cheap. Price the voice before the model.
A companion in an action game. Do what Krafton did: a behaviour tree for reflexes, a 2B to 8B quantised model on the player's GPU for talk, state supplied through tools. Ship it as an optional mode.
A story-critical character. Write the lines. Use the model for the bark layer, small reactions and flavour.
Anything with a moderation burden. Add input and output filtering, validate every state change on the server, log conversations for review, and fill in Steam's live-generated content disclosure honestly.
For the development side, see our posts on the best AI for coding games and on whether AI can play video games.
How we reviewed this
We read Artificial Analysis' model leaderboard (a table of 267 rows, read on 9 October 2026) and used the Latency First Chunk, speed and index columns as published, plus list prices from our shared model snapshot. We read vendor documentation (Inworld, Smallest.ai, NVIDIA), developer write-ups, Steam store pages, and press reports on the games, and two papers. We compute an 80-token line time and a per-1,000-line cost ourselves, from stated assumptions, and label them as estimates.
We did not run any NPC system, measure any latency ourselves, play any of these games, or read a player-count figure we could trust. Gaps: no published end-to-end latency for any shipped game (PUBG Ally, Teammates and Where Winds Meet give none); no model named for most; Convai's pricing page did not load; Inworld's latency and pricing are vendor figures (with the prices read through a third party); the Where Winds Meet 80% rating comes from a secondary summary of a page we could not open; AA's test-prompt length is unconfirmed; and no benchmark measures NPC believability. Several roundups on this topic are published by vendors, including Cinevva, which sells a game tool, and Arcanum, which sells developer services; we used them for facts we could cross-check or flagged them.
FAQ
Which model is best for AI NPC dialogue? There is no single best, but for text chat the fast, cheap non-reasoning models are the practical set: GPT-6 Luna, DeepSeek V4.1 Flash, Mistral Small 4 and Nemotron 3 Nano all start answering in about a second on Artificial Analysis' data. Claude Sonnet 5.5 at low effort is better but about twenty times the cost per line.
How fast does an AI NPC need to respond? For text chat, the first token within about a second, streamed, is workable. For voice, published guides target roughly 800 ms to first audio, which leaves the language model about 400 ms; most hosted models are slower than that.
Can I run an NPC model on the player's PC? Yes. Krafton's PUBG Ally runs a quantised 2-billion-parameter model on GPUs with as little as 8 GB of VRAM, and inZOI uses a 0.5-billion-parameter model. The trade-offs are quality, memory and the engineering work.
How much does AI NPC dialogue cost? By our estimate, 1,000 lines cost about $0.16 on GPT-6 Luna and about $3.20 on Claude Sonnet 5.5 at 1,200 input and 80 output tokens, before caching. Voice synthesis can cost more than the model: roughly $2 to $8 per 1,000 lines at Inworld's listed rates. See the full cost breakdown.
Do players like AI NPCs? Reception is split. Games built around AI conversation sit between 48% and 63% positive on Steam, and Ubisoft's NEO NPC demo drew strong criticism. Optional chat with side characters, as in Where Winds Meet, divided players more than it united them.
Are AI NPCs safe to ship? Only with guardrails. Players exploited Where Winds Meet's chatbots with invented claims within days, and Fortnite's AI Vader was made to swear. Valve requires disclosure of the guardrails on live-generated content.
If you are curious about AI-built games more broadly, the waitlist for CreateGame.ai is at creategame.ai.
Sources
- Artificial Analysis, model leaderboard (Latency First Chunk, speed, Intelligence Index v4.3.2), read 2026-10-09: https://artificialanalysis.ai/leaderboards/models
- CreateGame.ai shared model snapshot (AA prices and speeds), 2026-10-09 (internal data file derived from source 1)
- Stivers et al., "Universals and cultural variation in turn-taking in conversation", PNAS 106(26), 2009, doi:10.1073/pnas.0903616106: https://experts.illinois.edu/en/publications/universals-and-cultural-variation-in-turn-taking-in-conversation/
- Brysbaert, "How many words do we read per minute?", Journal of Memory and Language, 2019: https://www.bps.org.uk/research-digest/most-comprehensive-review-date-finds-average-persons-reading-speed-slower
- Smallest.ai, "Designing Voice Assistants: STT, LLM, TTS, Tools, and Latency Budget", 28 Sep 2026 (vendor): https://smallest.ai/blog/designing-voice-assistants-stt-llm-tts-tools-and-latency-budget
- Inworld, model documentation: https://docs.inworld.ai/models
- Arcanum RPGs, "Inworld AI (2026): What Happened to the AI NPC Company", pricing checked 27 Sep 2026: https://arcanumrpgs.com/blog/inworld-ai/
- Arcanum RPGs, "Games With AI NPCs You Can Actually Talk To (2026)", 26 Sep 2026: https://arcanumrpgs.com/blog/games-with-ai-npcs/
- Cinevva, "AI NPCs and Dialogue in Games: Tools and Reality Check (2026)", July 2026 (vendor): https://app.cinevva.com/guides/ai-npcs-dialogue
- NVIDIA Developer Blog, "How Krafton built PUBG Ally, a co-playable character powered by NVIDIA ACE", June 2026: https://developer.nvidia.com/blog/how-krafton-built-pubg-ally-a-co-playable-character-powered-by-nvidia-ace/
- NVIDIA, "NVIDIA ACE autonomous game characters debut in inZOI and Naraka: Bladepoint Mobile PC version": https://www.nvidia.com/en-eu/geforce/news/nvidia-ace-naraka-bladepoint-inzoi-launch-this-month/
- GamesBeat, "Nvidia ACE adds open source software for on-device AI NPCs in PC games" (Qwen3-8B): https://gamesbeat.com/?p=308553
- Ubisoft News, "Teammates", 21 Nov 2025: https://news.ubisoft.com/de-de/article/3mWlITIuWuu0MoVuR6o8ps
- NME, "Ubisoft's new NPC AI has been trashed by fans", 20 Mar 2024: https://www.nme.com/news/ubisofts-new-npc-ai-has-been-trashed-by-fans-3604577
- Hitmarker on Teammates; Outlook Respawn (May 2026) and Mezha on Ubisoft's annual report (secondary): https://hitmarker.net/news/ubisoft-reveals-teammates-a-generative-ai-experiment-featuring-voice-controlled-npcs-1582502 , https://respawn.outlookindia.com/gaming/gaming-news/ubisoft-reportedly-using-far-cry-7-for-ai-tests , https://mezha.ua/news/ubisoft-testuye-generativniy-shi-u-far-cry-7-311577/
- AllThings.how, "Where Winds Meet's AI NPCs are already being pushed to their breaking point", updated 18 Nov 2025: https://allthings.how/where-winds-meets-ai-npcs-are-already-being-pushed-to-their-breaking-point/
- Notebookcheck on Where Winds Meet (via search summary); Wccftech, KitGuru, 3DJuegos coverage: https://www.notebookcheck.net/New-free-RPG-impresses-with-innovation-a-feature-I-ve-always-wished-games-had.1164695.0.html , https://wccftech.com/where-winds-meet-ai-chatbot-npcs-are-unironically-the-games-best-feature/amp/ , https://www.3djuegos.com/juegos/where-winds-meet/noticias/he-hablado-npc-chino-siglo-ix-mejor-tortilla-patatas-cebolla-como-casi-todo-este-rpg-respuesta-me-ha-decepcionado
- Rock Paper Shotgun, "Horrible, boring and cheap: experts pan new chatbot NPCs" (via search summary): https://www.rockpapershotgun.com/horrible-boring-and-cheap-experts-pan-new-chatbot-npcs-but-some-leave-room-for-optimism
- Steam, Suck Up!, read 2026-10-09: https://store.steampowered.com/app/2726370/Suck_Up/
- Steam, Vaudeville, read 2026-10-09: https://store.steampowered.com/app/2240920/Vaudeville/
- Game Developer, "Vaudeville, a pre-mortem": https://www.gamedeveloper.com/design/vaudeville-pre-mortem
- Notebookcheck on Whispers from the Star's Steam launch; 36Kr review; Steam page: https://www.notebookcheck.net/Steam-launch-New-interactive-fiction-game-debuts-to-Very-Positive-reviews-may-divide-gamers-over-heavy-AI-integration.1087821.0.html , https://www.36kr.com/p/3428919332048259 , https://store.steampowered.com/app/3730100/_/
- Park et al., "Generative Agents: Interactive Simulacra of Human Behavior", arXiv 2304.03442: https://arxiv.org/abs/2304.03442
- Valve live-generated AI content rules, via GamingOnLinux (Jan 2024) and PCGamesN: https://gamingonlinux.com/2024/01/valve-announces-new-rules-for-games-with-ai-content-on-steam , https://www.pcgamesn.com/steam/ai-content-reporting
- Decrypt, "Fortnite fixes AI-powered Darth Vader after it starts saying slurs", May 2025; Insider Gaming on the SAG-AFTRA charge: https://decrypt.co/320500/fortnite-ai-powered-darth-vader-saying-slurs , https://insider-gaming.com/fortnite-faces-unfair-labor-practice-charge-for-use-of-ai-darth-vader-voice/
- Frisson Labs, Charles Niu, "It's 2026... where are all the AI NPCs?", 21 May 2026 (author builds AI companions): https://www.frisson-labs.com/ai-npcs-2026
- Gravitee, "Securing LLM-powered NPC dialogue and tool calls" (vendor blog): https://www.gravitee.io/corpus/gen-1150/simulation-video-game/securing-llm-powered-npc-dialogue-and-tool-calls-in-simulation-games-with-gravitee-ai-gateway.html