Inworld vs Convai: building AI NPCs in 2026

Published · Researched 2026-09-21

Research current as of September 20, 2026. Written from observed developer-community evidence — press hands-ons, forums, review aggregations, official docs — not from firsthand development work on either platform.

The one-sentence version: if you want the smartest-talking character you can demo, you want Inworld. If you want the character you can actually ship into a game this month, you want Convai. And if you're betting your studio's next three years on either, you'd better read the Soul Machines section before you sign anything.

What each platform actually is

Convai is the most complete full stack for putting a talking 3D character into a game engine. Unity plugin, Unreal Engine plugin (via FAB), web plugins, modding frameworks, Pixel Streaming Embed for browser-based characters — the delivery mechanism is "install the plugin, open the dashboard, have a character that talks." Its signature features are about embodiment: Convai Actions turn spoken commands into physical behavior (move to, follow, stop, then custom Blueprint-wired actions), and Dynamic Context gives characters three levels of live scene awareness. The pitch isn't a chatbot with a face. It's an NPC with a body.

Inworld is a character engine first and an infrastructure company second — a distinction that matters more every quarter. Founded by ex-Google DeepMind and Dialogflow engineers (CEO Kylan Gibbs and chairman Ilya Gelfenbeyn, who previously built and sold API.AI to Google), it started as the NPC-maker's brain: Character Brain (personality, backstory, emotions, motivations configured in natural language), long-term memory, relationship retention, and a Contextual Mesh of shared world lore every character draws on. In 2025–2026 it visibly repositioned toward real-time voice-AI infrastructure — TTS-2 and TTS-2 Flash, STT 1 with voice profiling, voice cloning and voice design, a Realtime API, and an LLM Router with 220+ models. Per the company's June 2026 press release, three of the ten largest token-consuming consumer apps run LLM workloads on Inworld infrastructure, one processing over 600 billion tokens a day. That's not an NPC company talking. That's a utility.

Both ship real products as of September 2026. Neither is a demo graveyard.

The showdown everyone keeps rerunning

The developer community has essentially been A/B-testing these two platforms in public for two years, and the results are remarkably consistent.

At NVIDIA's demo events — the ramen shop, the bellhop — Inworld kept winning on personality. Inverse's reviewer famously told Inworld's bellhop character Tae to poison the hotel guests; Tae "gladly" riffed along, sending "a crowd of Nvidia employees" into gasps. In the same coverage, Convai's characters were judged "dull," "predictable," and "uninspired." The Verge's Sean Hollister found Convai's Jin and Nova "effectively generative AI chatbots." Kotaku and Gizmodo panned the Convai demos as "dry, colorless, emotionless." Gizmodo's verdict on the whole exercise: "I won't remember Noodle Shop Jin. Though, he will remember me, the Dreaded Ramen Bounty Hunter."

But the GDC 2025 version of the story was worse for everyone. Journalists came away broadly negative — characters were fluent but hollow. A widely-cited developer research paper (zbbsdsb/macha on GitHub, updated September 2026) summarized the pattern: characters "forgot key details moments after they were stated. They responded to verbal abuse as though it were a compliment. They agreed with everything — 'yes and… my very stupid idea.'" Nobody's brain survives production scrutiny unscathed. Inworld's characters talk too much and can't be interrupted gracefully; Convai's characters agree with your worst ideas like an improv partner who never says no.

And Inworld wins on the thing players feel first: speed. NewAtlas's hands-on with Inworld Arcade: "There's virtually no delay involved in a verbal interaction, and that fact alone absolutely blows me away." Latency is Convai's most-repeated complaint — "Latency is big" shows up in the comments of every viral Convai reel. The counterweight: "Facial animations in this tech demo are terrible… characters tend to talk way too much, and they don't stop if you interrupt them." The brain is good. The face is not.

Pricing: what it actually costs

Both platforms use credit models that sound simple on the pricing page and get complicated the moment you try to predict a game's monthly bill.

Convai (prices from saasworthy.com, crawled ~September 14, 2026; the monthly/yearly orientation looks inverted and is UNVERIFIED against Convai's own pricing page):

  • Free: $0 — 100 interactions a month, 1 MB knowledge bank, 1 character concurrency, forum support.
  • Indie Dev: $22/month — 3,000 interactions/month, 5 MB knowledge, long-term memory, 10 hrs Convai Sim/Avatar Studio.
  • Professional: $69/month — 10,000 interactions/month, 20 MB knowledge, 3 concurrency, 40 hrs Pixel Streaming server.
  • Scale: $329/month — 50,000 interactions/month, 100 MB knowledge, 15 concurrency, conversation evaluation, vision via API.
  • Enterprise: custom — SLAs, data ownership, white-label, on-premises.

The part that scares developers is the credit math inside each turn. Convai's own docs (updated ~September 16, 2026) give an example: one typical spoken turn — 8 seconds of user speech, GPT-4o mini, an ElevenLabs voice, a 16-second reply — costs about 19 credits, including a 6-credit platform fee per turn. That's the line item Instagram commenters were reacting to with "and this is how games will become a subscription." Every spoken turn is four bills stapled together: STT + LLM tokens + TTS seconds + platform fee. Budget 3× your prototype math before you commit — the single most repeated piece of advice in this space.

Inworld (official inworld.ai/pricing, crawled ~September 3, 2026):

  • On-Demand: free — prototyping, up to 70 minutes of TTS, 100 custom voices, voice cloning and design, Realtime API, all 220+ LLM router models, commercial license. Rates: TTS-2 $25 per million characters, TTS-2 Flash $15 per million, STT 1 $0.15 per hour.
  • Creator: $25/month (includes $25 in credits, up to 33% off TTS/STT)
  • Builder: $100/month (includes $100 in credits, up to 40% off, 3,000 custom voices)
  • Developer: $300/month (includes $300 in credits, up to 47% off, 10,000 custom voices, priority email support)
  • Growth: $1,500/month (includes $1,500 in credits, up to 53% off)
  • Enterprise: custom — TTS as low as $5 per million characters, custom LLMs, SLAs, EU/India data residency.

Inworld's June 2026 pitch was explicit: "take down the biggest wall in consumer AI: cost." Dedicated GPUs from $5/GPU-hour versus roughly $11 hyperscaler on-demand; open models at up to 50% below public third-party rates. Independent aggregator techreviewer (July 2026) noted Inworld's TTS quality as "on par with top-tier alternatives at a fraction of the cost."

Honest comparison: Inworld's meter is more legible — characters in, minutes out, dollars attached. Convai's meter is a black box with four inputs and a platform fee per turn. But legible doesn't mean cheap at scale, and neither company will tell you your production bill before you've built the thing. Budget 3× your prototype math — the single most repeated piece of advice in this space.

The community signal, platform by platform

Convai has the mindshare. That's not close. Social searches surfaced 28 engagement-ranked posts, nearly all from Convai's own brand account @conv.ai (~1.9K followers) — brand-run content with real community reaction in the comments. On the June 2026 Convai Actions reel: "Hope gta's npc are like this." The Convai Developer Forum is active, GitHub docs were updated within four days of September 20, 2026, and tutorials publish regularly. Tommy Thompson of AI and Games — an actual AI-for-games specialist, not a hype account — reviewed Convai's production-ready NPC builder in 2026 and came away convinced, having been "rather cynical" about AI NPC companies failing to deliver practical production tools. (Convai republished the review, so treat the endorsement as marketing-adjacent; the quote is Thompson's.)

The complaints are the complaints of a platform people are actually using: latency, shallow characters, credit economics, and the cloud-dependency question that never goes away. "Does it work with tokens, or can it run in a package without internet?" asked one IG commenter. Per the research, the answer is no — it's cloud-dependent, and developers asking for offline mode are asking for a different product.

Inworld has the connoisseurs. Press and experienced devs consistently rate its Character Brain as more alive and its voice stack as better value. But the community footprint is thin — techreviewer's July 2026 aggregation found six reviews total (one G2, five Product Hunt): "The review base is thin and warrants caution before drawing firm conclusions." Most damning: "no reviewer commentary on reliability, uptime, or performance under production load — a critical gap." Support and documentation quality are "entirely unaddressed" in the independent record — UNVERIFIED beyond official claims.

Then the pivot anxiety. A German-language dev-community guide (eesel.ai) reports Inworld is "increasingly focusing on large enterprise customers and discontinuing personal accounts" — indie devs worry the NPC engine is being deprioritized for enterprise voice infrastructure. Single-source and contested (Inworld's official 2026 pages still show self-serve tiers), but the anxiety exists because the repositioning is real and visible. When a company whose founding story is NPCs starts marketing to "the biggest wall in consumer AI: cost," developers are right to ask whether the NPC engine is still a first-class product or a legacy funding the voice business.

Where NVIDIA ACE fits

ACE is the third option and the odd one out: it's not a character company, it's a chip company's developer toolkit. NVIDIA ACE (Avatar Cloud Engine) is a suite of real-time AI microservices — speech (Riva ASR/TTS), intelligence (NeMo/Nemotron LLMs, ACE Agent orchestration), animation (Audio2Face-3D, Audio2Emotion) — delivered as cloud NIMs, on-device RTX models, and Unreal Engine 5 plugins. The June 2026 Game Agent SDK is open source: a lightweight C/C++ agentic framework with Agent, Chat, and RAG APIs, plus a UE5 plugin suite covering ASR, a small on-device language model, and TTS with Blueprint integration.

The shipped evidence is real but narrow: KRAFTON's PUBG Ally (an in-game AI teammate, limited beta) and Creative Assembly's experimental Total War: PHARAOH AI advisor. The community is GitHub, the NVIDIA Developer Discord, dev forums, and dev press — working game and tech developers, zero consumer-companion crossover.

ACE's role: the build-it-yourself maximum-control option. You get the components and the orchestration, and you get the bill in the form of GPU hours and engineering time rather than per-turn credits. A third-party 2026 estimate puts self-hosted production around $4,500 per GPU per year — flagged as an estimate, not an official NVIDIA figure. The real cost of ACE is the team that can operate it. Convai and Inworld sell you the character; ACE sells you the engine room and expects you to bring the crew.

The honest positioning: if you're a small team that needs a talking character now, ACE is overkill. If you need on-device inference, custom pipelines, or characters that must work inside your existing engine architecture at scale, ACE is the only one of the three that doesn't route your players' voices through someone else's cloud as a design constraint.

The ghost at the table: Soul Machines

Here's why the pricing and pivot sections above matter more than they look. On February 5, 2026, Soul Machines — the photorealistic digital-human company, $135M raised from Horizons Ventures, Temasek, Salesforce Ventures, SoftBank Vision Fund 2 — entered receivership, with KPMG appointed to run a sale process. Headcount had fallen from 253 to 45. No buyer outcome has been announced in any observed source as of September 2026 — UNVERIFIED.

Soul Machines was the enterprise reference point for this entire category: the realism bar, the six-figure-services price bar. In 2026 the developer community's view of it is dominated by the receivership. The conversation about the company is now about it — as a cautionary tale and a pricing benchmark in competitor blogs — rather than with it.

The lesson isn't "all platforms die." It's more specific and more uncomfortable: the enterprise digital-human business model was the fragile one. Big funding, big services contracts, big headcount, slow sales cycles — and when the money ran out, even $135M didn't buy a soft landing. The platforms that survive this category will be the ones whose unit economics work for indie-scale developers, not just flagship brand programs. That's the subtext of Inworld's entire June 2026 price-cut campaign, and it's the anxiety behind every Convai credit-pricing thread. Build on a platform whose cheapest viable customer looks like you. If the platform's economics only make sense for customers ten times your size, you are not the customer — you're the portfolio padding, and portfolio padding gets cut. And remember your character's face and brain can come from different vendors: Ready Player Me — the free avatar layer many products built on — shut down January 31, 2026 after the Netflix acquisition. Architect accordingly.

Who should pick which

Pick Convai if: you're building in Unity or Unreal and you need a talking, moving, scene-aware character working this month. The plugins are the product. Accept that the characters will need real writing and guardrail work (Convai's own CEO says so: "We need more high quality writers and artists, not less"), that latency will be your first production fight, and that you'll want to model your credit burn at 3× your prototype numbers before you commit.

Pick Inworld if: character depth and voice quality are your differentiators and you have the engineering to evaluate production reliability yourself. The Character Brain is the best-regarded in the category, the voice stack is genuinely good value, and the LLM router gives you model choice the others don't. But confirm — in writing — that the NPC engine is still a supported first-class product. The pivot anxiety isn't proof of anything, but it exists for a reason.

Pick NVIDIA ACE if: you're a studio with engine programmers and you need on-device or hybrid inference, or you want the components without a character company's roadmap between you and your players. Budget for engineering time instead of credit bills, and accept that "fully customizable" means "you build the parts they didn't."

Pick neither, and self-host, if: your character's continuity matters more than its polish. Project AIRI — free, open-source, MIT-licensed, ~49,000 GitHub stars as observed in September 2026 — is the strongest "can't be taken away from me" option. It's more work and less magic. That's the trade.

Bottom line

Inworld builds the better brain. Convai builds the better toolbox. ACE builds the better engine room. None of them has solved the two problems that actually kill AI NPC projects: characters that are fluent but hollow, and bills that are predictable only in retrospect. The technology is genuinely impressive and genuinely unfinished: the faces are rough, the memories are leaky, and the pricing models are all still experiments with your game's margins.

Build the character your writers can make interesting, because neither platform will do that part for you. Budget triple. And after Soul Machines, verify the company's pulse before you marry its roadmap: a living Discord, current docs, recent releases. In 2026, the most important feature of an AI character platform isn't the character. It's whether the company will still be there when your players are.

Inworld-branded "AI-driven virtual characters" artwork (from WCCFTech's Inworld Q&A coverage).
Inworld-branded "AI-driven virtual characters" artwork (from WCCFTech's Inworld Q&A coverage). from WCCFTech's Inworld Q&A coverage
SUBSTITUTE: generic AI-NPC robot illustration from 2025 NPC-AI coverage (not Inworld product imagery; the docs-site candidate 404'd, swap
SUBSTITUTE: generic AI-NPC robot illustration from 2025 NPC-AI coverage (not Inworld product imagery; the docs-site candidate 404'd, swap noted). not Inworld product imagery; the docs-site candidate 404'd, swap noted
Genuine Inworld Studio character-authoring UI (branded sidebar: Characters / Interactions / Narrative / Scenes; avatar creator panel), from
Genuine Inworld Studio character-authoring UI (branded sidebar: Characters / Interactions / Narrative / Scenes; avatar creator panel), from a tutorial thumbnail. Credit unknown — owner-supplied research image
Convai Unreal Engine editor with a custom AI character (skeleton tree / morph targets visible), from Convai's official blog "Create AI
Convai Unreal Engine editor with a custom AI character (skeleton tree / morph targets visible), from Convai's official blog "Create AI Characters / Custom Avatars in Unreal Engine." Credit unknown — owner-supplied research image
Convai web playground demo with the "Aria" AI character (convai branding, chat panel), from Convai's official blog.
Convai web playground demo with the "Aria" AI character (convai branding, chat panel), from Convai's official blog. Credit unknown — owner-supplied research image
Community demo screenshot of Convai NPCs in Unreal, from the Convai Developer Forum.
Community demo screenshot of Convai NPCs in Unreal, from the Convai Developer Forum. Credit unknown — owner-supplied research image