Customer-facing voice agents
Create support, service, tutoring, or companion experiences that adjust delivery to the conversation.
Independent tool overview
Inworld Realtime TTS-2 is a low-latency speech model that uses prior conversation audio to adapt its delivery to a user's tone, pace, and emotional state.
Visit the official Realtime TTS-2 site ↗
Overview
Realtime TTS-2 is Inworld AI's expressive text-to-speech model for live voice agents and conversational applications. It produces a first audio chunk in under 200 milliseconds at the TTS layer and can run through REST streaming or Inworld's persistent Realtime API.
Its differentiator is conversational awareness: inside a Realtime session, the model can condition its next response on audio from prior turns rather than relying only on a transcript. Developers can also direct delivery with natural-language stage directions, add non-verbal cues, preserve a voice across more than 100 languages, and create voices from reference audio or a written description.
The model is still a research preview, and a good demo does not guarantee production behavior. Teams should test pronunciation, interruptions, noisy audio, emotional appropriateness, language quality, latency under load, and the consent and disclosure rules that apply to cloned or synthetic voices.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Create support, service, tutoring, or companion experiences that adjust delivery to the conversation.
Direct tone and pacing with prose instructions and inline non-verbal cues.
Keep one speaker identity while moving between supported languages, including within one utterance.
Clone a permitted voice from reference audio or design a new saved voice from a text description.
Capabilities
Uses audio from prior turns in a Realtime session so response delivery can reflect the user's tone, pace, and emotional state.
Accepts descriptive bracketed instructions for delivery rather than limiting developers to a fixed emotion list.
Preserves the selected speaker's identity across more than 100 languages and supports mid-utterance switching.
Supports instant cloning from clean reference audio and generates reusable voices from written descriptions.
Streams audio through REST or a WebSocket-based Realtime session with sub-200ms median TTS first-chunk latency.
Can combine Inworld speech recognition, model routing, LLMs, and TTS on one persistent connection.
Process
Step 1
Select a stock voice, clone an authorized speaker, or design a voice, then test representative scripts and languages.
Step 2
Use REST streaming for directed synthesis or the Realtime API when prior conversation audio should influence delivery.
Step 3
Add voice directions, pronunciation controls, speaking-rate settings, and safe fallbacks for uncertain output.
Step 4
Measure end-to-end latency and concurrency, monitor synthesis failures, document consent, and disclose synthetic voices where required.
Cost
Pricing is based on characters, with about 1,000 characters estimated as one minute of audio. Paid subscriptions provide monthly credits and progressively lower TTS-2 rates.
Start free; $25 per 1M characters
For evaluation and prototyping.
$25/month; $20 per 1M characters
For creators and small projects.
$100/month; $17.50 per 1M characters
For growing products and small teams.
$300/month; $15 per 1M characters
For production applications.
$1,500/month; $12.50 per 1M characters
For larger deployments and compliance needs.
Custom; as low as $5 per 1M characters
For custom volume, limits, support, or deployment.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Miscellaneous
Choose GPT Realtime 2 for an OpenAI-native speech-to-speech model that combines reasoning and audio interaction.
Explore GPT-Realtime-2 →Miscellaneous
Choose TADA by Hume AI when open-source text-audio alignment and controllable speech are the priority.
Explore TADA by Hume AI →Sales
Choose Vapi when you want a broader voice-agent platform that can orchestrate multiple speech and model providers.
Explore Vapi →Questions
It is Inworld AI's expressive, low-latency text-to-speech model for live conversation, available through streaming REST and Realtime APIs.
Inworld reports sub-200ms median time to the first audio chunk for the TTS layer. Complete agent response time also depends on every other part of the stack.
Within an Inworld Realtime session, prior user audio remains in context so the model can adapt response delivery to the preceding tone, pacing, and emotional cues.
The On-Demand rate is $25 per million characters, and paid tiers reduce the rate as low as $12.50 before custom enterprise pricing. Plans include credits and different limits.
Yes. It supports cloning from a short, clean reference recording and also offers Advanced Voice Design for creating a saved voice from a text description.
Yes. Inworld says the model can switch among supported languages within one generation while preserving the speaker identity, though quality should be tested language by language.
Bottom line
Realtime TTS-2 is a strong fit when delivery should respond to the person speaking, not merely read generated text. Its differentiators are meaningful for live agents, but teams should treat the preview label seriously and validate language quality, latency, safety, and consent on their own production traffic.
Visit Realtime TTS-2 website ↗
Custom Voices - xAI's new tool to clone your voice from short clips for use within Grok's applications

Autoscientist - Adaption's new tool for automating AI model training

Deep Max - Exa's new SOTA agentic search tool

Miso One - Open-source text-to-speech model that reads a speaker’s tone for expressive responses

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.