The Rundown AI homepage

Independent tool overview

Realtime TTS-2 at a glance

Inworld Realtime TTS-2 is a low-latency speech model that uses prior conversation audio to adapt its delivery to a user's tone, pace, and emotional state.

Visit the official Realtime TTS-2 site ↗
Realtime TTS-2 product preview
Best for
Realtime voice agents and conversational apps
Latency
Sub-200ms median first audio chunk for TTS
Languages
100+ with crosslingual switching
Availability
Research preview via REST and Realtime APIs
Last reviewed
August 29, 2026

Overview

What Realtime TTS-2 is

Realtime TTS-2 is Inworld AI's expressive text-to-speech model for live voice agents and conversational applications. It produces a first audio chunk in under 200 milliseconds at the TTS layer and can run through REST streaming or Inworld's persistent Realtime API.

Its differentiator is conversational awareness: inside a Realtime session, the model can condition its next response on audio from prior turns rather than relying only on a transcript. Developers can also direct delivery with natural-language stage directions, add non-verbal cues, preserve a voice across more than 100 languages, and create voices from reference audio or a written description.

The model is still a research preview, and a good demo does not guarantee production behavior. Teams should test pronunciation, interruptions, noisy audio, emotional appropriateness, language quality, latency under load, and the consent and disclosure rules that apply to cloned or synthetic voices.

Use cases

Who Realtime TTS-2 is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Customer-facing voice agents

Create support, service, tutoring, or companion experiences that adjust delivery to the conversation.

Expressive character voices

Direct tone and pacing with prose instructions and inline non-verbal cues.

Multilingual conversations

Keep one speaker identity while moving between supported languages, including within one utterance.

Custom voice prototyping

Clone a permitted voice from reference audio or design a new saved voice from a text description.

Capabilities

Core Realtime TTS-2 features

1

Conversational awareness

Uses audio from prior turns in a Realtime session so response delivery can reflect the user's tone, pace, and emotional state.

2

Natural-language voice direction

Accepts descriptive bracketed instructions for delivery rather than limiting developers to a fixed emotion list.

3

Crosslingual identity

Preserves the selected speaker's identity across more than 100 languages and supports mid-utterance switching.

4

Voice cloning and design

Supports instant cloning from clean reference audio and generates reusable voices from written descriptions.

5

Low-latency streaming

Streams audio through REST or a WebSocket-based Realtime session with sub-200ms median TTS first-chunk latency.

6

Integrated realtime stack

Can combine Inworld speech recognition, model routing, LLMs, and TTS on one persistent connection.

Process

How the Realtime TTS-2 workflow works

  1. Step 1

    Prototype the voice

    Select a stock voice, clone an authorized speaker, or design a voice, then test representative scripts and languages.

  2. Step 2

    Choose an API path

    Use REST streaming for directed synthesis or the Realtime API when prior conversation audio should influence delivery.

  3. Step 3

    Tune the experience

    Add voice directions, pronunciation controls, speaking-rate settings, and safe fallbacks for uncertain output.

  4. Step 4

    Load-test and govern

    Measure end-to-end latency and concurrency, monitor synthesis failures, document consent, and disclose synthetic voices where required.

Cost

Realtime TTS-2 pricing and free plan

Pricing is based on characters, with about 1,000 characters estimated as one minute of audio. Paid subscriptions provide monthly credits and progressively lower TTS-2 rates.

On-Demand

Start free; $25 per 1M characters

For evaluation and prototyping.

  • Up to 70 minutes of TTS included
  • 100 custom voices
  • Voice cloning, voice design, Realtime API, and commercial license

Creator

$25/month; $20 per 1M characters

For creators and small projects.

  • $25 in monthly credits
  • 500 custom voices
  • Workspace and team features

Builder

$100/month; $17.50 per 1M characters

For growing products and small teams.

  • $100 in monthly credits
  • 3,000 custom voices
  • Higher concurrency limits

Developer

$300/month; $15 per 1M characters

For production applications.

  • $300 in monthly credits
  • 10,000 custom voices
  • Priority email support
  • Professional cloning available as an add-on

Growth

$1,500/month; $12.50 per 1M characters

For larger deployments and compliance needs.

  • $1,500 in monthly credits
  • 30,000 custom voices
  • Higher API limits
  • ZDR, HIPAA, and BAA options as add-ons

Enterprise

Custom; as low as $5 per 1M characters

For custom volume, limits, support, or deployment.

  • Custom SLAs and terms
  • On-premises and data-residency options
  • Dedicated support

Pricing checked . Check current pricing at the source ↗

Assessment

Realtime TTS-2 strengths and limitations

Where it stands out

  • Conditions speech delivery on prior conversation audio, not only text
  • Combines low-latency streaming with detailed natural-language direction
  • Supports voice continuity across a large multilingual set
  • Offers stock voices, cloning, and prompt-based voice design
  • Provides a clear volume-discount path from prototype to enterprise

What to consider

  • Realtime TTS-2 is presented as a research preview, so behavior and interfaces can change
  • Language quality varies, with long-tail languages described as experimental during the launch window
  • Conversational audio awareness depends on the Realtime API rather than a simple isolated REST request
  • End-to-end voice-agent latency also depends on speech recognition, reasoning, tools, networking, and client playback
  • Voice cloning and emotional adaptation require careful consent, disclosure, abuse prevention, and human review

Compare

Realtime TTS-2 alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Miscellaneous

GPT-Realtime-2

Choose GPT Realtime 2 for an OpenAI-native speech-to-speech model that combines reasoning and audio interaction.

Explore GPT-Realtime-2

Miscellaneous

TADA by Hume AI

Choose TADA by Hume AI when open-source text-audio alignment and controllable speech are the priority.

Explore TADA by Hume AI

Sales

Vapi

Choose Vapi when you want a broader voice-agent platform that can orchestrate multiple speech and model providers.

Explore Vapi

Questions

Realtime TTS-2 FAQs

What is Realtime TTS-2?

It is Inworld AI's expressive, low-latency text-to-speech model for live conversation, available through streaming REST and Realtime APIs.

How fast is Realtime TTS-2?

Inworld reports sub-200ms median time to the first audio chunk for the TTS layer. Complete agent response time also depends on every other part of the stack.

How does conversational awareness work?

Within an Inworld Realtime session, prior user audio remains in context so the model can adapt response delivery to the preceding tone, pacing, and emotional cues.

How much does Realtime TTS-2 cost?

The On-Demand rate is $25 per million characters, and paid tiers reduce the rate as low as $12.50 before custom enterprise pricing. Plans include credits and different limits.

Does Realtime TTS-2 support voice cloning?

Yes. It supports cloning from a short, clean reference recording and also offers Advanced Voice Design for creating a saved voice from a text description.

Can it switch languages mid-sentence?

Yes. Inworld says the model can switch among supported languages within one generation while preserving the speaker identity, though quality should be tested language by language.

Bottom line

Our Realtime TTS-2 verdict

Realtime TTS-2 is a strong fit when delivery should respond to the person speaking, not merely read generated text. Its differentiators are meaningful for live agents, but teams should treat the preview label seriously and validate language quality, latency, safety, and consent on their own production traffic.

Visit Realtime TTS-2 website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.