The Rundown AI homepage

Independent tool overview

Gemini 3.1 Flash TTS at a glance

Gemini 3.1 Flash TTS is Google's preview text-to-speech model for low-latency, controllable single- and multi-speaker audio in more than 70 languages.

Visit the official Gemini 3.1 Flash TTS site ↗
Gemini 3.1 Flash TTS product preview
Model ID
gemini-3.1-flash-tts-preview
Status
Preview
Input and output
Text in, audio out
Languages
70+
Speaker modes
Single and multi-speaker
Last reviewed
August 30, 2026

Overview

What Gemini 3.1 Flash TTS is

Gemini 3.1 Flash TTS is a developer-focused speech model that turns a written script into natural, directed audio. It supports more than 70 languages, single- and multi-speaker output, natural-language voice direction and inline audio tags that can change delivery within a sentence.

This is a preview API model, not a full voiceover editor or a real-time conversational agent. It is strongest when a team needs scripted narration with precise control over style, accent, pace and tone. Developers can test it in Google AI Studio, call it through the Gemini API, stream audio from the 3.1 model or submit high-volume jobs through the Batch API.

Use cases

Who Gemini 3.1 Flash TTS is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Directed narration and voiceover

Create scripted audio where tone, pacing, accent and performance need more control than a basic text reader provides.

Multilingual content

Generate localized speech across more than 70 languages while keeping the same API-centered production workflow.

Two-speaker scripts

Produce podcast-style exchanges, training scenarios and character dialogue with distinct speaker assignments.

Low-latency application speech

Stream generated audio into an application when users should hear playback before the entire response finishes.

High-volume batch generation

Render larger libraries of prerecorded narration at the lower Batch API rate when immediate delivery is unnecessary.

Capabilities

Core Gemini 3.1 Flash TTS features

1

Inline audio tags

Place modifiers such as [whispers], [laughs], [sighs] or [shouting] inside a transcript to change a passage's delivery or add supported non-verbal sounds.

2

Natural-language direction

Describe the desired style, accent, pace, tone and emotional energy in ordinary language instead of relying only on fixed sliders.

3

Audio profiles and scene context

Define a speaker archetype, the surrounding scene and director's notes so the model has coherent performance context.

4

Single-speaker generation

Select a prebuilt voice and generate one continuous performance from a written transcript.

5

Native multi-speaker audio

Assign voices to multiple named speakers for dialogue, interviews and other scripted exchanges.

6

70+ language support

Generate speech across more than 70 languages, with Google's guidance recommending English-language audio tags even when the transcript is not English.

7

Streaming output

Gemini 3.1 Flash TTS can stream audio chunks as they are generated, a capability not supported by the earlier Gemini TTS versions.

8

Batch API support

Submit non-urgent generation jobs at half the standard input and output token rates.

9

Google AI Studio testing

Experiment with voices and performance controls in a browser, then export matching Gemini API parameters for implementation.

10

SynthID watermarking

Google embeds an imperceptible SynthID watermark in all audio produced by Gemini 3.1 Flash TTS.

Process

How the Gemini 3.1 Flash TTS workflow works

  1. Step 1

    Write and segment the script

    Prepare the exact words to be spoken and split long narration into sections of a few minutes to reduce voice and quality drift.

  2. Step 2

    Choose the speaker setup

    Select a supported prebuilt voice for one speaker or map separate voices to each character in a dialogue.

  3. Step 3

    Direct the performance

    Add an audio profile, scene context and director's notes covering style, accent, pacing and emotional tone.

  4. Step 4

    Insert local audio tags

    Use inline tags only where delivery should change, and keep the directions consistent with the script and selected voice.

  5. Step 5

    Test and review

    Generate samples in AI Studio or the API, then check pronunciation, speaker consistency, pacing, names, numbers and required disclosures.

  6. Step 6

    Stream or batch the final audio

    Use streaming for responsive playback, standard requests for immediate files or Batch API jobs for lower-cost offline production.

Cost

Gemini 3.1 Flash TTS pricing and free plan

Gemini Developer API pricing is token-based. Google counts generated audio at 25 tokens per second, making the paid Standard output rate approximately $0.03 per minute and Batch approximately $0.015 per minute, before the separate text-input charge. Free-tier requests are subject to rate limits and may be used to improve Google products; Google says paid-tier data is not.

Free Tier

Free

Limited Gemini Developer API access for testing.

  • Text input and audio output are free
  • Subject to free-tier rate limits
  • Google says free-tier data may be used to improve its products

Standard Paid API

$1 input / $20 audio output per 1M tokens

On-demand production pricing for the Gemini Developer API.

  • Text input: $1.00 per 1 million tokens
  • Audio output: $20.00 per 1 million tokens
  • About $0.03 per generated minute for audio output at 25 audio tokens per second
  • Google says paid-tier data is not used to improve its products

Batch API

$0.50 input / $10 audio output per 1M tokens

Lower-cost asynchronous generation for work that does not need an immediate response.

  • No free Batch tier
  • Text input: $0.50 per 1 million tokens
  • Audio output: $10.00 per 1 million tokens
  • About $0.015 per generated minute for audio output at 25 audio tokens per second

Pricing checked . Check current pricing at the source ↗

Assessment

Gemini 3.1 Flash TTS strengths and limitations

Where it stands out

  • Fine-grained direction through natural-language prompts and inline tags
  • More than 70 supported languages
  • Native single- and multi-speaker generation
  • Streaming support for responsive playback
  • Low effective audio-output price at the published token rate
  • Batch pricing for large offline workloads
  • Browser-based prototyping in Google AI Studio
  • SynthID watermarking on generated audio

What to consider

  • The model remains in Preview, so behavior, limits and the API surface can change
  • It accepts text only and produces audio only; it does not listen to or understand an audio input
  • It is designed for scripted speech, not open-ended real-time conversation
  • A selected voice may not always follow the requested speaker profile consistently
  • Speech quality and voice consistency can drift on outputs longer than a few minutes
  • Some requests can fail unpredictably, so production integrations need retry handling
  • Vague prompts may be rejected or cause performance directions to be spoken aloud
  • Audio tags are flexible rather than fully deterministic and do not have an exhaustive supported list
  • Pronunciation of names, numbers and specialized terms still requires human review
  • The API does not provide the complete editing, timeline and asset-management workflow of a dedicated voiceover studio
  • Teams should obtain permission before imitating identifiable people and disclose synthetic speech where context calls for it
  • Medical, legal, financial and other high-stakes narration should be checked against the approved source text before publication

Compare

Gemini 3.1 Flash TTS alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Marketing

ElevenLabs

Choose ElevenLabs for a mature end-to-end voice platform with broader creator tooling, voice management and deployment options.

Explore ElevenLabs

Content Creator

Inworld Realtime TTS

Consider Inworld Realtime TTS when interactive character and game-oriented low-latency speech is the primary use case.

Explore Inworld Realtime TTS

Content Creator

Qwen3-TTS

Use Qwen3-TTS when an open-source model and greater control over self-hosted infrastructure matter more than a managed API.

Explore Qwen3-TTS

Business Operations

Voxtral TTS

Compare Voxtral TTS for another developer API focused on multilingual speech and voice-agent applications.

Explore Voxtral TTS

Questions

Gemini 3.1 Flash TTS FAQs

What is Gemini 3.1 Flash TTS?

It is Google's preview text-to-speech model for turning exact written scripts into controllable audio. The model supports natural-language performance direction, inline audio tags, more than 70 languages and single- or multi-speaker generation.

Is Gemini 3.1 Flash TTS free?

Google offers a rate-limited free tier for Gemini Developer API testing. Paid Standard requests cost $1 per million text-input tokens and $20 per million audio-output tokens. Batch costs half those rates.

How much does Gemini 3.1 Flash TTS cost per minute?

Google counts audio at 25 tokens per second. At the current paid output rate, that works out to about $0.03 per generated minute for Standard requests or $0.015 per minute through Batch, plus the much smaller text-input charge.

Can Gemini 3.1 Flash TTS stream audio?

Yes. Version 3.1 supports streaming generated audio chunks by setting the API request to stream. Earlier Gemini TTS models do not support TTS streaming.

Does it support multiple speakers?

Yes. The Gemini TTS API can generate both single-speaker and multi-speaker audio, with a selected voice assigned to each named speaker.

What are Gemini audio tags?

Audio tags are inline directions such as [whispers], [laughs] or [sighs]. They can change the delivery of a line or add supported non-verbal sounds. Google says there is no exhaustive list, so results should be tested.

Is this the same as Gemini 3.1 Flash Live?

No. Flash TTS is for controlled recitation of a supplied script. Gemini 3.1 Flash Live is a separate audio-to-audio model built for interactive, real-time dialogue.

Is Gemini 3.1 Flash TTS ready for production?

It can be integrated into applications, but it is still labeled Preview. Production teams should pin the intended model ID, split long scripts, review every public-facing result and implement retries and monitoring.

Bottom line

Our Gemini 3.1 Flash TTS verdict

Gemini 3.1 Flash TTS is a compelling API choice for teams that want expressive, multilingual scripted speech without paying dedicated voice-platform prices. Inline tags, multi-speaker output, streaming and unusually low published rates make it especially attractive for application developers. Its Preview status and documented consistency issues mean it still needs careful prompt design, short generation segments, automated retries and human quality control.

Visit Gemini 3.1 Flash TTS website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.