The Rundown AI homepage

Independent tool overview

Gemini 3.5 Live Translate at a glance

Gemini 3.5 Live Translate is Google's low-latency speech-to-speech translation model for more than 70 languages. It streams translated audio a few seconds behind the speaker while attempting to preserve intonation, pacing, and pitch. Consumers can encounter it through Google Translate, while developers can build with the public-preview Gemini Live API at an effective paid rate of about $0.0368 per translated audio minute.

Visit the official Gemini 3.5 Live Translate site ↗
Gemini 3.5 Live Translate product preview
Model type
Low-latency streaming speech-to-speech translation
Languages
More than 70 supported source and target languages
Developer status
Public preview via Gemini Live API
Consumer access
Rolling out globally in Google Translate for Android and iOS
Paid API price
Approximately $0.0368 per minute for input plus translated output
Output disclosure
Generated audio carries an imperceptible SynthID watermark
Last reviewed
August 28, 2026

Overview

What Gemini 3.5 Live Translate is

Unlike turn-based interpreters that wait for a speaker to stop, Gemini 3.5 Live Translate continuously processes incoming speech and starts producing translated speech while the speaker continues. The model balances latency against context, automatically detects supported source languages, and is designed to handle multilingual input and noisy real-world environments.

Google launched the model across three channels: a public preview in the Gemini Live API and Google AI Studio, a global rollout in the Google Translate mobile apps, and a private preview for selected Google Workspace customers in Google Meet. The developer version is an interpreter pipeline rather than a general-purpose voice agent: it accepts audio only, produces translated audio plus a transcript, and does not support tools, search grounding, structured output, or text input.

The model is useful for live calls, meetings, lessons, broadcasts, travel, and multilingual services, but it is not a certified human interpreter. Google documents failure cases involving accents, similar languages, rapid language switching, multiple speakers, voice consistency, and background audio. High-stakes medical, legal, safety, or financial conversations still need qualified human interpretation and a way to recover when the model mishears or mistranslates.

Use cases

Who Gemini 3.5 Live Translate is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Live multilingual conversations

Translate ongoing speech between supported languages without forcing each participant into a strict speak-then-wait rhythm.

Translation-enabled calls and meetings

Build near-real-time interpretation into conferencing, customer support, marketplace, or collaboration experiences.

Tours, lessons, and broadcasts

Stream a translation to listeners a few seconds behind a guide, instructor, speaker, or program.

Mobile travel use

Use the Google Translate app with headphones, or the Android listening mode where available, for personal translated audio.

Capabilities

Core Gemini 3.5 Live Translate features

1

Continuous translation

The model translates while speech is still arriving, keeping output only a few seconds behind instead of waiting for a complete turn.

2

Voice-character preservation

Google designed the translated voice to retain aspects of the speaker's intonation, pacing, and pitch rather than outputting a generic flat voice.

3

Automatic language detection

Supported source languages can be detected without requiring users to manually switch the source-language configuration each time.

4

Broad language pairing

More than 70 languages enable over 2,000 language combinations rather than restricting translation to pairs that include English.

5

Noise robustness

The model is designed to filter noise and music and keep speech usable in less controlled environments, although background audio can still cause artifacts.

6

Streaming API

Developers send 16-bit mono PCM audio at 16kHz in roughly 100ms chunks over the Live API and receive 24kHz mono PCM translated audio plus text transcripts.

7

SynthID watermarking

Google embeds an imperceptible SynthID watermark in model-generated audio to help identify AI-produced speech.

Process

How the Gemini 3.5 Live Translate workflow works

  1. Step 1

    1. Pick the consumer or developer route

    Use Google Translate for personal listening, evaluate Google Meet under its Workspace rollout, or use the Live API for a product integration.

  2. Step 2

    2. Set the target language

    Configure the BCP-47 target-language code and decide how the system should handle input already spoken in the target language.

  3. Step 3

    3. Stream clean audio

    Capture mono 16kHz PCM, send small regular chunks, minimize overlapping speakers, and provide headphones or echo control where possible.

  4. Step 4

    4. Design for correction

    Show transcripts, make the active languages visible, let participants pause or repeat, and provide a human fallback for consequential communication.

  5. Step 5

    5. Test real conditions

    Evaluate representative accents, code-switching, background noise, names, technical vocabulary, rapid exchanges, and long sessions before launch.

Cost

Gemini 3.5 Live Translate pricing and free plan

The Gemini Developer API offers a free tier and paid Standard usage. Paid input audio costs $3.50 per million tokens, estimated at $0.0053 per minute, while translated output audio costs $21 per million tokens, estimated at $0.0315 per minute. With both streams, Google estimates about $0.0368 per minute. Free-tier data may be used to improve Google's products; paid-tier data is not. Google Translate consumer access has no separate per-minute API charge.

Google Translate app

Free consumer access

Personal Live Translate access through the Android and iOS app as the rollout reaches the user and language pair.

  • More than 70 languages
  • Headphone-based translated audio
  • Android listening mode rolling out
  • Not an embeddable commercial API plan

Gemini API Free tier

$0

For prototyping the preview model in Google AI Studio or through the Gemini Live API within free-tier limits.

  • Audio input and output free within limits
  • Preview availability and quotas apply
  • Google marks free-tier data as used to improve its products

Gemini API paid tier

About $0.0368/minute

Usage-based Standard pricing for production-oriented developer evaluation and integrations.

  • $3.50/M input audio tokens
  • $21/M output audio tokens
  • Estimated $0.0053/min input plus $0.0315/min output
  • Paid-tier data is not used to improve Google's products

Google Meet

Private preview

Selected Google Workspace business customers can evaluate the expanded 70-language speech-translation experience.

  • Broader rollout announced for later in 2026
  • Workspace plan and preview terms apply
  • Confirm availability with Google

Pricing checked . Check current pricing at the source ↗

Assessment

Gemini 3.5 Live Translate strengths and limitations

Where it stands out

  • Continuous speech translation produces a more natural conversational rhythm than strict turn-by-turn systems.
  • Supports more than 70 languages and over 2,000 possible language combinations.
  • Attempts to preserve the speaker's vocal delivery rather than replacing it with a generic voice.
  • Available both as a consumer experience and a developer API.
  • Official documentation publishes audio formats, model limits, pricing, supported languages, and known failure modes.
  • SynthID watermarking adds a disclosure mechanism to generated audio.

What to consider

  • The developer model is still a preview and has no announced shutdown date, so version stability and support terms can change.
  • It accepts audio only for translation; text prompts, search grounding, tools, function calling, and structured output are not supported.
  • Voice replication can drift after long pauses, misassign a voice or gender, or stick to one voice during rapid multi-speaker exchanges.
  • Heavy accents, closely related languages, and rapid language switching can confuse source-language detection or transcripts.
  • Background noise and music are filtered imperfectly and can create artifacts, especially when echoing audio already in the target language.
  • Near-real-time still means a delay of a few seconds, which affects interruptions and very fast conversations.
  • Free-tier API interactions may be used to improve Google's products; sensitive product testing should use the appropriate paid account and review Google's data terms.
  • Translation errors remain unacceptable in some medical, legal, emergency, immigration, and financial settings without qualified human review.

Compare

Gemini 3.5 Live Translate alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Consumer

Translate with ChatGPT

A general consumer translation tool for text, images, and voice when live streaming interpretation is not the only need.

Explore Translate with ChatGPT

Consumer

TranslateGemma

A Google open-model option for developers who need controllable text translation and self-managed deployment rather than a hosted live voice pipeline.

Explore TranslateGemma

Marketing

LiveAvatar by HeyGen

A better fit for produced multilingual avatar video and presentation workflows rather than low-latency live conversation translation.

Explore LiveAvatar by HeyGen

Questions

Gemini 3.5 Live Translate FAQs

What is Gemini 3.5 Live Translate?

It is Google's streaming speech-to-speech translation model. It automatically detects supported spoken languages and produces translated speech a few seconds behind while trying to preserve intonation, pacing, and pitch.

How many languages does Gemini 3.5 Live Translate support?

Google documents more than 70 supported languages, enabling over 2,000 possible source-and-target combinations rather than requiring English on one side.

How much does the Gemini Live Translate API cost?

The paid Standard tier costs $3.50 per million input audio tokens and $21 per million output audio tokens. At Google's estimate of 25 audio tokens per second, the combined effective cost is about $0.0368 per minute.

Is there a free Gemini Live Translate API tier?

Yes, within Google's free limits. The pricing table says free-tier interactions may be used to improve Google's products, while paid-tier interactions are not. The model itself is still a public preview.

Can I use Gemini 3.5 Live Translate in Google Translate?

Yes. Google began a global rollout in the Google Translate app for Android and iOS. Headphones provide the translated audio experience, and an earpiece listening mode is also rolling out on Android.

Is Gemini 3.5 Live Translate available in Google Meet?

Google launched the expanded Meet experience in private preview for selected business Workspace customers in June 2026 and announced a broader rollout later in the year. Confirm current availability for the specific Workspace account.

Does Gemini Live Translate clone the speaker's voice?

It tries to preserve vocal qualities such as intonation, pacing, and pitch, but Google documents inconsistent voice replication, especially after long pauses or during rapid multi-speaker conversations. It should not be treated as a guaranteed identity-preserving voice clone.

Is the translated audio watermarked?

Yes. Google says all model-generated audio is marked with an imperceptible SynthID watermark to support detection of AI-generated content.

Bottom line

Our Gemini 3.5 Live Translate verdict

Gemini 3.5 Live Translate is one of the most practical real-time translation models for broad language coverage, natural delivery, and developer access. The low estimated API cost makes experimentation unusually accessible. The preview label and documented voice, detection, and background-audio failures matter, though: deploy it with visible transcripts, correction paths, and human escalation rather than presenting it as infallible interpretation.

Visit Gemini 3.5 Live Translate website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.