Live multilingual conversations
Translate ongoing speech between supported languages without forcing each participant into a strict speak-then-wait rhythm.
Independent tool overview
Gemini 3.5 Live Translate is Google's low-latency speech-to-speech translation model for more than 70 languages. It streams translated audio a few seconds behind the speaker while attempting to preserve intonation, pacing, and pitch. Consumers can encounter it through Google Translate, while developers can build with the public-preview Gemini Live API at an effective paid rate of about $0.0368 per translated audio minute.
Visit the official Gemini 3.5 Live Translate site ↗
Overview
Unlike turn-based interpreters that wait for a speaker to stop, Gemini 3.5 Live Translate continuously processes incoming speech and starts producing translated speech while the speaker continues. The model balances latency against context, automatically detects supported source languages, and is designed to handle multilingual input and noisy real-world environments.
Google launched the model across three channels: a public preview in the Gemini Live API and Google AI Studio, a global rollout in the Google Translate mobile apps, and a private preview for selected Google Workspace customers in Google Meet. The developer version is an interpreter pipeline rather than a general-purpose voice agent: it accepts audio only, produces translated audio plus a transcript, and does not support tools, search grounding, structured output, or text input.
The model is useful for live calls, meetings, lessons, broadcasts, travel, and multilingual services, but it is not a certified human interpreter. Google documents failure cases involving accents, similar languages, rapid language switching, multiple speakers, voice consistency, and background audio. High-stakes medical, legal, safety, or financial conversations still need qualified human interpretation and a way to recover when the model mishears or mistranslates.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Translate ongoing speech between supported languages without forcing each participant into a strict speak-then-wait rhythm.
Build near-real-time interpretation into conferencing, customer support, marketplace, or collaboration experiences.
Stream a translation to listeners a few seconds behind a guide, instructor, speaker, or program.
Use the Google Translate app with headphones, or the Android listening mode where available, for personal translated audio.
Capabilities
The model translates while speech is still arriving, keeping output only a few seconds behind instead of waiting for a complete turn.
Google designed the translated voice to retain aspects of the speaker's intonation, pacing, and pitch rather than outputting a generic flat voice.
Supported source languages can be detected without requiring users to manually switch the source-language configuration each time.
More than 70 languages enable over 2,000 language combinations rather than restricting translation to pairs that include English.
The model is designed to filter noise and music and keep speech usable in less controlled environments, although background audio can still cause artifacts.
Developers send 16-bit mono PCM audio at 16kHz in roughly 100ms chunks over the Live API and receive 24kHz mono PCM translated audio plus text transcripts.
Google embeds an imperceptible SynthID watermark in model-generated audio to help identify AI-produced speech.
Process
Step 1
Use Google Translate for personal listening, evaluate Google Meet under its Workspace rollout, or use the Live API for a product integration.
Step 2
Configure the BCP-47 target-language code and decide how the system should handle input already spoken in the target language.
Step 3
Capture mono 16kHz PCM, send small regular chunks, minimize overlapping speakers, and provide headphones or echo control where possible.
Step 4
Show transcripts, make the active languages visible, let participants pause or repeat, and provide a human fallback for consequential communication.
Step 5
Evaluate representative accents, code-switching, background noise, names, technical vocabulary, rapid exchanges, and long sessions before launch.
Cost
The Gemini Developer API offers a free tier and paid Standard usage. Paid input audio costs $3.50 per million tokens, estimated at $0.0053 per minute, while translated output audio costs $21 per million tokens, estimated at $0.0315 per minute. With both streams, Google estimates about $0.0368 per minute. Free-tier data may be used to improve Google's products; paid-tier data is not. Google Translate consumer access has no separate per-minute API charge.
Free consumer access
Personal Live Translate access through the Android and iOS app as the rollout reaches the user and language pair.
$0
For prototyping the preview model in Google AI Studio or through the Gemini Live API within free-tier limits.
About $0.0368/minute
Usage-based Standard pricing for production-oriented developer evaluation and integrations.
Private preview
Selected Google Workspace business customers can evaluate the expanded 70-language speech-translation experience.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Consumer
A general consumer translation tool for text, images, and voice when live streaming interpretation is not the only need.
Explore Translate with ChatGPT →Consumer
A Google open-model option for developers who need controllable text translation and self-managed deployment rather than a hosted live voice pipeline.
Explore TranslateGemma →Marketing
A better fit for produced multilingual avatar video and presentation workflows rather than low-latency live conversation translation.
Explore LiveAvatar by HeyGen →Questions
It is Google's streaming speech-to-speech translation model. It automatically detects supported spoken languages and produces translated speech a few seconds behind while trying to preserve intonation, pacing, and pitch.
Google documents more than 70 supported languages, enabling over 2,000 possible source-and-target combinations rather than requiring English on one side.
The paid Standard tier costs $3.50 per million input audio tokens and $21 per million output audio tokens. At Google's estimate of 25 audio tokens per second, the combined effective cost is about $0.0368 per minute.
Yes, within Google's free limits. The pricing table says free-tier interactions may be used to improve Google's products, while paid-tier interactions are not. The model itself is still a public preview.
Yes. Google began a global rollout in the Google Translate app for Android and iOS. Headphones provide the translated audio experience, and an earpiece listening mode is also rolling out on Android.
Google launched the expanded Meet experience in private preview for selected business Workspace customers in June 2026 and announced a broader rollout later in the year. Confirm current availability for the specific Workspace account.
It tries to preserve vocal qualities such as intonation, pacing, and pitch, but Google documents inconsistent voice replication, especially after long pauses or during rapid multi-speaker conversations. It should not be treated as a guaranteed identity-preserving voice clone.
Yes. Google says all model-generated audio is marked with an imperceptible SynthID watermark to support detection of AI-generated content.
Bottom line
Gemini 3.5 Live Translate is one of the most practical real-time translation models for broad language coverage, natural delivery, and developer access. The low estimated API cost makes experimentation unusually accessible. The preview label and documented voice, detection, and background-audio failures matter, though: deploy it with visible transcripts, correction paths, and human escalation rather than presenting it as infallible interpretation.
Visit Gemini 3.5 Live Translate website ↗
Miso One - Open-source text-to-speech model that reads a speaker’s tone for expressive responses

DiffusionGemma - Google's open diffusion model that can quadruple text generation speed

Autoscientist - Adaption's new tool for automating AI model training

Sonic-3.5 & Ink-2 - Cartesia's new top-ranked speech and transcription models for voice agents

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.