The Rundown AI homepage

Independent tool overview

Gemini 3.5 Transcribe at a glance

Gemini 3.5 Transcribe is Google's public-preview speech-to-text model for low-latency live captions and prerecorded audio, with language auto-detection, custom vocabulary, smart formatting, diarization, and word timestamps.

Visit the official Gemini 3.5 Transcribe site ↗
Gemini 3.5 Transcribe product preview
Release status
Public preview
File model
gemini-3.5-transcribe
Live model
gemini-3.5-transcribe-live
Languages
85+ with code-switching
Estimated file price
About $0.005 per audio minute
Estimated live price
About $0.009 per audio minute

Overview

What Gemini 3.5 Transcribe is

Gemini 3.5 Transcribe is a dedicated speech-to-text model rather than a general Gemini model prompted to analyze audio. Google provides separate endpoints for live WebSocket streaming and prerecorded file processing, allowing developers to choose low latency or richer transcript metadata.

The model automatically detects more than 85 languages and can follow code-switching within a session. Custom vocabulary biasing helps with names, acronyms, product terms, and domain jargon. Prerecorded processing can add speaker labels and word-level timestamps, while the live model emphasizes sub-second interactive transcription.

Its distinctive feature is Smart transcription. Instead of preserving every spoken sound, it can remove filler words, repetitions, and disfluencies, resolve self-corrections, add punctuation, and normalize spoken values such as money or identification numbers into formatted text.

That cleanup is useful for dictation but inappropriate for every record. Legal, journalistic, research, support, compliance, and evidentiary workflows may require a verbatim transcript that preserves hesitation, false starts, or exact wording. Teams should store the original audio, expose whether cleanup is enabled, and review critical names, numbers, speakers, and quotations.

Use cases

Who Gemini 3.5 Transcribe is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Voice interfaces and live captions

Stream microphone or call audio into an interactive application with sub-second transcription latency.

Meetings and call analytics

Process recorded audio with speaker labels, timestamps, custom terminology, and clean formatting.

Polished dictation

Turn natural speech, self-corrections, and filler-heavy ideas into readable text for drafts, messages, and commands.

Multilingual products

Support language auto-detection, regional accents, and conversations that switch languages.

Domain-specific transcription

Bias recognition toward product names, acronyms, people, locations, and specialized terms with a custom vocabulary.

Capabilities

Core Gemini 3.5 Transcribe features

1

Live streaming transcription

The Live API uses a bidirectional WebSocket connection to return continuous low-latency text for interactive voice applications.

2

Prerecorded audio processing

The Interactions API transcribes uploaded meetings, calls, interviews, and other audio files with richer metadata.

3

Smart transcription

Optionally removes filler words, repetitions, and disfluencies, resolves self-corrections, and applies readable formatting.

4

Language auto-detection

Detects more than 85 languages and supports language switching within and between utterances.

5

Custom vocabulary

Accepts up to 1,000 terms to bias recognition toward domain vocabulary, with Google recommending a smaller focused list for best results.

6

Speaker diarization

Prerecorded processing can attribute segments to as many as eight speakers, although attribution for three or more remains experimental.

7

Word-level timestamps

Prerecorded transcripts can include start and end offsets for individual words for search, playback, clips, subtitles, and review.

8

Formatting and normalization

Adds capitalization and punctuation and can convert spoken numbers or money into concise formatted values.

9

Google product integrations

The same model powers voice experiences in the Gemini app on macOS, Rambler on supported Android devices, Google Antigravity, and other announced Google surfaces.

Process

How the Gemini 3.5 Transcribe workflow works

  1. Step 1

    Choose verbatim or cleaned output

    Decide whether the transcript must preserve fillers, repetitions, self-corrections, and exact spoken wording before enabling Smart transcription.

  2. Step 2

    Select live or file processing

    Use the Live model for interactive latency; use the file model when diarization, word timestamps, or longer recordings matter.

  3. Step 3

    Prepare audio and vocabulary

    Capture clear consented audio, normalize levels, reduce avoidable noise, and provide only the highest-value names, acronyms, and technical terms.

  4. Step 4

    Configure metadata features

    Request speaker labels and timestamps only when needed because they shorten the file limit and timestamps can reduce recognition accuracy.

  5. Step 5

    Validate consequential details

    Check names, identifiers, addresses, dates, currencies, negations, medical or legal terms, speaker attribution, and direct quotations against the audio.

  6. Step 6

    Store provenance and corrections

    Keep the original audio, model name, cleanup setting, vocabulary, transcript version, human edits, and word offsets so downstream users can audit the record.

Cost

Gemini 3.5 Transcribe pricing and free plan

Gemini 3.5 Transcribe has a free developer tier and token-based paid pricing. Google estimates an effective blended cost of about $0.005 per audio minute for prerecorded transcription and $0.009 per minute for Live transcription. Actual invoices are based on audio-input and text-output tokens, not the estimate.

Free developer tier

Free within limits

Evaluation access through Google AI Studio and unpaid Gemini API quota.

  • Live and prerecorded input and output are listed as free within quota
  • Free-tier content can be used to improve Google products
  • Human reviewers may read, annotate, and process inputs and outputs
  • Do not submit sensitive, confidential, or personal information

Prerecorded Transcribe

About $0.005/audio minute

Paid Gemini API processing for uploaded audio.

  • Audio input: $2 per 1 million tokens, estimated at $0.003 per minute
  • Text output: $12 per 1 million tokens, estimated at $0.002 per minute
  • Estimated blended rate assumes 25 audio tokens per second and 175 text tokens per minute
  • Speaker diarization and word timestamps are available
  • Batch, Flex, and Priority inference are not supported

Live Transcribe

About $0.009/audio minute

Paid real-time streaming over the Gemini Live API.

  • Audio input: $3.50 per 1 million tokens, estimated at $0.005 per minute
  • Text output: $21 per 1 million tokens, estimated at $0.004 per minute
  • Only final committed tokens are billed
  • Live streaming does not support speaker diarization or word-level timestamps

Enterprise delivery

Usage-based or contracted

Access through Gemini Enterprise Agent Platform and related Google Cloud products.

  • Available in public preview through Gemini Enterprise Agent Platform
  • Google Cloud pricing and contractual terms can differ from the developer API
  • Confirm region, SLA, support, retention, logging, and compliance requirements before production

Pricing checked . Check current pricing at the source ↗

Assessment

Gemini 3.5 Transcribe strengths and limitations

Where it stands out

  • Purpose-built live and file models avoid forcing one endpoint to serve incompatible latency and metadata needs.
  • Smart transcription can turn messy spoken dictation into polished, usable text without a separate cleanup pass.
  • More than 85 languages, code-switching, accents, and custom vocabulary support a broad set of products.
  • Prerecorded processing combines diarization, word timestamps, and vocabulary biasing in one low-cost API.
  • Estimated per-minute prices are highly competitive for large transcription workloads.
  • Paid Gemini API usage is not used to improve Google's products under the current developer terms.

What to consider

  • The model is in public preview, so behavior, limits, availability, naming, and pricing can change before general availability.
  • Smart transcription intentionally changes the spoken record by removing fillers, repetitions, and corrections; it is not a verbatim mode for evidentiary use.
  • Live sessions are limited to 10 minutes and do not support speaker diarization or word-level timestamps.
  • Prerecorded requests support up to one hour, but enabling diarization or word timestamps reduces the limit to 30 minutes.
  • Speaker attribution for three or more speakers is experimental, even though the file model supports up to eight labels.
  • Word-level timestamps can reduce transcription accuracy and should not be enabled by default when they are not needed.
  • Custom vocabulary accepts up to 1,000 terms, but Google says typical results are best with about 100 focused terms.
  • The published per-minute figures are estimates; actual cost depends on audio and output tokenization and transcript verbosity.
  • Free-tier inputs and outputs can be used for product improvement and reviewed by humans, making unpaid access unsuitable for confidential recordings.
  • Paid API data is not used for product improvement, but Google still logs prompts and responses for a limited period for abuse prevention and legal or regulatory obligations.
  • Transcripts can still misrecognize names, numbers, accents, overlapping speech, negations, quiet speakers, and noisy audio, while polished formatting can make errors look more authoritative.
  • Recording, diarization, storage, and downstream analysis require appropriate consent, retention, access, deletion, and jurisdiction-specific privacy controls.

Compare

Gemini 3.5 Transcribe alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Content Creator

Scribe v2

Choose ElevenLabs Scribe v2 when transcription accuracy, language coverage, diarization, and the ElevenLabs audio ecosystem are the priority.

Explore Scribe v2

Miscellaneous

Voxtral Transcribe 2

Choose Voxtral Transcribe 2 for Mistral's batch and real-time speech stack, including an open-weight realtime option.

Explore Voxtral Transcribe 2

Consumer

Cohere Transcribe

Choose Cohere Transcribe when a free open-source speech-recognition model and self-managed deployment matter most.

Explore Cohere Transcribe

Questions

Gemini 3.5 Transcribe FAQs

What is Gemini 3.5 Transcribe?

Gemini 3.5 Transcribe is Google's dedicated public-preview speech-to-text model. It supports live streaming and prerecorded audio with automatic language detection, custom vocabulary, smart formatting, and file-only speaker and timestamp metadata.

What is the difference between Gemini 3.5 Transcribe and Transcribe Live?

gemini-3.5-transcribe processes audio files and supports speaker diarization and word timestamps. gemini-3.5-transcribe-live streams audio over WebSockets with sub-second latency but does not support those two metadata features.

How much does Gemini 3.5 Transcribe cost?

Google estimates about $0.005 per minute for prerecorded transcription and $0.009 per minute for Live. Billing is actually calculated from separate audio-input and text-output token rates.

Does Gemini 3.5 Transcribe remove filler words?

Smart transcription can remove fillers, repetitions, and disfluencies, resolve self-corrections, and apply punctuation and normalization. Disable cleanup when exact spoken wording must be preserved.

How many languages does it support?

Google lists more than 85 supported languages and locales, with automatic detection and code-switching during a session.

Can Gemini 3.5 Transcribe identify speakers?

The prerecorded model can label up to eight speakers, but attribution for three or more is experimental. Live Transcribe does not currently support diarization.

Does it provide word-level timestamps?

The prerecorded model can provide word start and end offsets. Live Transcribe cannot, and Google notes that word timestamps can reduce transcription accuracy.

What is the maximum recording length?

Prerecorded audio can be up to one hour, or 30 minutes when diarization or word timestamps are enabled. A Live session is limited to 10 minutes.

Can I add custom vocabulary?

Yes. The model accepts up to 1,000 vocabulary phrases for names, acronyms, and domain terms, though Google says customers generally see the best results with a focused list of roughly 100.

Does Google train on Gemini 3.5 Transcribe audio?

For unpaid Google AI Studio and Gemini API access, Google can use inputs and outputs to improve products and human reviewers may process them. For paid Gemini API access through an active billing project, Google says prompts and responses are not used for product improvement, though limited safety logging remains.

Bottom line

Our Gemini 3.5 Transcribe verdict

Gemini 3.5 Transcribe is a compelling new option for developers who want both live voice interaction and low-cost file transcription in the Gemini ecosystem. Language detection, code-switching, vocabulary biasing, diarization, timestamps, and especially Smart transcription cover a wide range of products. Its preview status and the distinction between polished dictation and verbatim records are the main caveats: production teams should use paid data terms, benchmark their own accents and audio conditions, preserve source audio, and require human verification for consequential transcripts.

Visit Gemini 3.5 Transcribe website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.