The Rundown AI homepage

Independent tool overview

Muse Voice Transcribe at a glance

Muse Voice Transcribe is Meta’s real-time speech-to-text model for streaming transcription, speaker diarization, endpointing and multilingual code-switching.

Visit the official Muse Voice Transcribe site ↗
Muse Voice Transcribe product preview
Best for
Live multilingual transcription and speaker tracking
API price
$0.18 per audio hour
Languages
70+ trained; 25 extensively verified at launch
Speakers
20+ in long-form audio
Access
Meta Model API, Meta AI for Mac and Muse Code

Overview

What Muse Voice Transcribe is

Muse Voice Transcribe combines streaming automatic speech recognition with speaker diarization and speech endpoint detection. It processes audio continuously, adapts how long it waits before committing each word and can label more than 20 speakers without a separate post-processing step.

The model is trained across more than 70 languages, with 25 languages extensively verified for the initial release. It is available through Meta Model API, Meta AI for Mac and Muse Code, making it usable in custom transcription products as well as Meta’s own desktop workflows.

Use cases

Who Muse Voice Transcribe is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Live meeting and event transcription

Products that need low-latency text, speaker changes and turn endpoints while a conversation is still happening.

Multilingual conversations

Workflows where speakers change languages within a sentence or meeting and a single-language model would lose important context.

Multi-speaker recordings

Panels, classrooms, interviews and long group discussions that benefit from integrated speaker labels instead of a separate diarization pass.

Voice-enabled applications

Developers building dictation, assistants, call workflows or other products through Meta Model API’s streaming and file-transcription paths.

Capabilities

Core Muse Voice Transcribe features

1

Streaming speech recognition

Processes audio in real time and emits text as the speaker talks rather than waiting for the entire recording to finish.

2

Integrated speaker diarization

Marks speaker changes and assigns speaker labels within the same model workflow, including long audio with more than 20 speakers.

3

Speech endpoint detection

Identifies when speech starts and when a speaker has finished, which helps voice applications decide when to respond or close a turn.

4

Multilingual code-switching

Handles language changes within or between sentences instead of requiring the developer to split the conversation by language first.

5

Language, keyword and context biasing

Uses provided language or context hints to improve recognition of expected names, terms and domain-specific words.

6

Adaptive transcription delay

Dynamically waits for more context on difficult words while returning easier portions faster to balance accuracy and latency.

7

Long-form audio support

Meta says the model natively supports audio longer than one hour with no required diarization post-processing.

Process

How the Muse Voice Transcribe workflow works

  1. Step 1

    Choose streaming or file transcription

    Open a real-time connection for live audio or send a completed recording through the supported file-transcription workflow.

  2. Step 2

    Configure language and context hints

    Provide likely languages, keywords, names or domain context when available to improve recognition accuracy.

  3. Step 3

    Consume transcript events

    Use returned text, speaker labels and speech endpoints to update the interface or trigger the next application action.

  4. Step 4

    Review and retain appropriately

    Show that the transcript is AI-generated, allow correction and apply the required consent, privacy and retention rules for recorded audio.

Cost

Muse Voice Transcribe pricing and free plan

Meta Model API bills Muse Voice Transcribe by processed audio time rather than tokens. Streaming and file transcription cost the same, and Meta lists zero-data-retention service at the standard rate.

API audio processing

$0.18 per audio hour

Usage-based pricing for both streaming and non-streaming transcription through Meta Model API.

  • Metered by minutes of audio processed
  • Streaming and file transcription use the same rate
  • No token-based charge for the transcription itself

Zero data retention

Same $0.18 per audio hour rate

Meta lists zero-data-retention service at parity with Standard pricing.

  • Eligibility and implementation requirements should be confirmed
  • Use appropriate consent and privacy controls for audio
  • Platform free-tier credits may apply

Contributor discount

Not available at launch

The training-eligible discounted tier offered for Muse Spark is not available for Muse Voice Transcribe at launch.

  • Standard audio pricing applies
  • No training-data discount for this model at launch
  • Recheck current pricing before forecasting large volumes

Pricing checked . Check current pricing at the source ↗

Assessment

Muse Voice Transcribe strengths and limitations

Where it stands out

  • Combines transcription, speaker labeling and turn endpoint detection in one real-time model.
  • Supports natural code-switching and context biasing for multilingual or domain-specific conversations.
  • Handles long recordings and more than 20 speakers without a required diarization post-processing stage.
  • The published $0.18-per-hour API rate makes the cost structure easy to estimate from audio volume.

What to consider

  • Transcripts are AI-generated and may be inaccurate, so consequential records still need human review and correction.
  • Although trained with more than 70 languages, Meta extensively verified 25 for the initial release and recommends starting with those.
  • Speaker labels can still be challenged by overlapping speech, noise, accents or poor microphone placement despite integrated diarization.
  • The contributor-priced training-data discount available for Muse Spark is not offered for Muse Voice Transcribe at launch.
  • Applications that capture other people’s audio must implement appropriate notice, consent, security and retention policies.

Compare

Muse Voice Transcribe alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Content Creator

AssemblyAI

Choose AssemblyAI when you want a speech-focused API platform with a broader established transcription and audio-intelligence product set.

Explore AssemblyAI

Content Creator

Gemini 3.5 Transcribe

Choose Gemini 3.5 Transcribe when Google’s AI ecosystem and its transcription model fit the rest of your application stack.

Explore Gemini 3.5 Transcribe

Consumer

Cohere Transcribe

Choose Cohere Transcribe when Cohere’s platform, deployment requirements or broader enterprise language-model workflow is the better match.

Explore Cohere Transcribe

Questions

Muse Voice Transcribe FAQs

What is Muse Voice Transcribe?

Muse Voice Transcribe is Meta’s real-time audio model for speech-to-text, speaker diarization and speech endpoint detection.

How much does Muse Voice Transcribe cost?

Meta lists a price of $0.18 per hour of audio processed. Streaming and non-streaming transcription use the same rate, and platform free-tier credits may apply.

How many languages does Muse Voice Transcribe support?

Meta says the model was trained with more than 70 languages and extensively verified 25 for the initial release. It recommends starting with those 25 validated languages.

Can Muse Voice Transcribe identify speakers?

Yes. It includes streaming speaker diarization and can handle long audio with more than 20 speakers, although labels should still be reviewed in difficult recordings.

Does Muse Voice Transcribe support code-switching?

Yes. Meta says it natively supports switching languages within or between sentences and can use language, keyword and context biasing.

Where can I use Muse Voice Transcribe?

It is available through Meta Model API, Meta AI for Mac and Muse Code. Developers can use the API for streaming or file-based transcription.

Is this a hands-on Muse Voice Transcribe review?

No. This independent overview is based on Meta’s official launch, speech-to-text and pricing documentation; The Rundown has not completed a controlled hands-on transcription comparison for this page.

Bottom line

Our Muse Voice Transcribe verdict

Muse Voice Transcribe stands out for combining live text, speaker tracking, endpoint detection and multilingual code-switching at a simple published hourly rate. It is especially relevant to developers already considering Meta Model API or building multi-speaker voice experiences. Buyers should still test their languages, acoustic conditions and overlap patterns, and compare AssemblyAI or other speech-specialist platforms when mature audio analytics and operational tooling are a priority.

Visit Muse Voice Transcribe website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.