Live meeting and event transcription
Products that need low-latency text, speaker changes and turn endpoints while a conversation is still happening.
Independent tool overview
Muse Voice Transcribe is Meta’s real-time speech-to-text model for streaming transcription, speaker diarization, endpointing and multilingual code-switching.
Visit the official Muse Voice Transcribe site ↗
Overview
Muse Voice Transcribe combines streaming automatic speech recognition with speaker diarization and speech endpoint detection. It processes audio continuously, adapts how long it waits before committing each word and can label more than 20 speakers without a separate post-processing step.
The model is trained across more than 70 languages, with 25 languages extensively verified for the initial release. It is available through Meta Model API, Meta AI for Mac and Muse Code, making it usable in custom transcription products as well as Meta’s own desktop workflows.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Products that need low-latency text, speaker changes and turn endpoints while a conversation is still happening.
Workflows where speakers change languages within a sentence or meeting and a single-language model would lose important context.
Panels, classrooms, interviews and long group discussions that benefit from integrated speaker labels instead of a separate diarization pass.
Developers building dictation, assistants, call workflows or other products through Meta Model API’s streaming and file-transcription paths.
Capabilities
Processes audio in real time and emits text as the speaker talks rather than waiting for the entire recording to finish.
Marks speaker changes and assigns speaker labels within the same model workflow, including long audio with more than 20 speakers.
Identifies when speech starts and when a speaker has finished, which helps voice applications decide when to respond or close a turn.
Handles language changes within or between sentences instead of requiring the developer to split the conversation by language first.
Uses provided language or context hints to improve recognition of expected names, terms and domain-specific words.
Dynamically waits for more context on difficult words while returning easier portions faster to balance accuracy and latency.
Meta says the model natively supports audio longer than one hour with no required diarization post-processing.
Process
Step 1
Open a real-time connection for live audio or send a completed recording through the supported file-transcription workflow.
Step 2
Provide likely languages, keywords, names or domain context when available to improve recognition accuracy.
Step 3
Use returned text, speaker labels and speech endpoints to update the interface or trigger the next application action.
Step 4
Show that the transcript is AI-generated, allow correction and apply the required consent, privacy and retention rules for recorded audio.
Cost
Meta Model API bills Muse Voice Transcribe by processed audio time rather than tokens. Streaming and file transcription cost the same, and Meta lists zero-data-retention service at the standard rate.
$0.18 per audio hour
Usage-based pricing for both streaming and non-streaming transcription through Meta Model API.
Same $0.18 per audio hour rate
Meta lists zero-data-retention service at parity with Standard pricing.
Not available at launch
The training-eligible discounted tier offered for Muse Spark is not available for Muse Voice Transcribe at launch.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Content Creator
Choose AssemblyAI when you want a speech-focused API platform with a broader established transcription and audio-intelligence product set.
Explore AssemblyAI →Content Creator
Choose Gemini 3.5 Transcribe when Google’s AI ecosystem and its transcription model fit the rest of your application stack.
Explore Gemini 3.5 Transcribe →Consumer
Choose Cohere Transcribe when Cohere’s platform, deployment requirements or broader enterprise language-model workflow is the better match.
Explore Cohere Transcribe →Questions
Muse Voice Transcribe is Meta’s real-time audio model for speech-to-text, speaker diarization and speech endpoint detection.
Meta lists a price of $0.18 per hour of audio processed. Streaming and non-streaming transcription use the same rate, and platform free-tier credits may apply.
Meta says the model was trained with more than 70 languages and extensively verified 25 for the initial release. It recommends starting with those 25 validated languages.
Yes. It includes streaming speaker diarization and can handle long audio with more than 20 speakers, although labels should still be reviewed in difficult recordings.
Yes. Meta says it natively supports switching languages within or between sentences and can use language, keyword and context biasing.
It is available through Meta Model API, Meta AI for Mac and Muse Code. Developers can use the API for streaming or file-based transcription.
No. This independent overview is based on Meta’s official launch, speech-to-text and pricing documentation; The Rundown has not completed a controlled hands-on transcription comparison for this page.
Bottom line
Muse Voice Transcribe stands out for combining live text, speaker tracking, endpoint detection and multilingual code-switching at a simple published hourly rate. It is especially relevant to developers already considering Meta Model API or building multi-speaker voice experiences. Buyers should still test their languages, acoustic conditions and overlap patterns, and compare AssemblyAI or other speech-specialist platforms when mature audio analytics and operational tooling are a priority.
Visit Muse Voice Transcribe website ↗
Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.