Meeting and interview archives
Create searchable transcripts with speaker labels and word-level timestamps from long recordings.
Independent tool overview
Voxtral Transcribe 2 is Mistral's speech-to-text family, pairing a low-cost batch API with a 4B real-time model that is also available as Apache-licensed open weights for private or edge deployment.
Visit the official Voxtral Transcribe 2 site ↗
Overview
Voxtral Transcribe 2 includes two distinct products. Voxtral Mini Transcribe V2 handles recorded audio with speaker labels, word timestamps and vocabulary hints, while Voxtral Realtime transcribes a live audio stream with configurable latency.
Both models support 13 languages: English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian and Dutch. The batch model can process recordings up to three hours in one request; the real-time model is aimed at live captions, voice agents and call-assist systems.
The deployment choice matters. The batch model is a hosted Mistral API at $0.003 per audio minute. Realtime costs $0.006 per minute through the API or can be self-hosted from the 4B Apache 2.0 checkpoint, shifting cost and privacy responsibility to the operator.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Create searchable transcripts with speaker labels and word-level timestamps from long recordings.
Stream text with adjustable delay so downstream models can respond while a person is still speaking.
Self-host the Realtime checkpoint when audio must remain inside controlled infrastructure.
Capabilities
Processes recorded audio through Voxtral Mini Transcribe V2, including files up to three hours per request.
Voxtral Realtime uses a native streaming architecture with delay configurable for latency-versus-accuracy trade-offs.
Labels speakers and provides start and end times for multi-party recordings in the batch workflow.
Returns timing for individual words to support subtitles, search, playback alignment and review tools.
Accepts up to 100 names, technical terms or phrases to improve expected spellings; English support is the most mature.
Is designed for difficult environments such as call centers, factory floors and field recordings.
Makes the 4B Realtime weights available under Apache 2.0 for edge, private-cloud or on-premises use.
Process
Step 1
Use batch for completed recordings and richer metadata; use Realtime when partial text must arrive during the conversation.
Step 2
Normalize the input where appropriate and provide a compact list of important names or domain terms.
Step 3
Store timestamps, speaker labels and source-audio links so users can inspect uncertain passages.
Step 4
Verify names, numbers, commitments and regulated information against the audio before downstream action or publication.
Cost
Mistral bills the hosted models by audio minute. The open Realtime checkpoint has no model license fee, but self-hosting adds compute and operations costs.
$0.003/minute
Hosted batch transcription with diarization, timestamps and context biasing.
$0.006/minute
Hosted streaming transcription for live applications.
$0 license fee
Self-host the 4B BF16 checkpoint under Apache 2.0.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Content Creator
A mature hosted speech-to-text API with a broader menu of audio-intelligence features.
Explore AssemblyAI →Marketing
A broader voice platform that includes speech recognition alongside synthesis and agent tooling.
Explore ElevenLabs →Content Creator
A simpler creator workflow for transcribing and editing podcast episodes rather than building an API product.
Explore Transistor AI Transcription →Questions
Mini Transcribe V2 is the hosted batch model for completed recordings, with diarization, context biasing and word timestamps. Realtime is the low-latency streaming model and is also available as open weights.
The batch API costs $0.003 per audio minute and the Realtime API costs $0.006 per minute. That is about $3 or $6 respectively for 1,000 minutes, before other infrastructure or application costs.
English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian and Dutch.
Mistral releases the 4B Realtime weights under the Apache 2.0 license. The hosted batch model is accessed through Mistral's API rather than downloaded from that checkpoint.
Mistral advertises configurable delay down to below 200 milliseconds. Lower delay is not automatically better; choose a setting by measuring accuracy and responsiveness on your audio.
It can supply timestamps and speaker labels, and Mistral offers private deployment options. Human verification, recording consent, retention policy and the rest of a compliant process are still required.
Bottom line
Voxtral Transcribe 2 is a strong option for teams that need both economical batch transcription and a real-time path. The open 4B streaming model is the differentiator, while high-stakes transcripts still require source-audio review and a carefully governed deployment.
Visit Voxtral Transcribe 2 website ↗
SAM Audio - Meta's model to separate any sound from an audio or audio-visual source using natural language prompts

Model Council - Perplexity's new tool for querying and synthesizing outputs from multiple models into a single answer

GLM-4.6V - Zhipu AI's open-source multimodal model family with native tool use capabilities
.png)
Raven-1 - Tavus's real-time emotional perception model for AI conversations

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.