Conversational voice agents
Convert live caller speech into text quickly enough to feed an agent response loop.
Independent tool overview
Scribe v2 Realtime is ElevenLabs' streaming speech-to-text model for voice agents, live captions, meeting assistants, and other latency-sensitive applications. It returns partial transcripts in roughly 150 milliseconds, supports more than 90 languages, and is available through the ElevenLabs API and ElevenAgents.
Visit the official Scribe v2 Realtime site ↗
Overview
Scribe v2 Realtime is a developer-facing transcription model rather than a standalone meeting-notes app. Applications stream microphone or telephony audio over a WebSocket and receive partial text while a person is speaking, followed by committed transcript segments when speech is finalized.
The model is designed for conversational responsiveness. It combines automatic language detection, voice activity detection, manual commit controls, previous-text conditioning, precise word timestamps, and entity detection. ElevenLabs says it can handle language switching within a conversation and supports PCM audio from 8kHz to 48kHz plus mu-law encoding.
Choose it when low latency matters more than the fuller post-processing available from a batch model. Teams still need to build the surrounding capture, consent, storage, retry, and transcript-review experience themselves.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Convert live caller speech into text quickly enough to feed an agent response loop.
Generate partial and finalized captions for meetings, broadcasts, classes, or accessibility interfaces.
Add live transcripts to a custom application while retaining control over the user interface and downstream workflow.
Support conversations across more than 90 languages, including language changes within a session.
Capabilities
Returns partial transcript updates in roughly 150 milliseconds through a WebSocket connection.
Automatically recognizes supported languages and can follow multilingual speech within the same conversation.
Applications can show provisional words immediately and replace them with stable transcript segments after a commit.
Silence detection can segment speech automatically, while manual commit gives developers explicit finalization control.
Precise timing metadata supports synchronized captions and downstream audio navigation.
Previous transcript context can be supplied after a reconnect so the model continues with better linguistic continuity.
The current model reference lists detection across 65 entity types for applications that need structured information from live speech.
Process
Step 1
A server can create a single-use token for client-side streaming so the permanent ElevenLabs API key is not exposed in the browser.
Step 2
Send microphone or telephony audio chunks to the realtime WebSocket in a supported PCM or mu-law format.
Step 3
Display partial transcript events for immediacy, understanding that predictions can still change.
Step 4
Use automatic voice activity detection or manual commit, then store or process only the finalized transcript according to the application's privacy policy.
Cost
ElevenLabs lists Scribe v2 Realtime at a standard API rate of $0.39 per audio hour. Subscription plans range from Free to Business and include different amounts of usage; the product page also advertises rates of $0.28 per hour or lower on annual Business plans. Taxes and any surrounding agent, storage, or application infrastructure are separate.
$0.39/audio hour
The published Scribe v2 Realtime model rate for free and self-serve API accounts.
$0-$990/month
Free, Starter, Creator, Pro, Scale, and Business plans include different realtime transcription allowances alongside other ElevenLabs products.
$0.28/audio hour or lower
A lower rate advertised for annual Business volume; actual effective cost depends on the selected commitment and usage.
Custom
For negotiated limits, support, data controls, and compliance requirements.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Consumer
An open-source speech-recognition option for teams prioritizing model control and self-hosting.
Explore Cohere Transcribe →Content Creator
A Google transcription option for teams already building around the Gemini API ecosystem.
Explore Gemini 3.5 Transcribe →Miscellaneous
A Mistral speech-to-text family that includes an open-weights realtime model for greater deployment flexibility.
Explore Voxtral Transcribe 2 →Questions
It is ElevenLabs' streaming speech-to-text model for applications such as voice agents, live captions, and meeting assistants. It sends partial and committed transcripts over a realtime API connection.
ElevenLabs documents roughly 150 milliseconds of latency for partial transcription. End-to-end application latency will also include network, audio buffering, agent inference, and any speech-synthesis time.
The standard published API rate is $0.39 per audio hour as of August 29, 2026. Annual Business pricing can reach $0.28 per hour or lower, while subscription plans include different usage allowances.
ElevenLabs lists support for more than 90 languages with automatic language recognition and multilingual switching.
The realtime product FAQ says speaker diarization is not currently a priority. If speaker labels are essential, verify current support or use a batch transcription workflow designed for diarization.
Yes, client-side microphone streaming is supported, but the browser should use a short-lived single-use token created by a trusted server rather than exposing the permanent API key.
Bottom line
Scribe v2 Realtime is a compelling fit for developers who need responsive multilingual transcription and already plan to build the surrounding application. Test it with real accents, noise, telephony audio, and domain vocabulary before committing, and use batch Scribe v2 instead when finalized transcript richness matters more than immediacy.
Visit Scribe v2 Realtime website ↗
Atlas - OpenAI's new web browser built with ChatGPT at its core

Grok 4.1 - xAI's latest update to Grok with improved creative abilities, emotional intelligence, and real-world usability.

Apps SDK - Chat with and build apps directly in ChatGPT

Gemini 3 - Google's most intelligent model release to date, taking the top spot industry-wide across major benchmarks for multimodal understanding, agentic abilities, reasoning, and more.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.