The Rundown AI homepage

Independent tool overview

Scribe v2 at a glance

Scribe v2 is ElevenLabs' high-accuracy batch transcription model for uploaded audio and video. It supports more than 90 languages, speaker diarization, word-level timestamps, audio-event tags, multilingual files, up to 1,000 prompted keyterms, and optional sensitive-entity detection or redaction through the API.

Visit the official Scribe v2 site ↗
Scribe v2 product preview
Product type
Batch speech-to-text model and API
Developer
ElevenLabs
Languages
90+ supported
Speaker diarization
Up to 32 speakers
Keyterm prompting
Up to 1,000 terms
Entity categories
Up to 56 through the API
Standard upload limit
3GB and 10 hours
API base rate
$0.22 per audio hour

Overview

What Scribe v2 is

Scribe v2 is the batch and long-form member of ElevenLabs' speech-to-text family. It is designed for recordings such as interviews, podcasts, research sessions, training media, calls, and subtitle workflows where transcript quality and structure matter more than live response latency.

The model automatically detects languages within the same file, labels speakers, timestamps individual words, and tags non-speech events. The API can accept standard audio and video formats, handle files up to 3GB, and process recordings up to 10 hours in standard mode. A separate multichannel mode supports up to five channels with a one-hour maximum.

Two of its most useful controls are keyterm prompting and entity handling. Teams can supply up to 1,000 relevant words or phrases to improve domain vocabulary. The API can also detect up to 56 categories of personal, health, payment, or other sensitive information and can redact detected entities during transcription.

Scribe v2 should not be confused with Scribe v2 Realtime. The standard model is optimized for batch accuracy and long recordings; Realtime is a separate, higher-priced streaming model for live agents and interactive applications. Teams that need both usually route uploaded files to Scribe v2 and live audio to Realtime.

Use cases

Who Scribe v2 is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Long-form transcription

Convert interviews, podcasts, webinars, lectures, and recorded meetings into structured text.

Multilingual recordings

Transcribe files that contain multiple supported languages or code-switching without manually splitting each language segment.

Subtitles and captions

Use precise word-level timing and speaker labels to build synchronized caption workflows.

Technical vocabulary

Prompt the model with product names, specialist terms, people, or places that are likely to appear in the recording.

Sensitive-data workflows

Detect or redact supported entity types before storing or sending the transcript downstream, with human verification.

Automated media pipelines

Submit files through the API and deliver asynchronous results to a webhook for downstream processing.

Capabilities

Core Scribe v2 features

1

90+ language transcription

Supports a broad language set and can automatically detect language changes inside one recording.

2

Speaker diarization

Separates and labels up to 32 speakers in supported transcription workflows.

3

Word-level timestamps

Returns timing for individual words to support captions, search, playback highlighting, and editing.

4

Dynamic audio tagging

Identifies non-speech events such as laughter or footsteps alongside spoken content.

5

Context-aware keyterms

Accepts as many as 1,000 words or phrases and uses surrounding audio to determine when a supplied term applies.

6

Entity detection and redaction

Can find supported sensitive-data categories with timestamps or replace detected values using complete, categorized, or enumerated redaction.

7

No Verbatim mode

Optionally removes filler words, repeated phrases, and stuttering to produce a cleaner reading transcript.

8

Multichannel transcription

Processes up to five audio channels independently and assigns each channel a speaker ID.

9

Webhook delivery

Can send completed asynchronous transcription results to a configured endpoint.

Process

How the Scribe v2 workflow works

  1. Step 1

    Choose batch or realtime

    Use Scribe v2 for uploaded recordings and long-form accuracy; choose Scribe v2 Realtime when low-latency streaming is essential.

  2. Step 2

    Prepare the source media

    Confirm the format, file size, duration, channel layout, audio quality, and that you have consent and rights to process the recording.

  3. Step 3

    Configure transcript behavior

    Select language handling, speaker diarization, timestamps, audio tags, verbatim style, and any keyterms relevant to the recording.

  4. Step 4

    Configure sensitive-data controls

    Choose entity categories and redaction format where needed, and establish a second review because automated redaction can miss information.

  5. Step 5

    Submit through the UI or API

    Use ElevenLabs' speech-to-text interface for individual work or integrate the API and webhooks for repeatable pipelines.

  6. Step 6

    Review the result

    Check names, numbers, speaker changes, timestamps, code-switched segments, technical terms, and every critical redaction against the source audio.

  7. Step 7

    Export or route downstream

    Move approved text into subtitle, publishing, search, analytics, archive, or compliance systems with appropriate access controls.

Cost

Scribe v2 pricing and free plan

ElevenLabs sells Scribe v2 through Creative subscriptions with included UI transcription and through usage-based API pricing. Current API pricing is $0.22 per audio hour for Scribe v2, plus $0.05 per hour for keyterm prompting and $0.07 per hour for entity detection. Scribe v2 Realtime is a separate model at $0.39 per audio hour. Prices exclude taxes.

Free

$0 per month

Entry Creative plan with a small Scribe v2 UI allowance.

  • About 12 minutes of UI transcription per month
  • API overage rate shown at $0.22 per hour
  • Limits and concurrency are lower than paid plans

Starter

$6 per month

Creative plan for light recurring transcription.

  • About 1 hour 31 minutes of UI transcription per month
  • API overage rate shown at $0.22 per hour
  • Check the live plan table for all included ElevenLabs products

Creator

$22 per month

Creative plan with a larger monthly transcription allowance.

  • About 6 hours 7 minutes of UI transcription per month
  • API overage rate shown at $0.22 per hour
  • Promotional first-month pricing may be shown separately

Pro

$99 per month

Higher-volume Creative plan.

  • About 30 hours 18 minutes of UI transcription per month
  • API overage rate shown at $0.22 per hour
  • Higher listed concurrent-transcription allowance

Scribe v2 API

$0.22 per audio hour

Usage-based batch transcription for applications and automated pipelines.

  • Keyterm prompting adds $0.05 per audio hour
  • Entity detection adds $0.07 per audio hour
  • Requests using more than 100 keyterms have a 20-second minimum billable unit

Scribe v2 Realtime API

$0.39 per audio hour

Separate streaming model for live agents and low-latency transcription.

  • Not the same model path as batch Scribe v2
  • Designed for live audio rather than long uploaded files
  • Usage is billed by audio duration

Pricing checked . Check current pricing at the source ↗

Assessment

Scribe v2 strengths and limitations

Where it stands out

  • Broad language support with automatic multilingual transcription
  • Useful transcript structure from speakers, words, timestamps, and audio events
  • Large keyterm list for technical and brand vocabulary
  • Entity detection and multiple redaction formats built into the transcription request
  • Handles long and large uploaded audio or video files
  • Available through both a user interface and an automation-friendly API
  • Separate realtime model for applications that need live transcription

What to consider

  • No speech-to-text model is perfectly accurate, especially with noise, overlap, accents, names, and weak recordings
  • Entity detection and redaction can miss sensitive information and require human review
  • Keyterm prompting and entity detection add to API cost
  • More than 100 keyterms triggers a 20-second minimum billable unit
  • Standard mode is limited to 3GB and 10 hours per file; multichannel mode is limited to one hour
  • No Verbatim mode intentionally alters the spoken record and is unsuitable when every utterance must be preserved
  • Vendor benchmark and accuracy claims may not reflect a team's languages, microphones, speakers, or domain
  • HIPAA-related deployments require contacting ElevenLabs Sales and completing a Business Associate Agreement

Compare

Scribe v2 alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Content Creator

AssemblyAI

A developer-focused speech-to-text API worth comparing on language coverage, diarization, intelligence features, latency, and price.

Explore AssemblyAI

Project Management

Otter.ai

A meeting-oriented product with collaborative notes and a more packaged end-user workflow.

Explore Otter.ai

Content Creator

Descript

A transcript-based audio and video editor for creators who want editing tools around the transcription.

Explore Descript

Questions

Scribe v2 FAQs

What is ElevenLabs Scribe v2?

Scribe v2 is ElevenLabs' batch speech-to-text model for transcribing uploaded audio and video into structured text with language detection, speakers, timestamps, and optional advanced controls.

How many languages does Scribe v2 support?

ElevenLabs documents support for more than 90 languages, including automatic handling of multiple languages within one file.

What is the difference between Scribe v2 and Scribe v2 Realtime?

Scribe v2 is optimized for accurate batch and long-form transcription. Scribe v2 Realtime is a separate streaming model optimized for low-latency agents and live applications.

How much does the Scribe v2 API cost?

The current base API rate is $0.22 per audio hour before taxes. Keyterm prompting adds $0.05 per hour and entity detection adds $0.07 per hour.

Can Scribe v2 identify speakers?

Yes. Standard diarization supports up to 32 speakers, and a separate multichannel mode can process up to five channels independently.

Does Scribe v2 provide word-level timestamps?

Yes. The API can return start and end timing for words, which is useful for captions, playback synchronization, and searchable media.

What is keyterm prompting?

It lets a team provide up to 1,000 expected words or phrases, such as technical terms or names, so Scribe v2 can use the audio context to improve how those terms are transcribed.

Can Scribe v2 remove personal information?

The API can detect and redact supported entity types during transcription, but ElevenLabs warns that automated redaction may not identify everything. A human or secondary control should review sensitive transcripts.

What does No Verbatim mode do?

It removes filler words, repeated phrases, and stuttering to create a cleaner reading transcript. Leave it off when the exact spoken record matters.

What are the Scribe v2 file limits?

ElevenLabs documents a 3GB maximum file size and a 10-hour maximum duration in standard mode. Multichannel transcription has a one-hour maximum.

Can Scribe v2 transcribe video?

Yes. The API accepts multiple common video formats in addition to audio formats and transcribes the audio track.

Is Scribe v2 HIPAA compliant?

ElevenLabs says companies requiring HIPAA compliance must contact Sales and sign a Business Associate Agreement before using the service for HIPAA-related deployments.

Bottom line

Our Scribe v2 verdict

Scribe v2 is a strong fit for teams that need structured batch transcripts across many languages and want diarization, timestamps, vocabulary guidance, and sensitive-entity controls in one API. Its base usage price is straightforward, but advanced features increase cost and none of the accuracy or redaction features removes the need for review. Choose Realtime instead when live latency is the primary requirement, and compare packaged meeting or editing products if the team does not want to build its own workflow.

Visit Scribe v2 website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.