The Rundown AI homepage

Independent tool overview

AssemblyAI at a glance

AssemblyAI is a developer platform for transcribing pre-recorded, live, and short-form audio, then extracting structured information through diarization, speaker identification, summaries, sentiment, topics, entities, guardrails, and LLM workflows.

Visit the official AssemblyAI site ↗
AssemblyAI product preview
Product type
Speech AI APIs and managed voice-agent infrastructure
Core modes
Pre-recorded, realtime, sync short clips, and full voice agents
Current flagship
Universal-3.5 Pro for async and realtime transcription
Broad-language option
Universal-2 supports 99 languages
Developer access
REST, WebSocket, Python and JavaScript SDKs, and integrations
Billing
Pay as you go, generally prorated to the second
Regions
US and EU endpoints are documented
Last reviewed
August 30, 2026

Overview

What AssemblyAI is

AssemblyAI is primarily an API platform rather than a finished meeting-notes app. Developers can submit audio or video for asynchronous transcription, stream live speech over WebSocket, transcribe clips of up to two minutes in a synchronous request, or use a managed Voice Agent API that bundles speech recognition, a conversational model, text-to-speech, and session infrastructure.

The platform offers a broad path from audio to structured data, but model choice, add-ons, retention settings, region, and human review matter. Speech recognition errors can alter names, numbers, negation, technical terms, speaker attribution, and meaning, while downstream summaries and classifications can add a second layer of uncertainty.

Use cases

Who AssemblyAI is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Voice-agent developers

Use low-latency transcription, turn detection, contextual prompting, and a bundled agent stack for phone or in-app conversations.

Meeting and call products

Create transcripts with timestamps, speaker segmentation, summaries, chapters, entities, sentiment, and other structured outputs.

Media and podcast workflows

Generate searchable transcripts, captions, chapters, topics, and post-production metadata from recorded audio or video.

Multilingual transcription products

Choose a broad 99-language value model or newer 18-language and six-language models based on quality, latency, and cost.

Regulated teams with engineering resources

Evaluate EU processing, BAA, retention, deletion, opt-out, redaction, and self-hosted arrangements against formal requirements.

Capabilities

Core AssemblyAI features

1

Pre-recorded transcription

Submit an audio or video source, poll or receive a webhook, and retrieve a completed transcript with formatting and timing data.

2

Universal-3.5 Pro

The current high-accuracy async model supports 18 languages, code switching, and updated speaker diarization.

3

Universal-2

A lower-cost pre-recorded model supporting 99 languages.

4

Realtime transcription

Streams live audio over WebSocket with punctuation, casing, end-of-turn detection, and model-specific context controls.

5

Universal-Streaming

A lower-cost realtime option for English or a six-language multilingual model, priced by connected session duration.

6

Sync API

Returns a Universal-3.5 Pro transcript in one response for clips up to two minutes, with timing and confidence data.

7

Speaker diarization and identification

Segments a conversation by speaker and can map speakers to names or roles through separate capabilities.

8

Accuracy controls

Model-specific keyterms, prompting, custom spelling, conversation context, and medical terminology options help adapt recognition.

9

Speech understanding

Offers translation, custom formatting, entities, sentiment, chapters, key phrases, topics, and summarization.

10

PII redaction

Can mask configured sensitive categories in transcript text and generate audio with detected PII replaced by a beep or silence.

11

Content guardrails

Optional profanity filtering and content moderation can label or transform selected outputs.

12

LLM Gateway

Applies supported language models to audio-derived content without the customer separately hosting those models.

13

Voice Agent API

Bundles realtime speech-to-text, a voice-oriented LLM, text-to-speech, turn and interruption handling, recordings, transcripts, and hosting.

14

Project-scoped API keys

Projects isolate keys, uploads, transcripts, and other project data across environments or applications.

15

Enterprise options

Custom rate limits, concurrency, volume pricing, support, security arrangements, and self-hosted deployment are available through sales.

Process

How the AssemblyAI workflow works

  1. Step 1

    Define the audio contract

    Document languages, audio sources, latency target, speakers, accuracy thresholds, required outputs, consent, and the decisions the transcript will support.

  2. Step 2

    Set privacy terms first

    Choose US or EU processing, execute any required BAA or contract, configure TTL and deletion, and confirm model-training opt-out before sending sensitive audio.

  3. Step 3

    Benchmark real samples

    Create human-verified reference transcripts across accents, noise, channels, devices, languages, names, numbers, and difficult domain vocabulary.

  4. Step 4

    Choose the narrowest model

    Select async, realtime, sync, or the bundled agent API, then enable only the add-ons that improve the measured use case.

  5. Step 5

    Build a secure integration

    Keep keys server-side, isolate projects, validate URLs and webhooks, close streaming sessions, handle retries idempotently, and avoid logging raw sensitive content.

  6. Step 6

    Review consequential output

    Require human verification for names, numbers, quotations, commitments, diagnoses, legal statements, speaker identity, summaries, and automated actions.

  7. Step 7

    Monitor and delete

    Track error rates and cost by model and feature, test regressions after model updates, audit access, and verify deletion against the intended retention policy.

Cost

AssemblyAI pricing and free plan

AssemblyAI starts with free credit and then uses pay-as-you-go pricing. Pre-recorded media is charged from audio duration; streaming and Voice Agent sessions are charged from connected session duration. Rates are displayed per hour but prorated to the second, and optional understanding or guardrail features are usually additive.

Pre-recorded speech-to-text

$0.15–$0.21/hour

Universal-2 costs $0.15 per audio hour; Universal-3.5 Pro costs $0.21 per audio hour.

  • Universal-2 supports 99 languages
  • Universal-3.5 Pro supports 18 languages
  • Speaker diarization adds $0.02/hour
  • Medical Mode adds $0.15/hour

Realtime speech-to-text

$0.15–$0.45/session hour

Universal-Streaming costs $0.15 per connected hour; Universal-3.5 Pro Realtime costs $0.45 per connected hour.

  • Multilingual Universal-Streaming is also $0.15/hour
  • Realtime diarization adds $0.12/hour
  • Voice Focus adds $0.10/hour on Universal-3.5 Pro Realtime
  • Close sessions correctly to stop billing

Sync API

$0.45/audio hour

Single-request transcription for clips up to two minutes using Universal-3.5 Pro accuracy.

  • Keyterms and conversation context included
  • Prompting adds $0.05/hour
  • Billed on actual clip duration

Voice Agent API

$4.50/connected hour

Bundled speech-to-text, voice LLM, text-to-speech, hosting, orchestration, recording, and turn handling.

  • $0.075 per connected minute
  • Billed per second
  • No per-agent or concurrency fee listed
  • Bring your own Twilio account for telephony

Understanding and guardrails

$0.01–$0.15/audio hour per feature

Optional features are priced separately and can stack on the transcription rate.

  • Summarization $0.03/hour
  • Sentiment $0.02/hour
  • PII text redaction $0.08/hour
  • PII audio redaction $0.05/hour
  • Content moderation $0.15/hour

Custom and enterprise

Contact sales

For volume discounts, custom concurrency and limits, enterprise controls, specialized support, or self-hosted arrangements.

  • EU endpoints are documented at the same public transcription rates
  • Contract and data terms should be reviewed separately
  • LLM Gateway token charges depend on provider and model

Pricing checked . Check current pricing at the source ↗

Assessment

AssemblyAI strengths and limitations

Where it stands out

  • Covers async, realtime, synchronous short clips, and full voice agents in one platform
  • Transparent public base rates with per-second proration
  • Offers a 99-language value model and newer higher-accuracy models
  • Broad set of transcription, understanding, and guardrail features
  • Project-scoped keys and data help separate environments
  • US and EU processing endpoints are documented
  • Python, JavaScript, REST, WebSocket, and no-code integration paths
  • Custom keyterms, prompting, spelling, and context can improve domain performance
  • PII redaction can operate on both text and generated redacted audio
  • Enterprise and self-hosted paths exist for organizations with stricter requirements

What to consider

  • AssemblyAI is an API platform, so a production user experience, storage model, permissions, editing, and monitoring must still be built
  • Accuracy varies with language, accent, code switching, noise, channel quality, crosstalk, speed, and domain vocabulary
  • Provider benchmark claims should be validated on the application's own representative audio
  • Transcription mistakes in names, numbers, negation, codes, and technical terms can materially change meaning
  • Speaker diarization and identification can merge, split, or mislabel speakers
  • Summaries, sentiment, entities, topics, moderation, and LLM outputs add probabilistic errors beyond the transcript
  • PII detection is not complete and cannot guarantee regulatory compliance
  • AssemblyAI warns that text redaction applies to the transcript text property while entities, summaries, or other features may still expose PII
  • Language and feature support varies significantly by model
  • The Sync API is limited to clips of two minutes
  • Streaming charges run for session duration, and improperly closed sessions can create unexpected cost
  • Multiple add-ons can make the effective hourly rate much higher than the base model price
  • Without TTL, BAA, or customer deletion, async final transcript artifacts may be retained indefinitely
  • Certain submitted files may be used for AssemblyAI model improvement when permitted by contract unless the customer uses EU servers, has a BAA, or completes an applicable opt-out
  • Free users cannot opt out of the model-improvement program under the current policy
  • The public Playground is a demo and AssemblyAI advises against uploading sensitive data there
  • Medical Mode is terminology optimization, not a clinical system or a substitute for clinician verification
  • Recording and transcribing people can require notice, consent, labor review, or other legal compliance
  • LLM Gateway data handling and retention vary by selected provider and model
  • API outages, rate limits, model updates, and vendor changes can affect production behavior

Compare

AssemblyAI alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Miscellaneous

Voxtral Transcribe 2

A Mistral transcription family with a live model and an open-weight realtime option for more deployment control.

Explore Voxtral Transcribe 2

Consumer

Cohere Transcribe

An open-source speech-recognition model for teams that prioritize self-hosting and model access over a managed API suite.

Explore Cohere Transcribe

Questions

AssemblyAI FAQs

What is AssemblyAI?

AssemblyAI is a developer platform for speech-to-text, speech understanding, content guardrails, LLM analysis, and managed voice agents. It is an API service rather than a standalone consumer transcription editor.

How much does AssemblyAI cost?

Current public rates start at $0.15 per hour for Universal-2 async or Universal-Streaming realtime. Universal-3.5 Pro async is $0.21 per hour, its realtime model is $0.45 per connected hour, and the bundled Voice Agent API is $4.50 per connected hour. Add-ons cost extra.

Which model should I use?

Benchmark the application. Universal-3.5 Pro is the current higher-accuracy option for supported languages, Universal-2 offers 99-language coverage at lower cost, Universal-Streaming is the lower-cost realtime option, and Sync is for clips up to two minutes.

Does AssemblyAI identify speakers?

It offers speaker diarization to segment speakers and a separate Speaker Identification feature to map them to names or roles. Both are probabilistic and require verification for consequential use.

Can AssemblyAI redact PII?

Yes. It can replace selected PII categories in transcript text and generate audio with detected PII beeped or silenced. Redaction may miss data, and other outputs such as entities or summaries can still contain PII.

Does AssemblyAI use customer audio for training?

Its current policy says certain files may be used when the contract permits, after a PII-redaction process. It says files are not used when covered by a BAA, processed on EU servers, or after an applicable paid-plan opt-out. Free users cannot currently opt out.

How long does AssemblyAI retain transcripts?

It depends on product and account configuration. Streaming can have zero audio and transcript retention after opt-out, while async final transcripts may be indefinite without TTL, BAA, or deletion. Configure and verify retention before production.

Does AssemblyAI support healthcare audio?

It offers Medical Mode, BAAs, retention controls, redaction, and documented security features, but an organization must confirm eligibility and configure the service correctly. Transcripts and clinical content still require qualified review.

Can I test AssemblyAI without code?

Yes, through the Playground and supported no-code integrations. AssemblyAI describes the Playground as a public demo with limited functionality and advises users not to upload sensitive data.

Bottom line

Our AssemblyAI verdict

AssemblyAI offers one of the broadest current developer surfaces for turning speech into transcripts, structured signals, and live voice experiences. Its strongest fit is a team prepared to benchmark real audio, design privacy and retention deliberately, and keep humans responsible for consequential words, speakers, summaries, and actions.

Visit AssemblyAI website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.