The Rundown AI homepage

Independent tool overview

Cohere Transcribe at a glance

Cohere Transcribe is a 2-billion-parameter speech-to-text model that teams can run from open weights, test through Cohere's free rate-limited API, or deploy on dedicated Model Vault infrastructure.

Visit the official Cohere Transcribe site ↗
Cohere Transcribe product preview
Model
cohere-transcribe-03-2026
Size
2 billion parameters
Languages
14
License
Apache 2.0
Input / output
Audio to transcribed text
API file limit
25 MB
Deployment
Self-hosted, Cohere API, or Model Vault
Last reviewed
August 29, 2026

Overview

What Cohere Transcribe is

Cohere Transcribe is a dedicated audio-in, text-out automatic speech recognition model released under the Apache 2.0 license. It supports 14 languages and is designed for transcription pipelines such as searchable audio, meeting intelligence, call analytics, and voice automation.

The core model, cohere-transcribe-03-2026, uses a Conformer-based encoder-decoder architecture. Cohere ranked it first on the Hugging Face Open ASR Leaderboard at launch with a 5.42 average word error rate, but that vendor-reported aggregate should be treated as a starting point rather than proof that it will lead on every accent, microphone, industry term, or noise condition.

Its biggest practical advantage is deployment flexibility: developers can download the weights for local or private infrastructure, use Transformers for offline jobs, serve concurrent requests with vLLM, test Cohere's API, or move to a dedicated Model Vault deployment. Cohere has also released a separate Arabic-specialized model for dialects, Arabic-English code-switching, and domain vocabulary.

Use cases

Who Cohere Transcribe is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Private transcription pipelines

Teams that need to keep audio on controlled infrastructure can use the open weights instead of sending every recording to a shared hosted endpoint.

Multilingual business audio

The general model covers 14 specified languages for meetings, calls, media archives, and other monolingual transcription jobs.

Search and analytics

Transcripts can feed enterprise search, retrieval, summaries, quality review, and downstream analytics workflows.

High-volume engineering teams

Transformers supports offline and batched jobs, while Cohere recommends vLLM for production serving and concurrent requests.

Capabilities

Core Cohere Transcribe features

1

Open model weights

The 2B model is distributed through Hugging Face under Apache 2.0, allowing commercial use and self-managed deployment subject to the license terms.

2

Fourteen-language transcription

It supports English, French, German, Italian, Spanish, Portuguese, Greek, Dutch, Polish, Mandarin Chinese, Japanese, Korean, Vietnamese, and Arabic.

3

Long-form and batched processing

The Transformers processor can split longer recordings into chunks, reassemble the text, and process multiple audio files in one call.

4

Punctuation control

Developers can request punctuated output or lower-cased text without punctuation.

5

Flexible deployment

Use Transformers for offline inference, vLLM for production serving, Cohere's hosted API for evaluation, or a dedicated Model Vault for managed capacity.

6

Arabic-specialized option

Cohere Transcribe Arabic is a separate current model aimed at dialect variation, Arabic-English code-switching, and domain-specific vocabulary.

7

Audio preprocessing

The processor resamples audio to 16 kHz and averages stereo inputs into a single channel before transcription.

8

Benchmark-backed launch

Cohere reported a 5.42 average WER and the top position on the Open ASR Leaderboard at launch, alongside multilingual and human-preference evaluations.

Process

How the Cohere Transcribe workflow works

  1. Step 1

    Choose the model

    Use the general 14-language model for supported monolingual audio or evaluate the Arabic-specialized release for Arabic dialect and code-switching workloads.

  2. Step 2

    Run a representative test set

    Sample real accents, background noise, microphones, names, and domain vocabulary rather than relying only on leaderboard averages.

  3. Step 3

    Prepare the audio

    For the Cohere API, keep each upload within the documented 25 MB limit. For self-hosted jobs, use the processor's resampling and long-form chunking support.

  4. Step 4

    Specify the language

    Pass the correct supported language tag because the general model does not automatically detect language and performs best on single-language audio.

  5. Step 5

    Add speech detection

    Place voice activity detection or a noise gate before the model to reduce hallucinated text during silence or low-level background noise.

  6. Step 6

    Add missing post-processing

    Use a separate system when the product requires speaker labels or timestamps, because the model card says Cohere Transcribe does not provide them.

  7. Step 7

    Select deployment

    Start with the free rate-limited API or local weights, then compare the operating cost and controls of self-hosting against a dedicated Model Vault.

  8. Step 8

    Monitor production quality

    Track error rate, latency, failed files, hallucinations, and per-language performance as audio conditions and vocabulary change.

Cost

Cohere Transcribe pricing and free plan

The model weights are free under Apache 2.0, but self-hosting still incurs infrastructure and engineering costs. Cohere offers a free rate-limited API for experiments and custom-priced dedicated Model Vault capacity for production.

Open weights

Free

Download and run the model under Apache 2.0.

  • Hugging Face access requires accepting the repository conditions and sharing contact information
  • Compute, storage, observability, and engineering are not included
  • Transformers and vLLM integrations are documented

Cohere API evaluation

Free with rate limits

Hosted access for low-setup experimentation.

  • Subject to Cohere rate limits
  • Maximum file size is 25 MB
  • Intended as an evaluation path rather than unmetered production capacity

Model Vault

Custom

Dedicated managed inference without the shared API's rate limits.

  • Pricing is calculated per hour-instance
  • Longer-term commitments can receive discounts
  • Provisioning and current quotes are handled through Cohere

Private deployment

Contact sales

Enterprise deployment on controlled infrastructure.

  • Final cost depends on infrastructure and support requirements
  • Evaluate data residency, security, and operational needs with Cohere

Pricing checked . Check current pricing at the source ↗

Assessment

Cohere Transcribe strengths and limitations

Where it stands out

  • Open Apache 2.0 weights provide meaningful deployment flexibility.
  • Supports 14 languages in one general model.
  • Designed as a dedicated ASR model instead of a general language model with audio added on.
  • Supports long recordings, batching, punctuation control, and production serving through documented integrations.
  • Offers a low-friction hosted evaluation path and dedicated managed deployment option.
  • Cohere publishes benchmark, architecture, training, and model-limit information.

What to consider

  • The general model requires a specified language and does not provide automatic language detection.
  • It performs best on monolingual audio and can be inconsistent on code-switched recordings.
  • The model does not provide timestamps or speaker diarization.
  • Silence and low-volume noise can produce hallucinated text unless a noise gate or voice activity detector is used.
  • The hosted API limits files to 25 MB and the free endpoint is rate limited.
  • The Hugging Face repository is gated and requires users to share contact information before downloading files.
  • Self-hosting a 2B model still requires suitable compute, deployment work, monitoring, and security review.
  • Leaderboard word error rate does not guarantee accuracy on a team's own accents, terminology, or recording conditions.
  • Only 14 languages are supported by the general model.

Compare

Cohere Transcribe alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Consumer

Scribe v2 Realtime

Consider ElevenLabs when live, low-latency transcription and a broader language list are more important than open self-hosted weights.

Explore Scribe v2 Realtime

Project Management

Notta Showcase

Consider Notta for a ready-made meeting and recording workflow instead of building an ASR pipeline from model components.

Explore Notta Showcase

Consumer

Willow Frontier Mini

Consider Willow for end-user voice dictation rather than an enterprise speech-to-text model and API.

Explore Willow Frontier Mini

Questions

Cohere Transcribe FAQs

What is Cohere Transcribe?

It is Cohere's dedicated 2-billion-parameter automatic speech recognition model for converting audio into text across 14 supported languages.

Is Cohere Transcribe free?

The model weights are free under Apache 2.0, and Cohere provides a free rate-limited API for experiments. Self-hosted compute and dedicated Model Vault production capacity still cost money.

Can Cohere Transcribe be self-hosted?

Yes. Cohere publishes the weights on Hugging Face and documents offline inference through Transformers and production serving through vLLM.

Which languages does it support?

English, French, German, Italian, Spanish, Portuguese, Greek, Dutch, Polish, Mandarin Chinese, Japanese, Korean, Vietnamese, and Arabic.

Does Cohere Transcribe identify speakers?

No. The official model card says it does not include speaker diarization, so a separate diarization step is needed for speaker labels.

Does it return timestamps?

No. The official model card lists timestamps as unsupported.

Does it detect the spoken language automatically?

No. The model expects a language tag and performs best on one specified language at a time.

What is the Cohere Transcribe API upload limit?

Cohere's documentation lists a maximum file size of 25 MB for the hosted API.

How accurate is Cohere Transcribe?

Cohere reported a 5.42 average WER and first place on the Open ASR Leaderboard at launch. Accuracy can vary materially by language, accent, audio quality, noise, and industry vocabulary, so teams should evaluate their own recordings.

What is Cohere Transcribe Arabic?

It is a separate 2B model released after the general model for Arabic dialect variation, Arabic-English code-switching, and domain-specific vocabulary.

Bottom line

Our Cohere Transcribe verdict

Cohere Transcribe is a strong option for engineering teams that want accurate multilingual speech recognition without surrendering deployment control. Its open license, practical integrations, and hosted-to-dedicated deployment path are compelling, but buyers must plan around the lack of timestamps, speaker diarization, automatic language detection, and built-in silence handling.

Visit Cohere Transcribe website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.