Long-form speech research
Experiment with extended single- or multi-speaker narration where turn-taking and voice consistency matter across a long script.
Independent tool overview
VibeVoice is Microsoft's open-source research family for long-form speech generation and recognition, including separate multi-speaker and real-time TTS models.
Visit the official VibeVoice site ↗
Overview
VibeVoice is a Microsoft Research model family, not a single hosted voice app. For text-to-speech, the two most relevant releases solve different problems: VibeVoice-1.5B can generate long-form conversations up to about 90 minutes with as many as four speakers, while VibeVoice-Realtime-0.5B is a lighter single-speaker model designed to begin audible output in roughly 300 milliseconds.
The distinction matters. The 1.5B model is aimed at research on podcasts and extended dialogue, while the realtime model accepts streaming text and targets live narration for about 10 minutes. Microsoft publishes both under the MIT license but labels them research releases, warns against commercial or real-world deployment without further testing, and places explicit restrictions around impersonation, disinformation and other unsafe uses.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Experiment with extended single- or multi-speaker narration where turn-taking and voice consistency matter across a long script.
Turn a structured dialogue into conversational audio with as many as four assigned speakers using the 1.5B model.
Begin speaking while text is still arriving from an LLM or live data source with the 0.5B realtime model.
Inspect and adapt an open model rather than sending scripts to a closed hosted speech API.
Compare long-context diffusion-based TTS with commercial APIs under a controlled, consent-based research protocol.
Capabilities
VibeVoice-1.5B was trained for extended sequences and can generate roughly 90 minutes of speech in one pass.
The long-form model supports as many as four distinct speakers with conversational turn-taking.
VibeVoice-Realtime-0.5B incrementally encodes incoming text while continuing acoustic generation from previous context.
Microsoft reports about 300 milliseconds to first audible speech for the realtime model, with actual latency dependent on hardware and implementation.
The 0.5B streaming model targets robust single-speaker generation for approximately 10 minutes.
VibeVoice uses an acoustic tokenizer operating at 7.5 Hz to reduce the sequence length needed for extended speech.
A language model handles text and dialogue context while a diffusion head produces acoustic detail.
The published Hugging Face models can be loaded through supported Transformers text-to-speech classes and pipelines.
Microsoft distributes model weights on Hugging Face and source material through the public VibeVoice repository.
Microsoft's model cards describe an audible AI disclaimer and imperceptible watermark added to generated audio.
Process
Step 1
Use 1.5B for long multi-speaker scripts or Realtime 0.5B for low-latency single-speaker streaming.
Step 2
Review the model card before implementation and exclude impersonation, deception, unsupported voice conversion and other prohibited scenarios.
Step 3
Install the repository and supported Transformers dependencies, then provision compatible local or cloud compute.
Step 4
Resolve abbreviations, code, formulas, special symbols and ambiguous names before synthesis to reduce unpredictable speech.
Step 5
Start with short passages, compare pronunciation and speaker consistency, then expand only after the output is reliable for the chosen hardware.
Step 6
Listen to the final audio, verify the transcript, retain the model's transparency measures and clearly disclose AI-generated speech to the audience.
Cost
Microsoft publishes VibeVoice's code and model weights under the MIT license with no model-access fee. It is self-hosted research software, so the real cost is compute, storage, engineering, monitoring and safety review rather than a per-character subscription.
Free, self-hosted
Open long-form multi-speaker research model.
Free, self-hosted
Lightweight streaming single-speaker research model.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Miscellaneous
Choose Gemini 3.1 Flash TTS for a managed API with streaming, multi-speaker output, expressive tags and straightforward usage pricing.
Explore Gemini 3.1 Flash TTS →Marketing
Use ElevenLabs for a polished hosted platform with creator tools, deployment options and commercial support.
Explore ElevenLabs →Content Creator
Compare Qwen3-TTS when evaluating another open-source speech-model family with self-hosting flexibility.
Explore Qwen3-TTS →Content Creator
Consider Inworld Realtime TTS for managed low-latency character and interactive application speech.
Explore Inworld Realtime TTS →Questions
VibeVoice is a Microsoft Research family of open-source speech models. Its text-to-speech releases include a 1.5B long-form multi-speaker model and a separate 0.5B realtime single-speaker model.
The code and published model weights use the MIT license and do not have a model-access fee. You still pay for the hardware, hosting, storage and engineering needed to run them.
Microsoft lists approximately 90 minutes as the generation length for VibeVoice-1.5B. That is the long-form model, not the smaller realtime release, and the final audio still needs careful quality review.
VibeVoice-1.5B supports up to four speakers. VibeVoice-Realtime-0.5B supports only one speaker.
Microsoft reports roughly 300 milliseconds to first audible speech. Actual latency depends on the GPU, software stack, prompt and network architecture.
The 1.5B model card supports English and Chinese. The realtime model is intended for English, with nine additional experimental languages that Microsoft says may produce unpredictable results.
The repository uses an MIT license, but Microsoft's model cards say the releases are intended for research and do not recommend commercial or real-world use without further testing and development. Legal, safety and operational review is essential.
The current materials present VibeVoice TTS as models to run yourself. Microsoft does not list a managed pay-as-you-go VibeVoice TTS endpoint, so teams should plan for their own deployment or choose a hosted alternative.
Bottom line
VibeVoice is technically interesting because it tackles two difficult ends of TTS: very long multi-speaker audio and low-latency streaming. It is a strong research option for teams that can self-host and enforce strict consent and disclosure rules. It is not the simplest path to shipping a customer-facing voice feature; a managed API is usually the better choice when reliability, support, moderation and predictable production operations matter more than model access.
Visit VibeVoice website ↗
Kling 2.6 - Kling's new AI video model with native audio capabilities

Suno AI: Turns text prompts into studio-quality songs using cutting-edge music generation models.

Seedream 4.5 - ByteDance's upgraded AI image model with powerful editing, text rendering, typography, and realism

Kling AI: Enables you to generate imaginative videos and images with cutting-edge generative models.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.