Expressive voice-agent research
Prototype conversational agents where pauses, emphasis, emotion, and turn context matter more than simple narration.
Independent tool overview
Miso One is the launch name associated with Miso Labs' current open release, Miso TTS 8B. It is an English text-to-speech model for expressive conversational speech, dialogue, and voice continuation from optional prompt audio. The weights and local inference code are public under a modified MIT license, but Miso Labs has not launched or priced its promised production API. Running it yourself requires substantial hardware: the official repository recommends at least 24 GB of GPU memory for half-precision inference.
Visit the official Miso One site ↗
Overview
Miso TTS 8B turns text and optional audio context into conversational speech. Its architecture follows Sesame-style conversational speech modeling: a large Llama-style backbone predicts the first Mimi audio codebook and a smaller autoregressive decoder fills in the remaining codebooks. That design is intended to preserve timing, tone, and dialogue context rather than produce a flat sentence in isolation.
The practical product today is the downloadable model, inference repository, and Miso Labs demo. Developers can run the model locally, condition it with a transcript and a short reference recording, or use earlier generated speech as context for a continued exchange. This makes Miso most relevant to technical teams prototyping expressive agents, characters, and private voice systems—not buyers looking for a polished no-code voice platform.
The headline tradeoff is infrastructure. Miso is unusually open for an 8B-parameter voice model, but the official setup downloads roughly 30–40 GB and recommends a 24 GB GPU. Miso's advertised 110 ms figure refers to time to first audio on H100-class hosted infrastructure, not ordinary local hardware or complete response latency. There is no public API price, service-level commitment, or production usage allowance to compare yet.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Prototype conversational agents where pauses, emphasis, emotion, and turn context matter more than simple narration.
Keep model execution and sensitive voice data on infrastructure you control, subject to the model license and your own security practices.
Condition generation on transcript-plus-audio context or a short reference clip to continue a voice and conversational style.
Inspect the inference code and model architecture instead of relying exclusively on a closed hosted voice API.
Use the model when you already have a 24 GB or larger GPU, audio engineering expertise, and the capacity to build a production layer around raw inference.
Capabilities
Generates English speech with dialogue context, timing, and vocal character rather than treating each line as an unrelated narration clip.
Accepts prompt audio with a matching transcript for voice continuation and one-shot voice-cloning workflows.
Combines an 8B Llama-style backbone with a roughly 300M-parameter audio decoder to predict 32 Mimi audio codebooks.
Public code and model weights let technical teams run the system on their own GPU or private infrastructure.
Context segments can identify speakers and include prior text and audio so a later response follows the preceding exchange.
The reference pipeline applies Sony SilentCipher watermarking by default; production deployments should configure a private watermark key.
Miso Labs provides a browser demo for evaluating the model without first assembling a local environment.
Process
Step 1
Test representative English scripts, emotional directions, names, punctuation, and dialogue turns before committing infrastructure.
Step 2
Review the modified MIT terms, including the Miso Labs attribution requirement for exceptionally large commercial products, with counsel when appropriate.
Step 3
Plan for roughly 30–40 GB of initial downloads and at least 24 GB of VRAM for the recommended half-precision setup; CPU inference is possible but slow.
Step 4
Use the published environment and inference instructions, pin model and dependency versions, and isolate the service from public uploads during evaluation.
Step 5
Only clone voices when you have permission, retain provenance for every reference clip, and block impersonation or deceptive use.
Step 6
Measure time to first audio, real-time factor, total response delay, GPU utilization, concurrency, quality, and failure rates on your own hardware.
Step 7
Add authentication, quotas, moderation, watermark verification, logging, monitoring, retries, data retention rules, and human escalation around the raw model.
Cost
Miso TTS 8B model weights and inference code can be downloaded without a purchase price under a modified MIT license, but self-hosting compute, storage, engineering, and operations are separate costs. The public demo is available for evaluation. Miso Labs still describes API access as coming soon and has not published usage pricing; enterprise on-premises hosting and support are available by request.
Free download
Public Miso TTS 8B weights and inference code for local use under the published modified MIT license.
Free to try
Browser-based evaluation of Miso's speech output without setting up the model locally.
Coming soon
Miso Labs has announced future API access but has not published a launch date, rates, quotas, or service levels.
Custom
Hosting assistance and support contracts for organizations deploying the model in their own environment.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Business Operations
Choose PlayAI when a managed conversational voice platform and production APIs matter more than running open weights yourself.
Explore PlayAI →Content Creator
Consider Seed Audio 1.0 when you need a broader audio creation model that can generate speech, music, ambience, and effects in one scene.
Explore Seed Audio 1.0 →Content Creator
Use the ElevenLabs Text to Speech GPT for a simpler hosted workflow inside ChatGPT rather than managing an 8B model locally.
Explore ElevenLabs Text To Speech GPT →Questions
Miso One is the launch name used for Miso Labs' expressive voice model. The current official open release is labeled Miso TTS 8B.
The weights and inference code are free to download under a modified MIT license. GPU compute, storage, engineering, hosting, and operations are not included.
Miso Labs recommends at least 24 GB of VRAM for bf16 or fp16 inference. The model can run on a CPU with substantial RAM, but the official documentation warns that it will be slow.
Yes. Miso accepts optional transcript-and-audio context and promotes one-shot cloning from a short reference recording. Only use a voice with the speaker's informed permission.
No multilingual support is documented in the current official repository; it identifies the model language as English.
Not as a generally available, publicly priced product. Miso Labs says API access is coming soon, while the current usable options are the demo, local model, or a custom enterprise arrangement.
Miso Labs reports roughly 110 ms time to first audio for its hosted production setup on H100-class hardware. That is not full response latency and should not be assumed for local GPUs.
The modified MIT license permits broad commercial use, but products above 50 million monthly active users or $10 million in monthly revenue must prominently display Miso Labs in the user interface. Obtain legal advice for your use case.
Bottom line
Miso TTS 8B is compelling for experienced teams that want expressive English speech, voice conditioning, and the freedom to run an open model locally. It is not yet a turnkey replacement for a mature hosted voice platform: hardware demands are high, the API remains unpriced and unreleased, and buyers must supply the production, safety, and consent layers. Start with the demo, benchmark on the exact deployment hardware, and choose Miso primarily when model control and on-premises operation justify the engineering cost.
Visit Miso One website ↗
Autoscientist - Adaption's new tool for automating AI model training

Gemini 3.5 Live Translate - Google's real-time voice model for live translation across 70+ languages

GPT-Realtime-2 - OpenAI's new voice model that can think, call tools, and recover from interruptions in live calls

DiffusionGemma - Google's open diffusion model that can quadruple text generation speed

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.