Real-time voice agents
Low-latency streaming is suited to customer support, sales, recruiting, concierge, and other conversational applications.
Independent tool overview
Sonic-3.6 is Cartesia's real-time text-to-speech model for voice agents and applications that need natural, expressive streaming speech, broad language coverage, voice cloning, and low latency.
Visit the official Sonic-3.6 site ↗
Overview
Sonic-3.6 is the current model highlighted in Cartesia's Sonic text-to-speech product. It is designed for interactive voice agents, where the time to first playable audio, natural pacing, pronunciation, and continuity across streamed text all affect whether a conversation feels responsive.
Cartesia presents Sonic as natively multilingual across 44 languages with sub-90-millisecond model latency. The product supports automatic interpretation of emotional context, transcript-level non-verbal expressions such as laughter, instant voice cloning from a short sample, voice localization, and custom pronunciation dictionaries.
Sonic is delivered through Cartesia's hosted API and playground, with SDKs and streaming endpoints for developers. It can provide the speech layer inside a voice agent, but teams still need speech recognition, turn detection, an LLM or workflow engine, telephony, safety policies, monitoring, and fallback handling for a complete production system.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Low-latency streaming is suited to customer support, sales, recruiting, concierge, and other conversational applications.
Teams can use one voice platform across 44 supported languages and a range of regional accents.
Instant and professional cloning options can preserve an approved voice across products, markets, and use cases.
Custom dictionaries help control names, industry terms, products, medications, confirmation codes, and other difficult speech.
Capabilities
Sonic streams audio as it is generated through HTTP, server-sent events, or WebSocket workflows rather than waiting for a full recording.
The model interprets emotional subtext and pacing from the transcript and supports inserted non-verbal expressions such as laughter.
Cartesia lists native multilingual generation and localization across 44 languages, including several English, Spanish, Arabic, and Portuguese locales.
The product page says a voice can be cloned from 10 seconds of approved audio, while higher plans add professional voice-cloning options.
A source voice can be adapted to another supported language while retaining speaker identity, tone, and emotional character.
Developers can define how proper nouns, technical terms, brands, and other sensitive words should be spoken.
Cartesia provides JavaScript and Python libraries plus bytes, SSE, and WebSocket TTS interfaces for different streaming needs.
Process
Step 1
Choose languages, accents, tone, turn-taking goals, approved voices, and the maximum acceptable end-to-end response delay.
Step 2
Test representative transcripts, names, codes, emotional moments, interruptions, and noisy real-world wording before integration.
Step 3
Use the bytes endpoint for a simple streamed response, or SSE and WebSockets when timestamps, continuations, and persistent connections matter.
Step 4
Select an approved library or cloned voice and add dictionaries for important names, domain terms, and structured values.
Step 5
Track end-of-user-turn to first audible response, audio underruns, pronunciation errors, completion rate, regional latency, and cost per resolved call.
Cost
Cartesia bundles Sonic-3.6 usage into monthly credit plans. The published table translates each credit allowance into approximate TTS minutes and sets plan-specific concurrency and voice-cloning features. Enterprise terms are custom.
$0 per month
A small starting tier for testing text-to-speech and speech-to-text.
$5 per month
Adds commercial-use rights and instant voice cloning for small production workloads.
$49 per month
A higher-volume team tier with organizations and professional voice cloning.
$299 per month
For larger applications needing substantially more credits, concurrency, and support.
Custom
Custom usage, concurrency, compliance, deployment, and support terms.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Sales
Choose Vapi when you want a broader orchestration platform for assembling voice agents from speech, model, telephony, and tool providers.
Explore Vapi →Sales
Consider Bland AI for a managed phone-agent platform focused on automated calls and business workflows.
Explore Bland AI →Business Operations
Consider Speechify when the primary need is end-user text listening and content narration rather than a developer-first real-time TTS API.
Explore Speechify →Questions
Sonic-3.6 is Cartesia's real-time text-to-speech model for voice agents, applications, narration, localization, and other workflows that need expressive streaming audio.
Cartesia's current Sonic product page lists 44 languages and a range of regional accents. Teams should test the exact language and voice required for production.
Cartesia advertises sub-90-millisecond model latency. Actual end-to-end conversational delay will also include speech recognition, turn detection, LLM generation, network time, audio decoding, and buffering.
Yes. Cartesia says instant cloning can use 10 seconds of audio, and higher plans offer professional voice cloning. Only use voices with documented permission and appropriate disclosure.
Cartesia offers a Free plan with about 27 TTS minutes, Pro at $5 per month with about 133 minutes, Startup at $49 with about 1,667 minutes, and Scale at $299 with about 10,667 minutes. Enterprise pricing is custom.
No. Sonic-3.6 is the text-to-speech layer. Cartesia also offers speech-to-text and managed-agent products, but a production agent still needs orchestration, business tools, policy controls, monitoring, and often telephony.
Bottom line
Sonic-3.6 is a strong TTS candidate for teams building responsive multilingual voice agents, especially when natural delivery, cloning, localization, and pronunciation control matter together. The published entry pricing is accessible, but a production decision should be based on full-call latency, locale-specific quality, reliability, consent, and cost at the required concurrency—not a single model-speed claim or demo sample.
Visit Sonic-3.6 website ↗
Grok Bot - xAI's always-on agent teammates with their own cloud computer

Berd - Block's open-source desktop hub for running custom AI agents

Xirp - Spotify's vendor-neutral workspace for Claude, Gemini, and Codex coding agents

Search API that lets agents query and learn from the live web

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.