The Rundown AI homepage

Independent tool overview

Sonic-3.6 at a glance

Sonic-3.6 is Cartesia's real-time text-to-speech model for voice agents and applications that need natural, expressive streaming speech, broad language coverage, voice cloning, and low latency.

Visit the official Sonic-3.6 site ↗
Sonic-3.6 product preview
Product
Streaming text-to-speech API
Languages
44
Advertised model latency
Under 90 ms
Starting price
Free plan

Overview

What Sonic-3.6 is

Sonic-3.6 is the current model highlighted in Cartesia's Sonic text-to-speech product. It is designed for interactive voice agents, where the time to first playable audio, natural pacing, pronunciation, and continuity across streamed text all affect whether a conversation feels responsive.

Cartesia presents Sonic as natively multilingual across 44 languages with sub-90-millisecond model latency. The product supports automatic interpretation of emotional context, transcript-level non-verbal expressions such as laughter, instant voice cloning from a short sample, voice localization, and custom pronunciation dictionaries.

Sonic is delivered through Cartesia's hosted API and playground, with SDKs and streaming endpoints for developers. It can provide the speech layer inside a voice agent, but teams still need speech recognition, turn detection, an LLM or workflow engine, telephony, safety policies, monitoring, and fallback handling for a complete production system.

Use cases

Who Sonic-3.6 is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Real-time voice agents

Low-latency streaming is suited to customer support, sales, recruiting, concierge, and other conversational applications.

Multilingual customer experiences

Teams can use one voice platform across 44 supported languages and a range of regional accents.

Branded and cloned voices

Instant and professional cloning options can preserve an approved voice across products, markets, and use cases.

Pronunciation-sensitive workflows

Custom dictionaries help control names, industry terms, products, medications, confirmation codes, and other difficult speech.

Capabilities

Core Sonic-3.6 features

1

Real-time streaming speech

Sonic streams audio as it is generated through HTTP, server-sent events, or WebSocket workflows rather than waiting for a full recording.

2

Expressive delivery

The model interprets emotional subtext and pacing from the transcript and supports inserted non-verbal expressions such as laughter.

3

Forty-four languages

Cartesia lists native multilingual generation and localization across 44 languages, including several English, Spanish, Arabic, and Portuguese locales.

4

Instant voice cloning

The product page says a voice can be cloned from 10 seconds of approved audio, while higher plans add professional voice-cloning options.

5

Voice localization

A source voice can be adapted to another supported language while retaining speaker identity, tone, and emotional character.

6

Pronunciation dictionaries

Developers can define how proper nouns, technical terms, brands, and other sensitive words should be spoken.

7

Developer SDKs and endpoints

Cartesia provides JavaScript and Python libraries plus bytes, SSE, and WebSocket TTS interfaces for different streaming needs.

Process

How the Sonic-3.6 workflow works

  1. Step 1

    Define the voice experience

    Choose languages, accents, tone, turn-taking goals, approved voices, and the maximum acceptable end-to-end response delay.

  2. Step 2

    Prototype in the playground

    Test representative transcripts, names, codes, emotional moments, interruptions, and noisy real-world wording before integration.

  3. Step 3

    Integrate the right stream

    Use the bytes endpoint for a simple streamed response, or SSE and WebSockets when timestamps, continuations, and persistent connections matter.

  4. Step 4

    Configure pronunciation and voices

    Select an approved library or cloned voice and add dictionaries for important names, domain terms, and structured values.

  5. Step 5

    Measure the complete conversation

    Track end-of-user-turn to first audible response, audio underruns, pronunciation errors, completion rate, regional latency, and cost per resolved call.

Cost

Sonic-3.6 pricing and free plan

Cartesia bundles Sonic-3.6 usage into monthly credit plans. The published table translates each credit allowance into approximate TTS minutes and sets plan-specific concurrency and voice-cloning features. Enterprise terms are custom.

Free

$0 per month

A small starting tier for testing text-to-speech and speech-to-text.

  • 20,000 credits per month
  • Approximately 27 Sonic-3.6 minutes
  • Two concurrent TTS requests
  • No commercial-use license listed

Pro

$5 per month

Adds commercial-use rights and instant voice cloning for small production workloads.

  • 100,000 credits per month
  • Approximately 133 Sonic-3.6 minutes
  • Three concurrent TTS requests
  • Instant voice cloning

Startup

$49 per month

A higher-volume team tier with organizations and professional voice cloning.

  • 1.25 million credits per month
  • Approximately 1,667 Sonic-3.6 minutes
  • Five concurrent TTS requests
  • Two professional voice clones

Scale

$299 per month

For larger applications needing substantially more credits, concurrency, and support.

  • 8 million credits per month
  • Approximately 10,667 Sonic-3.6 minutes
  • Fifteen concurrent TTS requests
  • Four professional voice clones and priority support

Enterprise

Custom

Custom usage, concurrency, compliance, deployment, and support terms.

  • Volume pricing
  • Custom concurrency and voice-cloning capacity
  • SSO, DPAs, BAAs, security questionnaires, and shared support channel

Pricing checked . Check current pricing at the source ↗

Assessment

Sonic-3.6 strengths and limitations

Where it stands out

  • Designed around low-latency streaming rather than offline-only narration.
  • Broad 44-language coverage supports international voice-agent deployments.
  • Expressive delivery, non-verbal cues, cloning, localization, and pronunciation tools live in one platform.
  • A free plan and $5 commercial tier make initial evaluation inexpensive.
  • Multiple streaming interfaces and official SDKs support both simple and advanced integrations.

What to consider

  • The advertised sub-90-millisecond figure is model latency, not the full delay from a user finishing speech to hearing a response.
  • The Free plan does not list a commercial-use license and includes only about 27 TTS minutes per month.
  • Professional voice cloning begins on the $49-per-month Startup plan and requires appropriate speaker consent and governance.
  • Language availability does not guarantee equally strong accents, names, code-switching, or domain pronunciation in every locale.
  • A complete voice agent still requires STT, turn detection, LLM or workflow orchestration, telephony, tools, monitoring, and safeguards.
  • Model aliases and behavior can evolve, so production teams should regression-test updates and pin versions when stability matters.

Compare

Sonic-3.6 alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Sales

Vapi

Choose Vapi when you want a broader orchestration platform for assembling voice agents from speech, model, telephony, and tool providers.

Explore Vapi

Sales

Bland AI

Consider Bland AI for a managed phone-agent platform focused on automated calls and business workflows.

Explore Bland AI

Business Operations

Speechify

Consider Speechify when the primary need is end-user text listening and content narration rather than a developer-first real-time TTS API.

Explore Speechify

Questions

Sonic-3.6 FAQs

What is Cartesia Sonic-3.6?

Sonic-3.6 is Cartesia's real-time text-to-speech model for voice agents, applications, narration, localization, and other workflows that need expressive streaming audio.

How many languages does Sonic-3.6 support?

Cartesia's current Sonic product page lists 44 languages and a range of regional accents. Teams should test the exact language and voice required for production.

How fast is Sonic-3.6?

Cartesia advertises sub-90-millisecond model latency. Actual end-to-end conversational delay will also include speech recognition, turn detection, LLM generation, network time, audio decoding, and buffering.

Can Sonic-3.6 clone a voice?

Yes. Cartesia says instant cloning can use 10 seconds of audio, and higher plans offer professional voice cloning. Only use voices with documented permission and appropriate disclosure.

How much does Sonic-3.6 cost?

Cartesia offers a Free plan with about 27 TTS minutes, Pro at $5 per month with about 133 minutes, Startup at $49 with about 1,667 minutes, and Scale at $299 with about 10,667 minutes. Enterprise pricing is custom.

Is Sonic-3.6 a complete voice-agent platform?

No. Sonic-3.6 is the text-to-speech layer. Cartesia also offers speech-to-text and managed-agent products, but a production agent still needs orchestration, business tools, policy controls, monitoring, and often telephony.

Bottom line

Our Sonic-3.6 verdict

Sonic-3.6 is a strong TTS candidate for teams building responsive multilingual voice agents, especially when natural delivery, cloning, localization, and pronunciation control matter together. The published entry pricing is accessible, but a production decision should be based on full-call latency, locale-specific quality, reliability, consent, and cost at the required concurrency—not a single model-speed claim or demo sample.

Visit Sonic-3.6 website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.