The Rundown AI homepage

Independent tool overview

Sonic 3.6 & Ink 2 at a glance

Sonic 3.6 and Ink 2 are Cartesia's real-time text-to-speech and speech-to-text models for building low-latency voice agents through one API platform.

Visit the official Sonic 3.6 & Ink 2 site ↗
Sonic 3.6 & Ink 2 product preview
Product type
Real-time TTS and STT APIs
Developer
Cartesia
Current TTS model
Sonic 3.6
Current STT model
Ink 2
Sonic languages
44
Ink 2 languages
English
Starting price
Free
Paid plans
From $5 per month
Last reviewed
August 29, 2026

Overview

What Sonic 3.6 & Ink 2 is

Cartesia pairs Sonic 3.6 for text-to-speech with Ink 2 for streaming speech recognition. Together they cover the audio input and output layers of a real-time voice agent without requiring separate speech vendors.

This listing originally covered Sonic 3.5 and Ink 2. Sonic 3.6 is now the generally available TTS model and is fully backward-compatible with 3.5, so new projects should evaluate the current model ID, sonic-3.6.

Sonic 3.6 supports 44 languages, voice cloning, streaming, locale-aware text normalization, and stable dated snapshots. Ink 2 is currently English-only and focuses on low-latency transcription, structured data, and built-in turn detection.

Cartesia advertises sub-90ms TTS and 100ms transcript latency for the combined real-time stack. Those are vendor claims; production latency will also depend on geography, network conditions, buffering, audio transport, the language model, and application logic.

Pricing uses shared monthly credits. TTS is roughly one credit per input character, while Ink 2 costs three credits per second of audio, including silence. Teams should test realistic calls and set overage controls before forecasting costs from the headline plan price.

Use cases

Who Sonic 3.6 & Ink 2 is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Real-time voice agents

Use streaming recognition, built-in turn detection, and low-latency speech generation for natural back-and-forth conversations.

Customer-support automation

Handle confirmation codes, dates, phone numbers, emails, names, and conversational pacing in support-style calls.

Multilingual speech output

Generate speech across 44 languages with localized pronunciation and locale-aware rendering.

Developers consolidating vendors

Connect both speech recognition and synthesis through one provider, account, credit pool, and integration surface.

Voice cloning projects

Create instant or professional voice clones for approved speakers and supported use cases.

Latency-sensitive applications

Build phone, web, mobile, or avatar experiences where transcript and audio delays directly affect conversation quality.

Capabilities

Core Sonic 3.6 & Ink 2 features

1

Sonic 3.6 streaming TTS

Streams expressive synthesized speech through bytes, SSE, or WebSocket endpoints.

2

44-language speech generation

Covers major European and Asian languages plus Odia and Urdu added in Sonic 3.6.

3

Context-aware delivery

Adjusts pacing and intonation from the transcript and handles disfluencies such as 'uhm' and 'hmm' naturally.

4

Structured-text pronunciation

Reads confirmation codes, phone numbers, dates, emails, and heteronyms with less manual preprocessing.

5

Locale support

Uses locale codes such as en-GB to interpret dates and other localized formats correctly.

6

Voice cloning

Supports instant cloning and plan-dependent professional voice cloning, with older professional clones working on Sonic 3.6.

7

Ink 2 streaming STT

Transcribes English audio in real time for voice-agent and push-to-talk workflows.

8

Native turn detection

Emits start, update, eager-end, resume, and end events without requiring a separate voice-activity detector.

9

Stable model snapshots

Lets production teams pin dated Sonic releases while testing upcoming behavior through a preview alias.

10

Official SDKs and integrations

Provides Python and JavaScript libraries plus documented integrations with voice-agent and telephony frameworks.

11

Enterprise data controls

Offers zero data retention for eligible inference requests and dedicated regional deployments on Enterprise.

Process

How the Sonic 3.6 & Ink 2 workflow works

  1. Step 1

    Choose the speech layers

    Decide whether the application needs Sonic TTS, Ink STT, or both; confirm that English-only recognition fits the input audience.

  2. Step 2

    Start with a representative test set

    Collect real transcripts and audio covering accents, background noise, names, numbers, interruptions, silence, and difficult pronunciations.

  3. Step 3

    Create a server-side API key

    Keep permanent credentials off clients and use Cartesia's client authentication approach where browser or mobile access is required.

  4. Step 4

    Integrate realtime audio

    Stream microphone audio to Ink 2 and route agent text back through Sonic 3.6 using the appropriate WebSocket or streaming endpoint.

  5. Step 5

    Tune turn handling

    Use Ink 2's turn lifecycle events to decide when the agent should listen, prepare a response, interrupt, or resume.

  6. Step 6

    Select and test voices

    Compare several voices with real scripts and verify expressiveness, pronunciation, pace, and consistency across target languages.

  7. Step 7

    Pin production versions

    Use a dated Sonic snapshot when repeatable behavior matters, and evaluate stable alias changes before rolling them into critical workflows.

  8. Step 8

    Measure end-to-end performance

    Track time to transcript, turn-end accuracy, model response time, time to first audio, interruptions, word error rate, and call completion.

  9. Step 9

    Control spend and data

    Monitor credits, decide whether to enable overages, set concurrency expectations, and evaluate Enterprise retention or regional requirements.

Cost

Sonic 3.6 & Ink 2 pricing and free plan

Every plan includes a shared monthly credit balance. Approximate TTS minutes and STT hours assume the entire allowance is used on that one capability, so a production stack using both will split the credits. Sonic TTS costs roughly one credit per character; Ink 2 realtime STT costs three credits per second, including silence.

Free

$0 per month

For prototypes and light evaluation with 20,000 monthly credits.

  • About 27 Sonic 3.6 minutes if all credits go to TTS
  • About 1 hour 51 minutes of Ink 2 if all credits go to STT
  • 2 concurrent TTS requests
  • 8 concurrent STT requests

Pro

$5 per month

Entry paid plan with commercial-use rights and 100,000 monthly credits.

  • About 133 TTS minutes or 9 hours 16 minutes of Ink 2
  • 3 concurrent TTS requests
  • 12 concurrent STT requests
  • Instant voice cloning

Startup

$49 per month

For growing applications with 1.25 million monthly credits and organization features.

  • About 1,667 TTS minutes or 115 hours 44 minutes of Ink 2
  • 5 concurrent TTS requests
  • 20 concurrent STT requests
  • Professional voice cloning

Scale

$299 per month

For higher-volume production with 8 million monthly credits.

  • About 10,667 TTS minutes or 740 hours 44 minutes of Ink 2
  • 15 concurrent TTS requests
  • 60 concurrent STT requests
  • Priority support and higher concurrency

Enterprise

Custom

For negotiated volume, security, compliance, and deployment requirements.

  • Custom credits and concurrency
  • DPAs and BAAs
  • SSO and security reviews
  • Zero data retention and regional deployments available

Pricing checked . Check current pricing at the source ↗

Assessment

Sonic 3.6 & Ink 2 strengths and limitations

Where it stands out

  • Covers both speech input and output through one platform
  • Designed specifically for low-latency, interruptible voice conversations
  • Sonic 3.6 supports 44 languages and locale-aware output
  • Ink 2 includes native turn detection rather than requiring a separate VAD
  • Handles common structured strings such as phone numbers, dates, emails, and confirmation codes
  • Dated Sonic snapshots give production teams a stable versioning option
  • Free plan makes realistic prototyping possible before committing
  • Official SDKs, realtime endpoints, and integrations support common agent stacks
  • Enterprise customers can request zero data retention and regional inference

What to consider

  • Ink 2 currently supports English only even though Sonic 3.6 can speak 44 languages
  • Sonic 3.5 is no longer the newest TTS model, so old comparisons and model IDs can become stale
  • Headline latency claims do not include the full network, LLM, application, or telephony path
  • Credit pricing makes per-call cost less obvious than a single per-minute rate
  • Ink 2 bills every second of streamed audio, including silence
  • Approximate included minutes and hours cannot both be consumed from the same shared credits
  • Realtime Ink 2 is not yet available through the batch transcription endpoint
  • Voice quality and transcription accuracy still vary by speaker, accent, audio equipment, noise, and script
  • Production teams need their own evaluation set rather than relying only on public leaderboards
  • Voice cloning requires verified consent and controls against impersonation or deceptive use
  • Zero data retention and dedicated regional deployments require Enterprise
  • Overages can create unplanned spend if usage monitoring and limits are not configured

Compare

Sonic 3.6 & Ink 2 alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Marketing

ElevenLabs

Consider ElevenLabs for a broad voice platform spanning expressive TTS, cloning, dubbing, and conversational agents.

Explore ElevenLabs

Content Creator

AssemblyAI

Choose AssemblyAI when speech recognition and audio intelligence matter more than sourcing TTS and STT from one vendor.

Explore AssemblyAI

Miscellaneous

TADA by Hume AI

Consider TADA when an open-source TTS model and tight text-audio alignment are more important than a managed full-stack voice API.

Explore TADA by Hume AI

Questions

Sonic 3.6 & Ink 2 FAQs

What are Sonic 3.6 and Ink 2?

Sonic 3.6 is Cartesia's current text-to-speech model, while Ink 2 is its streaming speech-to-text model with built-in turn detection. They can form the speech layer of a real-time voice agent.

What happened to Sonic 3.5?

Sonic 3.6 is now generally available and Cartesia says it is fully backward-compatible with Sonic 3.5. New projects should evaluate sonic-3.6, while existing production teams can migrate with their normal regression tests.

How many languages does Sonic 3.6 support?

Sonic 3.6 supports 44 languages. It added Odia and Urdu and also includes locale codes for localized handling of formats such as dates.

How many languages does Ink 2 support?

Ink 2 is currently listed as English-only. Teams needing multilingual transcription should evaluate another Cartesia STT model or a different provider.

How much does Cartesia cost?

Plans start at $0, followed by Pro at $5 per month, Startup at $49, Scale at $299, and custom Enterprise pricing. Each plan includes credits shared across model usage.

How are Sonic and Ink usage billed?

Standard Sonic TTS is approximately one credit per input character. Realtime Ink 2 costs three credits per second of audio, including silence. Only successful requests consume credits.

Does Ink 2 require separate voice activity detection?

No. Its turn-detection endpoint emits a lifecycle of turn events so the application can identify when a speaker starts, pauses, resumes, and ends.

Should production apps use a stable alias or dated Sonic snapshot?

Use the stable sonic-3.6 alias if automatic stable updates are acceptable. Pin a dated snapshot when you need repeatable behavior and want to evaluate every model change before rollout.

Does Cartesia offer zero data retention?

Enterprise customers can enable zero data retention for eligible TTS and STT inference payloads. It does not cover voice cloning workflows, and operational metadata is still retained.

Bottom line

Our Sonic 3.6 & Ink 2 verdict

Cartesia is a strong fit for teams that want a tightly integrated, low-latency voice stack and can work within Ink 2's English-only recognition. Sonic 3.6 offers broad multilingual output and practical production versioning, while Ink 2's turn events simplify conversational orchestration. The deciding test should be an end-to-end pilot using real callers, realistic silence, noisy audio, and a full cost model—not the model latency claim alone.

Visit Sonic 3.6 & Ink 2 website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.