Real-time voice agents
Use streaming recognition, built-in turn detection, and low-latency speech generation for natural back-and-forth conversations.
Independent tool overview
Sonic 3.6 and Ink 2 are Cartesia's real-time text-to-speech and speech-to-text models for building low-latency voice agents through one API platform.
Visit the official Sonic 3.6 & Ink 2 site ↗
Overview
Cartesia pairs Sonic 3.6 for text-to-speech with Ink 2 for streaming speech recognition. Together they cover the audio input and output layers of a real-time voice agent without requiring separate speech vendors.
This listing originally covered Sonic 3.5 and Ink 2. Sonic 3.6 is now the generally available TTS model and is fully backward-compatible with 3.5, so new projects should evaluate the current model ID, sonic-3.6.
Sonic 3.6 supports 44 languages, voice cloning, streaming, locale-aware text normalization, and stable dated snapshots. Ink 2 is currently English-only and focuses on low-latency transcription, structured data, and built-in turn detection.
Cartesia advertises sub-90ms TTS and 100ms transcript latency for the combined real-time stack. Those are vendor claims; production latency will also depend on geography, network conditions, buffering, audio transport, the language model, and application logic.
Pricing uses shared monthly credits. TTS is roughly one credit per input character, while Ink 2 costs three credits per second of audio, including silence. Teams should test realistic calls and set overage controls before forecasting costs from the headline plan price.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Use streaming recognition, built-in turn detection, and low-latency speech generation for natural back-and-forth conversations.
Handle confirmation codes, dates, phone numbers, emails, names, and conversational pacing in support-style calls.
Generate speech across 44 languages with localized pronunciation and locale-aware rendering.
Connect both speech recognition and synthesis through one provider, account, credit pool, and integration surface.
Create instant or professional voice clones for approved speakers and supported use cases.
Build phone, web, mobile, or avatar experiences where transcript and audio delays directly affect conversation quality.
Capabilities
Streams expressive synthesized speech through bytes, SSE, or WebSocket endpoints.
Covers major European and Asian languages plus Odia and Urdu added in Sonic 3.6.
Adjusts pacing and intonation from the transcript and handles disfluencies such as 'uhm' and 'hmm' naturally.
Reads confirmation codes, phone numbers, dates, emails, and heteronyms with less manual preprocessing.
Uses locale codes such as en-GB to interpret dates and other localized formats correctly.
Supports instant cloning and plan-dependent professional voice cloning, with older professional clones working on Sonic 3.6.
Transcribes English audio in real time for voice-agent and push-to-talk workflows.
Emits start, update, eager-end, resume, and end events without requiring a separate voice-activity detector.
Lets production teams pin dated Sonic releases while testing upcoming behavior through a preview alias.
Provides Python and JavaScript libraries plus documented integrations with voice-agent and telephony frameworks.
Offers zero data retention for eligible inference requests and dedicated regional deployments on Enterprise.
Process
Step 1
Decide whether the application needs Sonic TTS, Ink STT, or both; confirm that English-only recognition fits the input audience.
Step 2
Collect real transcripts and audio covering accents, background noise, names, numbers, interruptions, silence, and difficult pronunciations.
Step 3
Keep permanent credentials off clients and use Cartesia's client authentication approach where browser or mobile access is required.
Step 4
Stream microphone audio to Ink 2 and route agent text back through Sonic 3.6 using the appropriate WebSocket or streaming endpoint.
Step 5
Use Ink 2's turn lifecycle events to decide when the agent should listen, prepare a response, interrupt, or resume.
Step 6
Compare several voices with real scripts and verify expressiveness, pronunciation, pace, and consistency across target languages.
Step 7
Use a dated Sonic snapshot when repeatable behavior matters, and evaluate stable alias changes before rolling them into critical workflows.
Step 8
Track time to transcript, turn-end accuracy, model response time, time to first audio, interruptions, word error rate, and call completion.
Step 9
Monitor credits, decide whether to enable overages, set concurrency expectations, and evaluate Enterprise retention or regional requirements.
Cost
Every plan includes a shared monthly credit balance. Approximate TTS minutes and STT hours assume the entire allowance is used on that one capability, so a production stack using both will split the credits. Sonic TTS costs roughly one credit per character; Ink 2 realtime STT costs three credits per second, including silence.
$0 per month
For prototypes and light evaluation with 20,000 monthly credits.
$5 per month
Entry paid plan with commercial-use rights and 100,000 monthly credits.
$49 per month
For growing applications with 1.25 million monthly credits and organization features.
$299 per month
For higher-volume production with 8 million monthly credits.
Custom
For negotiated volume, security, compliance, and deployment requirements.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Marketing
Consider ElevenLabs for a broad voice platform spanning expressive TTS, cloning, dubbing, and conversational agents.
Explore ElevenLabs →Content Creator
Choose AssemblyAI when speech recognition and audio intelligence matter more than sourcing TTS and STT from one vendor.
Explore AssemblyAI →Miscellaneous
Consider TADA when an open-source TTS model and tight text-audio alignment are more important than a managed full-stack voice API.
Explore TADA by Hume AI →Questions
Sonic 3.6 is Cartesia's current text-to-speech model, while Ink 2 is its streaming speech-to-text model with built-in turn detection. They can form the speech layer of a real-time voice agent.
Sonic 3.6 is now generally available and Cartesia says it is fully backward-compatible with Sonic 3.5. New projects should evaluate sonic-3.6, while existing production teams can migrate with their normal regression tests.
Sonic 3.6 supports 44 languages. It added Odia and Urdu and also includes locale codes for localized handling of formats such as dates.
Ink 2 is currently listed as English-only. Teams needing multilingual transcription should evaluate another Cartesia STT model or a different provider.
Plans start at $0, followed by Pro at $5 per month, Startup at $49, Scale at $299, and custom Enterprise pricing. Each plan includes credits shared across model usage.
Standard Sonic TTS is approximately one credit per input character. Realtime Ink 2 costs three credits per second of audio, including silence. Only successful requests consume credits.
No. Its turn-detection endpoint emits a lifecycle of turn events so the application can identify when a speaker starts, pauses, resumes, and ends.
Use the stable sonic-3.6 alias if automatic stable updates are acceptable. Pin a dated snapshot when you need repeatable behavior and want to evaluate every model change before rollout.
Enterprise customers can enable zero data retention for eligible TTS and STT inference payloads. It does not cover voice cloning workflows, and operational metadata is still retained.
Bottom line
Cartesia is a strong fit for teams that want a tightly integrated, low-latency voice stack and can work within Ink 2's English-only recognition. Sonic 3.6 offers broad multilingual output and practical production versioning, while Ink 2's turn events simplify conversational orchestration. The deciding test should be an end-to-end pilot using real callers, realistic silence, noisy audio, and a full cost model—not the model latency claim alone.
Visit Sonic 3.6 & Ink 2 website ↗
DiffusionGemma - Google's open diffusion model that can quadruple text generation speed

Antares - Cisco's small, open security models that scan code locally

Gemini 3.5 Live Translate - Google's real-time voice model for live translation across 70+ languages
.jpeg)
Lyria 3.5 - Google's new music model with more realistic vocals

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.