Consented brand voices
Organizations with a named speaker's written permission to deploy a consistent synthetic voice in approved customer experiences.
Independent tool overview
xAI Custom Voices lets developers clone a consenting speaker's voice from a short recording, then use the resulting voice ID with Grok text-to-speech and real-time voice APIs. Console creation is available in most of the United States, while API-based creation requires an Enterprise plan.
Visit the official xAI Custom Voices site ↗
Overview
Custom Voices is an xAI developer feature, not a general voice-cloning tool inside the consumer Grok chat experience. A team records or uploads reference speech, creates a private custom voice, and passes its voice ID to Grok text-to-speech, streaming TTS, or real-time speech-to-speech endpoints.
xAI's launch announcement describes a two-stage ownership check in the console: the speaker reads a live passphrase, then the system compares speaker embeddings from that check and the longer recording. The company says this prevents cloning from an old clip or cloning somebody else's voice. That is a meaningful control, but it does not remove the developer's duty to obtain durable, use-specific consent and prevent deceptive output.
Current documentation recommends 90–120 seconds of clean, expressive, single-speaker audio, although clips can be shorter and may not exceed 120 seconds. Custom voices are private to the creating team, can be used anywhere built-in xAI voices work, and inherit multilingual output and streaming capabilities.
The existing Rundown description said Custom Voices were for use across Grok applications. The current first-party documentation is narrower and more useful: this is primarily an API and developer-console capability for TTS and voice agents. Treat every cloned voice as sensitive identity data, disclose synthetic speech, secure its source audio and voice ID, and never use it for impersonation or unsanctioned reuse.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Organizations with a named speaker's written permission to deploy a consistent synthetic voice in approved customer experiences.
Creators scaling drafts, revisions, localization, or accessible versions in their own voice while clearly labeling synthetic audio.
People preserving their own vocal identity for assistive communication, with careful long-term consent and account recovery planning.
Engineering teams that need the same custom voice across generated narration and conversational agents.
Capabilities
Creates a reusable voice model from no more than 120 seconds of reference audio.
The announced console flow checks a live spoken passphrase and similarity between that recording and the main reference clip.
Custom voices appear only to the owning team and are not added to xAI's public built-in voice list.
A custom voice ID works with REST TTS, streaming TTS, and real-time speech-to-speech endpoints.
Custom voices inherit the multilingual output capabilities available to xAI's voice APIs.
Teams can label a voice with a name, description, language, accent, age, gender, use case, and tone.
Enterprise API users can list, inspect, update metadata, download reference audio, and delete a voice.
xAI says using a custom voice carries the same TTS or Voice Agent usage charge as a built-in voice.
Process
Step 1
Define where the voice will appear, what it may say, who can use it, how long permission lasts, and how consent can be withdrawn.
Step 2
Get permission from the actual speaker for cloning, storage, generation, distribution, languages, commercial contexts, and any downstream customer use.
Step 3
Capture 90–120 seconds of natural, expressive, single-speaker audio in a quiet, soft room with no music or private conversation.
Step 4
Use the console passphrase and speaker-similarity flow; Enterprise teams should confirm in writing how verification applies to API-created voices.
Step 5
Keep API keys, console access, reference audio, and voice IDs limited to people who need them, with logging and rapid revocation.
Step 6
Review pronunciation, emotional tone, accents, multilingual output, names, numbers, and high-risk phrases before launch.
Step 7
Tell listeners they are hearing an AI-generated or cloned voice and provide a clear route to a human for consequential interactions.
Step 8
Prevent identity verification, payment authorization, political persuasion, emergency instructions, deceptive endorsements, and unsupervised high-stakes advice.
Step 9
Audit generated scripts and distribution, respond to misuse, honor consent withdrawal, delete the voice and source audio, and remove deployed assets where possible.
Cost
Creating and using a custom voice does not add a separate cloning surcharge, but generated audio is billed at the applicable xAI Voice API rate. Console users can create up to 30 custom voices; direct API creation is restricted to Enterprise teams. Prices below were checked in xAI's current API documentation.
No extra creation charge
Create up to 30 custom voices in the xAI console, subject to availability and account access.
$15 per 1M characters
Usage-based synthesis through xAI's REST or streaming TTS endpoints.
$0.08/minute audio
Current pricing for grok-voice-think-fast-2.0, equivalent to $4.80 per audio hour.
Contact sales
Direct creation through the custom-voices API requires an Enterprise-enabled team.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Marketing
Choose ElevenLabs for a broader mature voice platform with multiple cloning tiers, creator tooling, dubbing, and a large ecosystem.
Explore ElevenLabs →Consumer
Choose Qwen3-TTS CustomVoice when model access and self-managed experimentation matter more than a hosted, verified cloning workflow.
Explore Qwen3-TTS CustomVoice 1.7B →Content Creator
Choose ElevenLabs Voice Design V3 when a new synthetic character voice is safer and more appropriate than copying a real person's identity.
Explore ElevenLabs Voice Design V3 →Questions
It is an xAI developer feature that creates a private reusable voice from a short reference recording for use in Grok text-to-speech and real-time voice APIs.
The current official documentation presents it as an xAI console and API capability for developers, not a general consumer cloning button across Grok chat.
As of September 1, 2026, xAI says Custom Voices is available only in the United States, excluding Illinois.
The maximum is 120 seconds. xAI accepts shorter clips but recommends 90–120 seconds; clips under 30 seconds may capture less detail.
No. xAI says the console requires a live passphrase and speaker-similarity check and that users cannot clone another person's voice. You also need explicit legal permission for the full planned use.
xAI says there is no extra custom-voice surcharge. Current usage is $15 per 1 million TTS characters or $0.08 per speech-to-speech audio minute for grok-voice-think-fast-2.0, with possible additional charges.
The default limit is 30 custom voices per team. xAI asks larger deployments to request a higher limit.
Only an Enterprise-enabled team can call the creation endpoint directly. Other eligible users create voices in the xAI console and then copy the voice ID.
No. xAI says they are scoped to the creating team and never appear in the built-in voice list available to other users.
xAI says custom voices inherit multilingual TTS capabilities. Test every target language with the speaker because pronunciation, accent, meaning, and implied endorsement can change.
Yes whenever a listener could reasonably think the person recorded or approved the exact message. Disclosure is also a core fraud-prevention control, and specific laws or platform rules may require it.
Technically, it can power a voice agent, but U.S. AI-generated calls are subject to FCC artificial or prerecorded voice rules. Obtain appropriate consent, provide required identification, disclose automation, and get legal review for telemarketing or scaled calling.
xAI says deletion removes the voice and its underlying reference audio and causes future calls with that voice ID to fail. It does not erase audio already exported or published outside xAI.
xAI's API security FAQ says it does not train on API inputs or outputs without explicit permission. API requests and responses are stored for 30 days by default for abuse auditing, with separate Zero Data Retention options for eligible teams.
Bottom line
xAI Custom Voices is a capable developer primitive with unusually clear console ownership checks, simple TTS pricing, and one voice ID across narration and live agents. The page's original description understated the product's API focus. Its safety depends on the deployment: obtain written and revocable consent, resolve the Enterprise API verification discrepancy, disclose synthetic speech, secure the source audio and credentials, and block every use where a listener could be deceived or harmed.
Visit xAI Custom Voices website ↗
Deep Max - Exa's new SOTA agentic search tool

Realtime TTS-2 - Inworld AI's new voice model that hears conversation audio to match user tone and emotion

HY-World 2.0 - Tencent's open-source world model that turns text, images, or video into interactive 3D scenes

GPT-Realtime-2 - OpenAI's new voice model that can think, call tools, and recover from interruptions in live calls

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.