Customer-support voice agents
Handling routine questions, looking up approved account information, and preparing or completing narrowly authorized support steps with clear human escalation.
Independent tool overview
Grok Voice Think Fast 2.0 is SpaceXAI's flagship realtime speech-to-speech model for voice assistants, phone agents, and interactive applications. It streams audio and text over WebSocket, supports reasoning while speaking, interruption handling, multiple languages and voices, and tool calls to search, document collections, MCP servers, or developer functions. Its $0.08-per-audio-minute price is simple, but production voice automation still requires consent, identity verification, narrow tool authority, latency and call-quality testing, human escalation, and rigorous monitoring.
Visit the official Grok Voice Think Fast 2.0 site ↗
Overview
SpaceXAI released Grok Voice Think Fast 2.0 on July 29, 2026 as the successor to Think Fast 1.0. The versioned API model is `grok-voice-think-fast-2.0`, and the `grok-voice-latest` alias began routing to it on August 5, 2026.
The model runs through the realtime Speech to Speech API at `wss://api.x.ai/v1/realtime`. Applications stream user audio or text to the session and receive incremental audio, transcript, tool-call, and lifecycle events in return.
SpaceXAI says the model reasons in parallel with speech, speaks in shorter conversational turns, asks one question at a time, and improves transcription under noise and telephony compression. The launch's benchmark and Starlink conversion claims are vendor-published or vendor-selected results and should be reproduced on the exact callers, accents, networks, tools, and business flow being deployed.
Sessions can use web search, X search, file or collection search, remote MCP servers, and custom functions. That enables live support, scheduling, account lookup, and guided workflows, but any tool that sends, books, refunds, cancels, changes records, or exposes private data needs separate authorization and confirmation controls.
The current API price for Think Fast 2.0 is $0.08 per minute of audio sent or received, plus $0.004 for a text input event. The direct model documentation lists 10 concurrent sessions per team, a 120-minute maximum session, and the us-east-1 cluster.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Handling routine questions, looking up approved account information, and preparing or completing narrowly authorized support steps with clear human escalation.
Building inbound SIP experiences that converse, accept DTMF input, call business tools, and transfer to a person when the policy or confidence threshold requires it.
Adding low-latency conversational help to browser, mobile, kiosk, or device experiences with short-lived client authentication.
Supporting the 20-plus documented language variants while testing dialect, code-switching, names, numbers, and domain vocabulary for each market.
Combining natural conversation with current search, internal knowledge, MCP tools, or developer functions.
Consolidating chat, search, document collections, and voice workloads behind the same API account and control plane.
Capabilities
Consumes live audio and returns streamed speech without requiring the application to assemble separate transcription, text-model, and speech-synthesis vendors.
SpaceXAI says Think Fast models reason while speaking so tool planning can begin before the agent finishes its initial sentence.
Exposes session, conversation, audio, transcript, response, tool-call, and error events for a custom client experience.
Server VAD can detect when a person starts and stops speaking, with adjustable threshold, silence duration, and prefix padding.
Applications can let a speaker interrupt the assistant and cancel or truncate playback for more natural turn-taking.
Sessions can use the documented high reasoning setting or disable reasoning when the workflow prioritizes speed and simplicity.
The current guide lists English plus Arabic variants, Bengali, Chinese, French, German, Hindi, Indonesian, Italian, Japanese, Korean, Portuguese variants, Russian, Spanish variants, Turkish, and Vietnamese.
The voice API offers a growing built-in roster and can use an eligible custom voice ID created from a reference recording.
The agent can request developer functions for account lookup, scheduling, escalation, or other controlled business operations.
Supported sessions can use web search, X search, document collections or file search, and remote MCP tools.
Short-lived credentials let browsers and mobile apps open realtime sessions without exposing the team's long-lived API key.
Opt-in resumption can replay transcripts, assistant tool calls, and function results after reconnecting, with documented history expiry after 30 minutes of inactivity.
The API can bridge PSTN, PBX, and contact-center calls through a registered direct SIP number and signed incoming-call webhook.
Clients can receive incremental and final assistant transcripts and can optionally configure input transcription events for observability and review.
Process
Step 1
State what the agent may answer or change, what data it can use, prohibited topics, caller verification requirements, escalation triggers, and the exact completion criteria.
Step 2
Tell people they are speaking with an AI and provide recording or data-use notices and consent choices required by the caller's jurisdiction, channel, and business policy.
Step 3
Keep long-lived API keys on the server. Issue scoped, short-lived ephemeral tokens to browsers and mobile clients, and validate signed SIP webhooks before joining a call.
Step 4
Do not treat voice, caller ID, familiarity, or knowledge of public facts as identity proof. Use approved step-up authentication before revealing records or changing an account.
Step 5
Use concise second-person instructions covering role, audience, tone, knowledge sources, one-question-at-a-time behavior, tool rules, confirmation language, handoff, and forbidden claims.
Step 6
Give each tool a typed schema, allowlisted targets, least-privilege account, timeout, budget, and idempotency key. Keep raw credentials and unrestricted HTTP, payment, messaging, or admin tools out of the model's reach.
Step 7
Repeat the exact recipient, booking, cancellation, refund, amount, address, or account change and obtain explicit confirmation immediately before execution.
Step 8
Tune VAD by device and channel, support barge-in, cap silence and retries, and prevent a stale tool result from being spoken after the caller has changed direction.
Step 9
Use a curated knowledge base or verified business system, state uncertainty, avoid inventing policy, and route unsupported medical, legal, financial, safety, or account decisions to a qualified person.
Step 10
Evaluate accents, languages, code-switching, background noise, telephony compression, packet loss, crosstalk, names, addresses, dates, money, serial numbers, DTMF, interruptions, and hostile prompt injection.
Step 11
Track latency, recognition and task accuracy, transfers, abandonments, repeated questions, incorrect actions, tool errors, caller sentiment, cost, and human-review findings—not just vendor benchmarks.
Step 12
Add health checks, reconnect and backoff, call-duration and spend caps, circuit breakers, live-agent transfer, incident logs, prompt and model versioning, and a tested rollback from the latest alias to a pinned model.
Cost
Grok Voice Think Fast 2.0 costs $0.08 per minute of audio sent to or received from the model, equivalent to $4.80 per audio hour, plus $0.004 for a text-only input event. A conversation can incur both input and output audio time. Telephony carriers, phone numbers, SIP infrastructure, search or external tools, storage, monitoring, and application hosting can add separate costs.
$0.08/minute
Usage-based pricing for audio sent or received through the realtime speech-to-speech model.
$0.004/event
Flat charge for a text `conversation.item.create` input event sent without audio.
$0.05/minute
Previous-generation version that remains documented for teams that deliberately pin it rather than follow the latest alias.
From $0.10/hour or $15/1M characters
Separate speech-to-text and text-to-speech APIs for modular pipelines rather than direct speech-to-speech interaction.
Contact sales
Negotiated capacity, regional processing, support, compliance, and service terms for larger or regulated deployments.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Miscellaneous
Choose GPT-Realtime-2 when OpenAI's realtime model behavior, tools, and surrounding API ecosystem better fit the application.
Explore GPT-Realtime-2 →Marketing
Choose ElevenLabs when voice quality, voice design, cloning, multilingual synthesis, or its conversational-agent stack is the main priority.
Explore ElevenLabs →Business Operations
Choose Retell AI for a managed conversational-phone platform with telephony workflow tooling rather than a lower-level model API.
Explore Retell AI →Sales
Choose Bland AI when the team wants a packaged phone-agent platform for sales, support, or service operations.
Explore Bland AI →Sales
Choose Vapi when model-provider flexibility and an orchestration layer for building and operating voice agents matter more than using one vendor end to end.
Explore Vapi →Questions
It is SpaceXAI's flagship realtime speech-to-speech model for building voice assistants, phone agents, and interactive applications with streaming audio, reasoning, and tool use.
Use `grok-voice-think-fast-2.0` to pin this release or `grok-voice-latest` to follow SpaceXAI's current voice model automatically.
The direct API price is $0.08 per minute of audio sent or received, plus $0.004 for a text input event. Carrier, tool, infrastructure, and other service costs can be additional.
Yes. It accepts text and audio and returns text and audio events through a realtime WebSocket session.
SpaceXAI says the model performs reasoning in parallel with spoken output, which can let it plan or initiate tool calls before finishing a short verbal preamble. It does not make the result infallible.
Yes. Current documentation lists developer functions, web search, X search, collection or file search, and remote MCP tools.
Yes. SpaceXAI documents SIP routing for PSTN, PBX, and contact-center calls, including signed incoming-call webhooks and DTMF events.
The current guide lists more than 20 language and regional variants, with additional languages potentially working at varying accuracy. Each target language and dialect still needs real-call testing.
The server should create a short-lived ephemeral token and pass that token to the client. A long-lived xAI API key should never be embedded in browser or mobile code.
The public model page lists 10 concurrent sessions per team and a 120-minute maximum session. Larger deployments should confirm current account limits before launch.
The voice overview says realtime audio is processed and never stored. Because the general API policy describes default retention for requests and responses, confirm the treatment of transcripts, text, tool payloads, session-resumption data, and logs for the exact account configuration.
It can handle well-bounded routine work, but production deployments still need human escalation for identity problems, emergencies, high-impact decisions, repeated misunderstanding, policy exceptions, and actions the caller disputes.
Use the alias only with regression monitoring and a pinned fallback. A versioned model name provides more predictable behavior when a change could affect callers or tool actions.
Not automatically. Those uses require qualified oversight, identity and consent controls, appropriate contracts, retention and regional review, narrow tools, audited escalation, and compliance with all applicable rules.
Bottom line
Grok Voice Think Fast 2.0 is a serious realtime voice model for teams that want low-latency conversation and tool use in one xAI session. Its $0.08-per-minute pricing, multilingual support, WebSocket event model, ephemeral tokens, and SIP path make it practical to prototype and potentially operate at scale. The decision should come from representative call testing, not benchmark tables. Pin the model, disclose the AI, verify identity outside the voice channel, lock tools down, confirm every consequential action, measure real task outcomes, and keep a fast human handoff.
Visit Grok Voice Think Fast 2.0 website ↗
Presence - OpenAI's managed enterprise service deploying chat and voice agents

Low-cost model converting complex documents into machine-readable data

🗣️ Mercury 2 - Inception's diffusion reasoning model built for realtime voice agents

The fastest email experience for ultimate productivity.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.