The Rundown AI homepage

Independent tool overview

Grok Voice Think Fast 2.0 at a glance

Grok Voice Think Fast 2.0 is SpaceXAI's flagship realtime speech-to-speech model for voice assistants, phone agents, and interactive applications. It streams audio and text over WebSocket, supports reasoning while speaking, interruption handling, multiple languages and voices, and tool calls to search, document collections, MCP servers, or developer functions. Its $0.08-per-audio-minute price is simple, but production voice automation still requires consent, identity verification, narrow tool authority, latency and call-quality testing, human escalation, and rigorous monitoring.

Visit the official Grok Voice Think Fast 2.0 site ↗
Grok Voice Think Fast 2.0 product preview
Product type
Realtime speech-to-speech model API
Current status
Active flagship voice model
Released
July 29, 2026
Model ID
grok-voice-think-fast-2.0
Latest alias
grok-voice-latest
Endpoint
wss://api.x.ai/v1/realtime
Modalities
Text and audio input; text and audio output
Reasoning
High or none; high is the documented default
Audio price
$0.08/minute sent or received
Text-event price
$0.004 per text input
Session limits
10 concurrent sessions per team; 120-minute maximum
Reviewed
August 31, 2026

Overview

What Grok Voice Think Fast 2.0 is

SpaceXAI released Grok Voice Think Fast 2.0 on July 29, 2026 as the successor to Think Fast 1.0. The versioned API model is `grok-voice-think-fast-2.0`, and the `grok-voice-latest` alias began routing to it on August 5, 2026.

The model runs through the realtime Speech to Speech API at `wss://api.x.ai/v1/realtime`. Applications stream user audio or text to the session and receive incremental audio, transcript, tool-call, and lifecycle events in return.

SpaceXAI says the model reasons in parallel with speech, speaks in shorter conversational turns, asks one question at a time, and improves transcription under noise and telephony compression. The launch's benchmark and Starlink conversion claims are vendor-published or vendor-selected results and should be reproduced on the exact callers, accents, networks, tools, and business flow being deployed.

Sessions can use web search, X search, file or collection search, remote MCP servers, and custom functions. That enables live support, scheduling, account lookup, and guided workflows, but any tool that sends, books, refunds, cancels, changes records, or exposes private data needs separate authorization and confirmation controls.

The current API price for Think Fast 2.0 is $0.08 per minute of audio sent or received, plus $0.004 for a text input event. The direct model documentation lists 10 concurrent sessions per team, a 120-minute maximum session, and the us-east-1 cluster.

Use cases

Who Grok Voice Think Fast 2.0 is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Customer-support voice agents

Handling routine questions, looking up approved account information, and preparing or completing narrowly authorized support steps with clear human escalation.

Phone and contact-center workflows

Building inbound SIP experiences that converse, accept DTMF input, call business tools, and transfer to a person when the policy or confidence threshold requires it.

Interactive app assistants

Adding low-latency conversational help to browser, mobile, kiosk, or device experiences with short-lived client authentication.

Multilingual conversations

Supporting the 20-plus documented language variants while testing dialect, code-switching, names, numbers, and domain vocabulary for each market.

Tool-using voice workflows

Combining natural conversation with current search, internal knowledge, MCP tools, or developer functions.

Teams already using xAI

Consolidating chat, search, document collections, and voice workloads behind the same API account and control plane.

Capabilities

Core Grok Voice Think Fast 2.0 features

1

Direct speech-to-speech

Consumes live audio and returns streamed speech without requiring the application to assemble separate transcription, text-model, and speech-synthesis vendors.

2

Parallel speech reasoning

SpaceXAI says Think Fast models reason while speaking so tool planning can begin before the agent finishes its initial sentence.

3

Realtime WebSocket events

Exposes session, conversation, audio, transcript, response, tool-call, and error events for a custom client experience.

4

Voice activity detection

Server VAD can detect when a person starts and stops speaking, with adjustable threshold, silence duration, and prefix padding.

5

Barge-in handling

Applications can let a speaker interrupt the assistant and cancel or truncate playback for more natural turn-taking.

6

Reasoning control

Sessions can use the documented high reasoning setting or disable reasoning when the workflow prioritizes speed and simplicity.

7

Multilingual operation

The current guide lists English plus Arabic variants, Bengali, Chinese, French, German, Hindi, Indonesian, Italian, Japanese, Korean, Portuguese variants, Russian, Spanish variants, Turkish, and Vietnamese.

8

Built-in and custom voices

The voice API offers a growing built-in roster and can use an eligible custom voice ID created from a reference recording.

9

Function calling

The agent can request developer functions for account lookup, scheduling, escalation, or other controlled business operations.

10

Search and knowledge tools

Supported sessions can use web search, X search, document collections or file search, and remote MCP tools.

11

Ephemeral client tokens

Short-lived credentials let browsers and mobile apps open realtime sessions without exposing the team's long-lived API key.

12

Session resumption

Opt-in resumption can replay transcripts, assistant tool calls, and function results after reconnecting, with documented history expiry after 30 minutes of inactivity.

13

SIP phone integration

The API can bridge PSTN, PBX, and contact-center calls through a registered direct SIP number and signed incoming-call webhook.

14

Transcript events

Clients can receive incremental and final assistant transcripts and can optionally configure input transcription events for observability and review.

Process

How the Grok Voice Think Fast 2.0 workflow works

  1. Step 1

    Define a narrow call purpose

    State what the agent may answer or change, what data it can use, prohibited topics, caller verification requirements, escalation triggers, and the exact completion criteria.

  2. Step 2

    Disclose the AI clearly

    Tell people they are speaking with an AI and provide recording or data-use notices and consent choices required by the caller's jurisdiction, channel, and business policy.

  3. Step 3

    Authenticate the application safely

    Keep long-lived API keys on the server. Issue scoped, short-lived ephemeral tokens to browsers and mobile clients, and validate signed SIP webhooks before joining a call.

  4. Step 4

    Verify the caller separately

    Do not treat voice, caller ID, familiarity, or knowledge of public facts as identity proof. Use approved step-up authentication before revealing records or changing an account.

  5. Step 5

    Write the spoken policy

    Use concise second-person instructions covering role, audience, tone, knowledge sources, one-question-at-a-time behavior, tool rules, confirmation language, handoff, and forbidden claims.

  6. Step 6

    Expose minimum tools

    Give each tool a typed schema, allowlisted targets, least-privilege account, timeout, budget, and idempotency key. Keep raw credentials and unrestricted HTTP, payment, messaging, or admin tools out of the model's reach.

  7. Step 7

    Confirm consequential actions

    Repeat the exact recipient, booking, cancellation, refund, amount, address, or account change and obtain explicit confirmation immediately before execution.

  8. Step 8

    Design for interruption and silence

    Tune VAD by device and channel, support barge-in, cap silence and retries, and prevent a stale tool result from being spoken after the caller has changed direction.

  9. Step 9

    Ground important answers

    Use a curated knowledge base or verified business system, state uncertainty, avoid inventing policy, and route unsupported medical, legal, financial, safety, or account decisions to a qualified person.

  10. Step 10

    Test real audio conditions

    Evaluate accents, languages, code-switching, background noise, telephony compression, packet loss, crosstalk, names, addresses, dates, money, serial numbers, DTMF, interruptions, and hostile prompt injection.

  11. Step 11

    Measure the whole outcome

    Track latency, recognition and task accuracy, transfers, abandonments, repeated questions, incorrect actions, tool errors, caller sentiment, cost, and human-review findings—not just vendor benchmarks.

  12. Step 12

    Operate with a fallback

    Add health checks, reconnect and backoff, call-duration and spend caps, circuit breakers, live-agent transfer, incident logs, prompt and model versioning, and a tested rollback from the latest alias to a pinned model.

Cost

Grok Voice Think Fast 2.0 pricing and free plan

Grok Voice Think Fast 2.0 costs $0.08 per minute of audio sent to or received from the model, equivalent to $4.80 per audio hour, plus $0.004 for a text-only input event. A conversation can incur both input and output audio time. Telephony carriers, phone numbers, SIP infrastructure, search or external tools, storage, monitoring, and application hosting can add separate costs.

Think Fast 2.0 audio

$0.08/minute

Usage-based pricing for audio sent or received through the realtime speech-to-speech model.

  • Equivalent to $4.80 per audio hour
  • Both directions can contribute billable audio
  • Model: grok-voice-think-fast-2.0

Text input

$0.004/event

Flat charge for a text `conversation.item.create` input event sent without audio.

  • Separate from audio duration
  • Useful for dynamic text context or text turns
  • Batching design affects event count

Think Fast 1.0

$0.05/minute

Previous-generation version that remains documented for teams that deliberately pin it rather than follow the latest alias.

  • Equivalent to $3.00 per audio hour
  • Deprecated in the current model guide
  • Test behavior and migration timing before relying on it

Supporting voice APIs

From $0.10/hour or $15/1M characters

Separate speech-to-text and text-to-speech APIs for modular pipelines rather than direct speech-to-speech interaction.

  • Batch STT: $0.10/hour
  • Streaming STT: $0.20/hour
  • TTS: $15 per million characters

Enterprise

Contact sales

Negotiated capacity, regional processing, support, compliance, and service terms for larger or regulated deployments.

  • Custom limits and SLAs may be available
  • BAA and data-processing requirements need contract confirmation
  • Infrastructure and carrier charges may be separate

Pricing checked . Check current pricing at the source ↗

Assessment

Grok Voice Think Fast 2.0 strengths and limitations

Where it stands out

  • One realtime model handles audio understanding, conversational reasoning, and spoken output.
  • The versioned model name allows controlled production pinning while the latest alias supports automatic upgrades.
  • Sub-second launch latency claims and streaming output are well suited to interactive voice experiences.
  • Parallel reasoning can make tool use feel faster by starting the action while the agent speaks a short preamble.
  • Configurable VAD and interruption support address common voice-interface friction.
  • Twenty-plus documented language variants cover many major customer markets.
  • Built-in voices, custom voice IDs, telephony codecs, and SIP support cover app and contact-center deployments.
  • Web, X, collection, MCP, and function tools support grounded, action-oriented conversations.
  • Ephemeral tokens reduce the risk of exposing a long-lived API key in browser or mobile code.
  • The $0.08-per-minute model price is easy to understand compared with token accounting during live calls.
  • Official docs publish concurrent-session, maximum-duration, cluster, event, and pricing details.
  • SpaceXAI reports material gains in conversation dynamics, transcription, and tool reliability over Think Fast 1.0.

What to consider

  • The model can mishear names, numbers, addresses, dates, account details, accents, quiet speech, overlapping speakers, or low-quality telephony audio.
  • It can hallucinate policies, diagnoses, legal conclusions, prices, availability, tool results, confirmations, or whether an action succeeded.
  • Headline benchmark results in the launch post come from a particular benchmark provider and setup; they do not establish performance for a specific production call flow.
  • The Starlink sales-conversion and support-containment claim is vendor-reported A/B-test evidence without enough public detail to predict another organization's outcome.
  • Reasoning while speaking can create a polished response before the underlying tool result is fully known, so clients must prevent premature factual claims or stale follow-ups.
  • Tool calls turn an audio mistake or prompt-injection attack into a possible real-world action.
  • Caller speech, background audio, retrieved documents, web pages, tool output, and transferred context can all contain malicious or conflicting instructions.
  • Voice and caller ID are not reliable identity factors and can be spoofed or cloned.
  • Recording, transcription, outbound calling, automated-decision, biometric, accessibility, and AI-disclosure rules vary by jurisdiction and use case.
  • Custom voices create consent, impersonation, ownership, and fraud risks and are currently restricted geographically.
  • The $0.08 rate applies to audio sent or received, so a two-sided conversation can cost more than wall-clock call duration suggests.
  • Carrier, SIP, phone-number, tool-provider, retrieval, application, observability, and human-escalation costs are not included in the model rate.
  • The direct model page lists 10 concurrent sessions per team and a 120-minute maximum, which may require a limit increase or capacity plan.
  • The current model page lists only us-east-1 for direct availability; geography, residency, and latency need confirmation for the target audience.
  • Session resumption is opt-in and documented history expires after 30 minutes of inactivity, so applications must own durable business state separately.
  • SpaceXAI says realtime audio is not stored, while its general API security FAQ describes 30-day default retention for API requests and responses; teams should confirm how transcripts, text events, tool payloads, and logs are treated for their configuration.
  • Zero Data Retention changes or disables stored-data features and must be validated at the team and application level.
  • The `grok-voice-latest` alias can change behavior without a code edit; production systems need a pinned fallback and regression tests.
  • A voice agent must provide an immediate path to a qualified human for emergencies, vulnerable users, high-impact decisions, authentication failures, and repeated misunderstanding.
  • No voice model should autonomously make medical, legal, financial, employment, security, or safety-critical decisions.

Compare

Grok Voice Think Fast 2.0 alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Miscellaneous

GPT-Realtime-2

Choose GPT-Realtime-2 when OpenAI's realtime model behavior, tools, and surrounding API ecosystem better fit the application.

Explore GPT-Realtime-2

Marketing

ElevenLabs

Choose ElevenLabs when voice quality, voice design, cloning, multilingual synthesis, or its conversational-agent stack is the main priority.

Explore ElevenLabs

Business Operations

Retell AI

Choose Retell AI for a managed conversational-phone platform with telephony workflow tooling rather than a lower-level model API.

Explore Retell AI

Sales

Bland AI

Choose Bland AI when the team wants a packaged phone-agent platform for sales, support, or service operations.

Explore Bland AI

Sales

Vapi

Choose Vapi when model-provider flexibility and an orchestration layer for building and operating voice agents matter more than using one vendor end to end.

Explore Vapi

Questions

Grok Voice Think Fast 2.0 FAQs

What is Grok Voice Think Fast 2.0?

It is SpaceXAI's flagship realtime speech-to-speech model for building voice assistants, phone agents, and interactive applications with streaming audio, reasoning, and tool use.

What model name should developers use?

Use `grok-voice-think-fast-2.0` to pin this release or `grok-voice-latest` to follow SpaceXAI's current voice model automatically.

How much does Think Fast 2.0 cost?

The direct API price is $0.08 per minute of audio sent or received, plus $0.004 for a text input event. Carrier, tool, infrastructure, and other service costs can be additional.

Does it transcribe and speak in the same session?

Yes. It accepts text and audio and returns text and audio events through a realtime WebSocket session.

What does 'thinks while it speaks' mean?

SpaceXAI says the model performs reasoning in parallel with spoken output, which can let it plan or initiate tool calls before finishing a short verbal preamble. It does not make the result infallible.

Can it call tools?

Yes. Current documentation lists developer functions, web search, X search, collection or file search, and remote MCP tools.

Can it handle phone calls?

Yes. SpaceXAI documents SIP routing for PSTN, PBX, and contact-center calls, including signed incoming-call webhooks and DTMF events.

How many languages does it support?

The current guide lists more than 20 language and regional variants, with additional languages potentially working at varying accuracy. Each target language and dialect still needs real-call testing.

How should a browser authenticate?

The server should create a short-lived ephemeral token and pass that token to the client. A long-lived xAI API key should never be embedded in browser or mobile code.

What are the current session limits?

The public model page lists 10 concurrent sessions per team and a 120-minute maximum session. Larger deployments should confirm current account limits before launch.

Does xAI store voice audio?

The voice overview says realtime audio is processed and never stored. Because the general API policy describes default retention for requests and responses, confirm the treatment of transcripts, text, tool payloads, session-resumption data, and logs for the exact account configuration.

Can it replace human support agents?

It can handle well-bounded routine work, but production deployments still need human escalation for identity problems, emergencies, high-impact decisions, repeated misunderstanding, policy exceptions, and actions the caller disputes.

Should production use the latest alias?

Use the alias only with regression monitoring and a pinned fallback. A versioned model name provides more predictable behavior when a change could affect callers or tool actions.

Is it safe for healthcare or financial calls?

Not automatically. Those uses require qualified oversight, identity and consent controls, appropriate contracts, retention and regional review, narrow tools, audited escalation, and compliance with all applicable rules.

Bottom line

Our Grok Voice Think Fast 2.0 verdict

Grok Voice Think Fast 2.0 is a serious realtime voice model for teams that want low-latency conversation and tool use in one xAI session. Its $0.08-per-minute pricing, multilingual support, WebSocket event model, ephemeral tokens, and SIP path make it practical to prototype and potentially operate at scale. The decision should come from representative call testing, not benchmark tables. Pin the model, disclose the AI, verify identity outside the voice channel, lock tools down, confirm every consequential action, measure real task outcomes, and keep a fast human handoff.

Visit Grok Voice Think Fast 2.0 website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.