Voice-to-action agents
Let callers make requests that require reasoning, API lookups, scheduling, account actions, or multi-step tool use.
Independent tool overview
GPT-Realtime-2 is OpenAI's reasoning-capable speech-to-speech model for live voice agents, with adjustable reasoning, function calling, interruption handling, expressive delivery, and a 128K context window. It remains available, but GPT-Realtime-2.1 is now the improved same-price update.
Visit the official GPT-Realtime-2 site ↗
Overview
GPT-Realtime-2 lets developers build agents that listen, reason, speak, and call tools during a live conversation. It is aimed at customer support, scheduling, commerce, guidance, and other workflows where a voice interface needs to complete tasks instead of merely answering questions.
The model can give a short spoken preamble before a tool call, run multiple tools in parallel, explain what it is checking, and recover verbally when a tool or request fails. Adjustable reasoning levels let builders trade latency and cost for stronger performance on harder turns.
GPT-Realtime-2 is still listed as an active API model. For new deployments, GPT-Realtime-2.1 is the better default because OpenAI says it improves alphanumeric recognition, silence and noise handling, and interruption behavior at the same token prices.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Let callers make requests that require reasoning, API lookups, scheduling, account actions, or multi-step tool use.
Handle natural corrections, interruptions, specialized vocabulary, and empathetic or reassuring delivery.
Turn changing system context into immediate explanations, alerts, and next-step instructions.
Maintain more conversation and tool context across complex sessions with the 128K context window.
Capabilities
Choose minimal, low, medium, high, or xhigh reasoning effort; higher settings can improve hard tasks while adding latency and token use.
Call multiple functions at once when a request needs several independent lookups or actions.
Tell the user what the agent is doing with short phrases such as checking a calendar or looking up an order.
Keep the exchange coherent when a user changes the request or speaks over the response.
Adjust tone and speaking style to the situation, from calm troubleshooting to upbeat confirmation.
Accept still images as additional context alongside live audio and text.
Connect browser, server, or telephony experiences through WebRTC, WebSocket, or SIP-based Realtime flows.
Process
Step 1
Specify what the agent may answer, which actions it may take, and which cases require a human or explicit confirmation.
Step 2
Use WebRTC for client-side voice, WebSocket for server-side control, or SIP for phone-call integration.
Step 3
Start with low reasoning effort, then raise it only for turns where better planning justifies added latency and cost.
Step 4
Expose well-described functions with validated parameters, clear success states, and safe failure messages.
Step 5
Evaluate noise, silence, interruptions, names, numbers, accents, tool failures, latency, disclosure, and escalation paths before launch.
Cost
GPT-Realtime-2 is billed by text, audio, and image tokens. There is no free API tier. GPT-Realtime-2.1 currently uses the same rates and is the stronger choice for new builds.
$4 input / $24 output per 1M tokens
Text context and generated text or reasoning usage.
$32 input / $64 output per 1M tokens
Live spoken input and generated speech.
$5 per 1M tokens
Still images supplied as context to the voice agent.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Business Operations
For Google's low-latency live voice model and Gemini ecosystem.
Explore Gemini 3.1 Flash Live →Sales
For a managed voice-agent platform that handles telephony and lets teams choose among model providers.
Explore Vapi →Business Operations
For a packaged conversational voice platform focused on production phone agents.
Explore Retell AI →Miscellaneous
For realtime speech synthesis that adapts delivery to conversational audio rather than running the full agent brain.
Explore Realtime TTS-2 →Questions
It is an OpenAI API model for live speech-to-speech agents that can reason, call functions, accept corrections, and respond with expressive audio.
Yes. OpenAI still lists it as an active model, but GPT-Realtime-2.1 is now the improved version and uses the same token prices.
OpenAI says version 2.1 improves alphanumeric recognition, silence and noise handling, and interruption behavior.
Current standard rates are $32 per million audio input tokens, $64 per million audio output tokens, $4 per million text input tokens, and $24 per million text output tokens.
Yes. It supports function calling and can make parallel tool calls during a conversation.
Yes. Realtime applications can use SIP for telephony, while WebRTC and WebSocket cover client-side and server-side connections.
Bottom line
GPT-Realtime-2 was a major step from fast voice response to genuinely agentic voice interaction. Its reasoning controls, tool use, recovery behavior, and long context still make it viable, but teams starting now should normally select GPT-Realtime-2.1 and evaluate the full call experience—not just model accuracy—across latency, interruptions, tool failures, safety, and cost.
Visit GPT-Realtime-2 website ↗
Custom Voices - xAI's new tool to clone your voice from short clips for use within Grok's applications

Autoscientist - Adaption's new tool for automating AI model training

Deep Max - Exa's new SOTA agentic search tool

Miso One - Open-source text-to-speech model that reads a speaker’s tone for expressive responses

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.