The Rundown AI homepage

Independent tool overview

GPT-Realtime-2 at a glance

GPT-Realtime-2 is OpenAI's reasoning-capable speech-to-speech model for live voice agents, with adjustable reasoning, function calling, interruption handling, expressive delivery, and a 128K context window. It remains available, but GPT-Realtime-2.1 is now the improved same-price update.

Visit the official GPT-Realtime-2 site ↗
GPT-Realtime-2 product preview
Model ID
gpt-realtime-2
Inputs
Audio, text, and images
Outputs
Audio and text
Context window
128,000 tokens
Max output
32,000 tokens
Current successor
GPT-Realtime-2.1

Overview

What GPT-Realtime-2 is

GPT-Realtime-2 lets developers build agents that listen, reason, speak, and call tools during a live conversation. It is aimed at customer support, scheduling, commerce, guidance, and other workflows where a voice interface needs to complete tasks instead of merely answering questions.

The model can give a short spoken preamble before a tool call, run multiple tools in parallel, explain what it is checking, and recover verbally when a tool or request fails. Adjustable reasoning levels let builders trade latency and cost for stronger performance on harder turns.

GPT-Realtime-2 is still listed as an active API model. For new deployments, GPT-Realtime-2.1 is the better default because OpenAI says it improves alphanumeric recognition, silence and noise handling, and interruption behavior at the same token prices.

Use cases

Who GPT-Realtime-2 is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Voice-to-action agents

Let callers make requests that require reasoning, API lookups, scheduling, account actions, or multi-step tool use.

Customer support calls

Handle natural corrections, interruptions, specialized vocabulary, and empathetic or reassuring delivery.

Live spoken guidance

Turn changing system context into immediate explanations, alerts, and next-step instructions.

Longer voice workflows

Maintain more conversation and tool context across complex sessions with the 128K context window.

Capabilities

Core GPT-Realtime-2 features

1

Configurable reasoning

Choose minimal, low, medium, high, or xhigh reasoning effort; higher settings can improve hard tasks while adding latency and token use.

2

Parallel tool calls

Call multiple functions at once when a request needs several independent lookups or actions.

3

Spoken preambles

Tell the user what the agent is doing with short phrases such as checking a calendar or looking up an order.

4

Interruption and correction handling

Keep the exchange coherent when a user changes the request or speaks over the response.

5

Expressive delivery

Adjust tone and speaking style to the situation, from calm troubleshooting to upbeat confirmation.

6

Image input

Accept still images as additional context alongside live audio and text.

7

Realtime transports

Connect browser, server, or telephony experiences through WebRTC, WebSocket, or SIP-based Realtime flows.

Process

How the GPT-Realtime-2 workflow works

  1. Step 1

    Define the call boundary

    Specify what the agent may answer, which actions it may take, and which cases require a human or explicit confirmation.

  2. Step 2

    Choose a connection

    Use WebRTC for client-side voice, WebSocket for server-side control, or SIP for phone-call integration.

  3. Step 3

    Configure voice and reasoning

    Start with low reasoning effort, then raise it only for turns where better planning justifies added latency and cost.

  4. Step 4

    Add narrow tools

    Expose well-described functions with validated parameters, clear success states, and safe failure messages.

  5. Step 5

    Test real conversations

    Evaluate noise, silence, interruptions, names, numbers, accents, tool failures, latency, disclosure, and escalation paths before launch.

Cost

GPT-Realtime-2 pricing and free plan

GPT-Realtime-2 is billed by text, audio, and image tokens. There is no free API tier. GPT-Realtime-2.1 currently uses the same rates and is the stronger choice for new builds.

Text tokens

$4 input / $24 output per 1M tokens

Text context and generated text or reasoning usage.

  • Cached text input is $0.40 per 1M tokens.
  • Higher reasoning effort can increase output-token use.

Audio tokens

$32 input / $64 output per 1M tokens

Live spoken input and generated speech.

  • Cached audio input is $0.40 per 1M tokens.
  • Actual call cost depends on turn length, interruptions, context, and tool behavior.

Image input

$5 per 1M tokens

Still images supplied as context to the voice agent.

  • Cached image input is $0.50 per 1M tokens.
  • The model does not output images or video.

Pricing checked . Check current pricing at the source ↗

Assessment

GPT-Realtime-2 strengths and limitations

Where it stands out

  • Combines low-latency speech with reasoning and reliable function calling.
  • Preambles and audible tool transparency make pauses feel less confusing.
  • A 128K context window supports longer, more agentic sessions.
  • Adjustable reasoning provides a practical latency-versus-intelligence control.
  • Supports browser, server, and telephony deployment patterns.

What to consider

  • GPT-Realtime-2.1 now improves several production voice behaviors at the same price, so GPT-Realtime-2 is no longer the best starting model.
  • Audio output is relatively expensive, and long conversations can accumulate substantial token cost.
  • Higher reasoning levels can increase latency, which may make a live call feel less natural.
  • Structured outputs, fine-tuning, video, and image output are not supported.
  • Developers remain responsible for AI disclosure, consent, recording laws, tool permissions, escalation, and domain-specific compliance.

Compare

GPT-Realtime-2 alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Sales

Vapi

For a managed voice-agent platform that handles telephony and lets teams choose among model providers.

Explore Vapi

Business Operations

Retell AI

For a packaged conversational voice platform focused on production phone agents.

Explore Retell AI

Miscellaneous

Realtime TTS-2

For realtime speech synthesis that adapts delivery to conversational audio rather than running the full agent brain.

Explore Realtime TTS-2

Questions

GPT-Realtime-2 FAQs

What is GPT-Realtime-2?

It is an OpenAI API model for live speech-to-speech agents that can reason, call functions, accept corrections, and respond with expressive audio.

Is GPT-Realtime-2 still available?

Yes. OpenAI still lists it as an active model, but GPT-Realtime-2.1 is now the improved version and uses the same token prices.

What changed in GPT-Realtime-2.1?

OpenAI says version 2.1 improves alphanumeric recognition, silence and noise handling, and interruption behavior.

How much does GPT-Realtime-2 cost?

Current standard rates are $32 per million audio input tokens, $64 per million audio output tokens, $4 per million text input tokens, and $24 per million text output tokens.

Can GPT-Realtime-2 call tools?

Yes. It supports function calling and can make parallel tool calls during a conversation.

Can it connect to phone calls?

Yes. Realtime applications can use SIP for telephony, while WebRTC and WebSocket cover client-side and server-side connections.

Bottom line

Our GPT-Realtime-2 verdict

GPT-Realtime-2 was a major step from fast voice response to genuinely agentic voice interaction. Its reasoning controls, tool use, recovery behavior, and long context still make it viable, but teams starting now should normally select GPT-Realtime-2.1 and evaluate the full call experience—not just model accuracy—across latency, interruptions, tool failures, safety, and cost.

Visit GPT-Realtime-2 website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.