How Does GPT Realtime Power Voice AI Agents?

TL;DR
GPT realtime enables low-latency voice agents through one model that natively understands and produces audio, allowing it to recognize vocal cues, express emotion, and switch languages. The generally available Realtime API adds image input, SIP telephony, MCP support, asynchronous function calling, EU data residency, and improved context management for production applications.
Transcript
Good morning and thanks for joining us today. We're taking a big step towards enabling AI agents that can talk and listen with a human level voice quality. We're excited to release a new advanced speech model GPT real time as well as an improved real-time API. Both are generally available for developers to build with starting today. Voice is one of... Read More
Key Insights
- GPT realtime is a speech-to-speech model that natively understands and produces audio. Unlike an architecture built from separate transcription, language, and voice models, its unified design supports fast responses and preserves meaningful vocal signals such as laughter, sighs, emotion, and changes in language.
- Native audio understanding enables expressive and multilingual conversations. The demonstration showed the model changing emotional delivery after a fictional lottery-ticket scenario and producing a short rhyming poem that switched among English, Spanish, and Japanese within the same response.
- Instruction following makes the model more steerable in multi-turn conversations. Developers can specify pace, tone, style, policies, or character roles, while users can provide conversational directions. In the refund demonstration, the agent repeatedly maintained a firm $10 limit despite pressure to approve $25.
- Image input allows a real-time voice agent to discuss visual information supplied during a conversation. In the demonstration, the model identified a child, a stuffed unicorn, a toy train track, a green hair clip, a rainbow mane and tail, sunlight, and a possible balance risk.
- Model training combines high-quality voice data with specialized reward models to improve naturalness. The research team also described more powerful models, a sample-efficient reinforcement learning algorithm, stronger data filtering, and a data flywheel based directly on real customer use cases.
- Benchmark results show measurable gains in instruction following and function calling. GPT realtime achieved over 30% accuracy on an audio version of the Scale MultiChallenge instruction-following benchmark and 66% accuracy on ComplexFuncBench Audio, which evaluates difficult function-calling scenarios.
- Reliability work targets practical speech-agent problems beyond general conversational quality. The team built focused evaluations and training data for handling long alphanumeric strings, including phone numbers and VINs, and for responding more appropriately when the model cannot hear a user clearly.
- The generally available Realtime API adds production features for multimodal and connected voice applications. These include image input, EU data residency, asynchronous function calling, cache-friendly context management, an updated Agents SDK, SIP telephony for phone scenarios, and MCP for adding pluggable capabilities.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: What is GPT realtime and how does it process speech?
GPT realtime is a speech-to-speech model that natively understands audio and produces audio through one model. It differs from a classic architecture that separates transcription, language processing, and voice generation. This design is fast and can retain audio information such as laughter or sighs. It can also speak with varied emotion and switch languages within a sentence.
Q: How does GPT realtime improve natural voice interactions?
GPT realtime improves voice interactions through higher audio quality, broader emotional expression, and stronger control over how it speaks. Its training combines high-quality voice data with specialized reward models designed to make speech sound more natural. Developers and users can direct its pace, tone, style, or character, while native audio processing helps it interpret nonverbal signals such as laughter and sighs.
Q: How well does GPT realtime follow developer instructions?
GPT realtime is trained to follow instructions supplied by developers and users across difficult, multi-turn conversations. The demonstration gave the agent a policy forbidding refunds above $10. When asked to approve a $25 refund, and then pressured several times, it consistently refused while remaining conversational. On an audio version of the Scale MultiChallenge instruction-following benchmark, the model scored over 30% accuracy.
Q: How does image input work in the Realtime API?
Image input lets developers send a picture to the Realtime API so the voice model can discuss what is visible. During the demonstration, it described a child standing on a stuffed unicorn, a wooden toy train track, scattered colorful pieces, a green hair clip, a rainbow mane and tail, and sunlight. It also identified that standing on the toy might be wobbly and suggested guiding the child down.
Q: How does GPT realtime use function calling?
GPT realtime was trained to decide when a function should be called, select the appropriate function, and pass the correct arguments. The team emphasized function calling as a major priority because it allows conversational agents to take useful actions. On ComplexFuncBench Audio, an evaluation designed around challenging function-calling scenarios, the new model achieved 66% accuracy and improved over earlier models.
Q: What production features were added to the Realtime API?
The generally available Realtime API adds image input, EU data residency, asynchronous function calling, and more tools for managing context in a cache-friendly way. The Agents SDK was updated to incorporate these changes. The API also supports SIP telephony for voice applications that operate over phone systems and MCP for connecting models to pluggable capabilities and tools.
Q: How was GPT realtime trained for reliability?
GPT realtime was developed with high-quality voice data, specialized reward models, a highly sample-efficient reinforcement learning algorithm, more powerful models, and a major investment in data quality. The team filtered speech-related data and built a data flywheel that trains on real customer use cases. It also created targeted evaluations for unclear speech and difficult alphanumeric strings such as phone numbers and VINs.
Q: What kinds of applications can developers build with GPT realtime?
Developers can use GPT realtime and the Realtime API for low-latency voice applications in areas mentioned in the presentation, including customer support, education, tutoring, healthcare, and phone-based services. SIP telephony supports voice-over-phone scenarios, image input adds visual context, and MCP lets the model take actions through connected tools. The platform was also described as capable of serving voice applications at very large scale.
Summary & Key Takeaways
-
GPT realtime uses a native speech-to-speech architecture instead of separate transcription, language, and voice models. This unified design reduces latency while preserving audio details such as laughter and sighs. It also supports expressive speech, multilingual switching within a sentence, adjustable delivery, roleplay, and more natural conversational interactions.
-
The model was improved using high-quality voice data, specialized reward models, reinforcement learning, stronger models, targeted evaluations, and customer feedback. Training emphasized instruction following, function calling, unclear audio, and long alphanumeric strings. Reported results include over 30% accuracy on an audio instruction benchmark and 66% on ComplexFuncBench Audio.
-
The generally available Realtime API expands production capabilities with image input, EU data residency, asynchronous function calling, cache-friendly context controls, an updated Agents SDK, SIP telephony, and MCP support. These additions help developers create scalable voice applications that can understand visual context, operate over phone systems, and take actions through connected tools.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from OpenAI 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator