OpenAI has launched three new real-time audio API models: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper. The Realtime API has also exited beta and now includes MCP server support, image input, and SIP phone calling.
New models offer improved accuracy and multilingual support
GPT-Realtime-2 is built on GPT-5-class reasoning and supports a 128K token context window. It scores 15.2% higher on Big Bench Audio and 13.8% higher on Audio Multichallenger compared to the previous version. GPT-Realtime-Translate supports over 70 input languages and 13 output languages. GPT-Realtime-Whisper provides live captions for speech.
Pricing for GPT-Realtime-2 is $32 per million audio input tokens and $64 per million audio output tokens, with cached input at $0.40 per million tokens. GPT-Realtime-Translate costs $0.034 per minute, and GPT-Realtime-Whisper costs $0.017 per minute. All three models are available now through the OpenAI API and the developer playground.
Zillow reported a 26-point lift in call success rate on its hardest adversarial benchmark, from 69% to 95%, after prompt optimization on GPT-Realtime-2. BolnaAI reported 12.5% lower word error rates on Hindi, Tamil, and Telugu compared to the previous translation approach.



Discussion
0 comments
Log in to join the thread with a thoughtful take, question, or correction.