OpenAI Launches Three New Real-Time Voice Models for Realtime API

OpenAI has released three new real-time voice models under its Realtime API. The models are GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper. Three models target reasoning, translation, and transcription GPT-Realtime-2 offers GPT-5 level reasoning and can handle interruptions and tool calls. GPT-Realtime-Translate supports over 70 input languages and 13 output languages. GPT-Realtime-Whisper is a streaming transcription model for […]

OpenAI Launches Three New Real-Time Voice Models for Realtime API

OpenAI has released three new real-time voice models under its Realtime API. The models are GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper.

Three models target reasoning, translation, and transcription

GPT-Realtime-2 offers GPT-5 level reasoning and can handle interruptions and tool calls. GPT-Realtime-Translate supports over 70 input languages and 13 output languages. GPT-Realtime-Whisper is a streaming transcription model for real-time speech-to-text.

Pricing for GPT-Realtime-2 is $32 per 1 million audio input tokens and $64 per 1 million audio output tokens. GPT-Realtime-Translate costs $0.034 per minute. GPT-Realtime-Whisper costs $0.017 per minute.

The new models expand OpenAI's voice capabilities for developers building real-time conversational applications. They compete with other real-time voice AI offerings in the market.

9to5Mac reported the launch. OpenAI has not confirmed the exact availability timeline.

Discussion

0 comments

Log in to join the thread with a thoughtful take, question, or correction.

Add to the discussion