OpenAI Launches Real-Time Audio Models for Multilingual Translation and Live Transcription
August 5, 2026
OpenAI unveils three real-time audio models in the API: GPT-Realtime-2 for reasoning-enabled live voice interactions, GPT-Realtime-Translate for live multilingual translation across 70+ input languages to 13 output languages, and GPT-Realtime-Whisper for streaming live transcription.
Safety and compliance are reinforced with safeguards, active classifiers, and guardrails via the Agents SDK, complemented by usage policies to prevent misuse.
The article showcases practical examples and use cases from industries and firms such as Zillow, Deutsche Telekom, and Priceline to illustrate voice-to-action, systems-to-voice, and voice-to-voice workflows.
GPT-Realtime-2 supports longer context up to 128K tokens, parallel tool calls, improved recovery, adjustable reasoning levels, and tone control to manage complex live conversations and actions within a session.
GPT-Realtime-Translate enables real-time, low-latency translation for multilingual voice conversations in customer support, education, and global events, including live product-education translations.
Pricing is set at $32 per 1M input tokens and $64 per 1M output tokens for GPT-Realtime-2, $0.034 per minute for GPT-Realtime-Translate, and $0.017 per minute for GPT-Realtime-Whisper.
All three models are available via the Realtime API, with guidance to get started using Codex or the Codex app.
GPT-Realtime-Whisper delivers low-latency streaming transcription suitable for live captions, meeting notes, and continuous speech-understanding workflows.
Summary based on 1 source
Get a daily email with more AI stories
Source

OpenAI • Aug 4, 2026
Advancing voice intelligence with new models in the API