OpenAI Launches Real-Time Audio Models for Multilingual Translation and Live Transcription

August 5, 2026
OpenAI Launches Real-Time Audio Models for Multilingual Translation and Live Transcription
  • OpenAI unveils three real-time audio models in the API: GPT-Realtime-2 for reasoning-enabled live voice interactions, GPT-Realtime-Translate for live multilingual translation across 70+ input languages to 13 output languages, and GPT-Realtime-Whisper for streaming live transcription.

  • Safety and compliance are reinforced with safeguards, active classifiers, and guardrails via the Agents SDK, complemented by usage policies to prevent misuse.

  • The article showcases practical examples and use cases from industries and firms such as Zillow, Deutsche Telekom, and Priceline to illustrate voice-to-action, systems-to-voice, and voice-to-voice workflows.

  • GPT-Realtime-2 supports longer context up to 128K tokens, parallel tool calls, improved recovery, adjustable reasoning levels, and tone control to manage complex live conversations and actions within a session.

  • GPT-Realtime-Translate enables real-time, low-latency translation for multilingual voice conversations in customer support, education, and global events, including live product-education translations.

  • Pricing is set at $32 per 1M input tokens and $64 per 1M output tokens for GPT-Realtime-2, $0.034 per minute for GPT-Realtime-Translate, and $0.017 per minute for GPT-Realtime-Whisper.

  • All three models are available via the Realtime API, with guidance to get started using Codex or the Codex app.

  • GPT-Realtime-Whisper delivers low-latency streaming transcription suitable for live captions, meeting notes, and continuous speech-understanding workflows.

Summary based on 1 source


Get a daily email with more AI stories

More Stories