At a glance
- What changed
- OpenAI says three Realtime API audio models—GPT‑Realtime‑2, GPT‑Realtime‑Translate, and GPT‑Realtime‑Whisper—support voice agents that reason, translate, and transcribe in real time.
- Why it matters
- Realtime voice apps often fail on long context, tool calls, or multilingual use. These models target lower-latency voice agents that can keep a conversation going while translating or transcribing, which can expand where voice interfaces are practical.
- Who is affected
- developers, operators, product teams
- What to do next
- Watch developer adoption in the Realtime API, how translation quality holds up in noisy settings, and whether voice agents reliably handle interruptions and tool calls in produc…
What changed
On May 7, 2026, OpenAI introduced three Realtime API audio models: GPT‑Realtime‑2 for voice interactions with stronger reasoning, GPT‑Realtime‑Translate for live speech translation, and GPT‑Realtime‑Whisper for streaming transcription.
Why it matters
Realtime voice apps often fail on long context, tool calls, or multilingual use. These models target lower-latency voice agents that can keep a conversation going while translating or transcribing, which can expand where voice interfaces are practical.
In plain English
Instead of recording audio and sending it later, an app can talk to the API continuously and get immediate speech-to-text, translation, and spoken replies.
What this means for you
Who is affected: developers, operators, product teams
Next move: Watch developer adoption in the Realtime API, how translation quality holds up in noisy settings, and whether voice agents reliably handle interruptions and tool calls in produc…
- GPT‑Realtime‑2 targets live conversations with stronger reasoning and longer context for agent workflows.
- GPT‑Realtime‑Translate supports live speech translation across 70+ input languages into 13 output languages.
- GPT‑Realtime‑Whisper provides low-latency streaming transcription priced per minute.
What remains uncertain
Watch developer adoption in the Realtime API, how translation quality holds up in noisy settings, and whether voice agents reliably handle interruptions and tool calls in production.