Observed Signal · May 7, 2026 · Product Launch · Source: OpenAI Blog · Impact: 4/5 · Sentiment: Positive

OpenAI launches three Realtime voice models

Executive Signal Summary

OpenAI announced the addition of three realtime voice-intelligence models to its Realtime API on May 7, 2026: GPT‑Realtime‑2, a GPT‑5‑class reasoning voice model for realistic conversational and agentic workflows; GPT‑Realtime‑Translate, which provides live translation with support for more than 70 input languages and 13 output languages; and GPT‑Realtime‑Whisper, a low-latency streaming speech-to-text transcription capability. The features are intended for customer service, education, media, events and creator platforms. Translate and Whisper are billed by the minute while GPT‑Realtime‑2 is billed by token consumption. OpenAI said it has embedded safety guardrails and active classifiers to halt conversations that violate harmful-content policies. The announcement follows OpenAI’s published evaluations showing gains versus prior realtime models and details pricing and safety controls in the Realtime API documentation.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

A major AI platform (OpenAI) released realtime voice models that expand conversational/voice capabilities (reasoning, translation, live transcription) and can materially affect voice interfaces, CX, and audio-driven product workflows across industries.

SIGNAL RADAR

Track Vimeo Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • OpenAI launched three new realtime audio models on May 7, 2026: GPT‑Realtime‑2, GPT‑Realtime‑Translate, and GPT‑Realtime‑Whisper.
  • GPT‑Realtime‑2 is described as a GPT‑5‑class reasoning voice model built for realistic conversational and agentic workflows.
  • GPT‑Realtime‑Translate supports over 70 input languages and 13 output languages for live conversational translation.
  • GPT‑Realtime‑Whisper provides live streaming speech-to-text transcription.
  • All three models are available in OpenAI’s Realtime API; Translate and Whisper are billed by the minute and GPT‑Realtime‑2 is billed by token consumption; OpenAI says it added safety guardrails and classifiers to prevent abuse.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: OpenAI Blog•Published: May 7, 2026
Original Coverage Title: “Advancing voice intelligence with new models in the API”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Conversational AI & Voice ModelsMay 8, 2026

OpenAI Releases GPT‑Realtime‑2, Translate, Whisper

OpenAI introduced a Chrome extension for Codex that lets the coding agent run in the browser background across tabs (extension currently in the Codex app but not yet available in the UK/EU) and reported that Codex sees over four million weekly users. Separately, OpenAI published three Realtime API voice models — GPT‑Realtime‑2, GPT‑Realtime‑Translate and GPT‑Realtime‑Whisper — bringing GPT‑5‑class reasoning and low‑latency streaming to voice agents. GPT‑Realtime‑2 supports interruption recovery and tool use; Translate covers 70+ input to 13 output languages; Whisper delivers streaming transcription. Pricing published: GPT‑Realtime‑2 — $32 per 1M audio‑input tokens and $64 per 1M audio‑output tokens; GPT‑Realtime‑Translate ~$0.00034/min; GPT‑Realtime‑Whisper ~$0.00017/min. OpenAI says ChatGPT will receive related voice updates in future releases.

Read assessment
Conversational AI & Voice ModelsMay 20, 2026

OpenAI GPT‑Realtime‑2 Adds GPT‑5‑Class Reasoning for Voice

OpenAI released three speech-focused models and highlighted GPT‑Realtime‑2 as the first voice model it describes as having “GPT‑5‑class” reasoning. The article examines what changes for voice agents when reasoning happens inside a real‑time speech model versus the traditional pipeline (speech-to-text → text LLM → text-to-speech). Native speech-to-speech with on-model reasoning promises lower latency and preserved nonverbal cues, but retains tradeoffs: loss of an auditable text transcript, harder deterministic tool-calling, and potential response latency that is audible in voice interactions. The author recommends structured, production-focused evaluations (multi-step retention, interruption handling, latency under load, tool-call accuracy, and graceful uncertainty) and advises teams to run the existing pipeline in shadow mode while measuring with their real prompts and tooling before migrating. Publication date: 2026-05-20.

Read assessment
Conversational AI & ChatbotsJul 9, 2026

OpenAI launches GPT‑Live voice models

OpenAI introduced GPT‑Live, a new full‑duplex family of voice models powering ChatGPT Voice that can listen and speak simultaneously and delegate complex tasks to other models in the background. GPT‑Live (including GPT‑Live‑1 and GPT‑Live‑1 mini) adds live translation, visual answer cards and improved noise suppression; the mini version will be available to free ChatGPT users. OpenAI says GPT‑Live can hand off deeper analyses to larger models such as GPT‑5.5 (and an upcoming GPT‑5.6) and then return results into the ongoing conversation. The models are rolling out on ChatGPT for web, iOS and Android, with API access planned and developer registration open. The article references hands‑on testing by AI expert Jens Polomski and notes more than 150 million people already use ChatGPT’s voice/dictation features.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.