Observed Signal · Mar 21, 2025 · Technical Release · Source: OnlineMarketing.de · Impact: 3/5 · Sentiment: Positive
Hello Voice Agents: OpenAI's New Speech Models
OpenAI has launched new Speech-to-Text and Text-to-Speech models in its API, built on GPT-4o and GPT-4o mini, claiming superior transcription and synthesis performance over Whisper and previous TTS models. The Text-to-Speech model enables voices with character via extensive customization, including 11 voices and five Vibes (tone options). Developers can instruct the model on how to say something as well as what to say, enabling more personalized voice experiences. A live interactive demo lets users pick voices and tones, and tailor pronunciation and linguistic features before reading scripts aloud. The two new Speech-to-Text models are designed to maintain high transcription reliability in challenging audio (noisy environments, strong accents, varying speeds), with potential use in call centers and meeting notes. OpenAI also notes integration with the Agents SDK to convert text-based Agents into Voice Agents, broadening automation capabilities for voice-enabled applications.
New OpenAI speech models enabling voice agents; integration with Agents SDK; relevant to voice technologies in modern marketing and customer experiences.
Track OpenAI Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- OpenAI released new Speech-to-Text and Text-to-Speech models in its API, based on GPT-4o and GPT-4o mini, claiming improvements over Whisper and existing TTS models.
- Text-to-Speech supports 11 voices and five Vibes, enabling configurable tone and character for voices.
- Developers can instruct the model on how to say it, not just what to say, enabling more customized experiences.
- Voice Agents can be created for use cases such as empathetic customer service and expressive storytelling, with configurable voice traits.
- OpenAI integrates with the Agents SDK to convert text-based Agents into Voice Agents; two new Speech-to-Text models aim for robust transcription in noisy environments and with accents.
Connected Companies & Entities
1 Entity mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
OpenAI launches three Realtime voice models
OpenAI announced the addition of three realtime voice-intelligence models to its Realtime API on May 7, 2026: GPT‑Realtime‑2, a GPT‑5‑class reasoning voice model for realistic conversational and agentic workflows; GPT‑Realtime‑Translate, which provides live translation with support for more than 70 input languages and 13 output languages; and GPT‑Realtime‑Whisper, a low-latency streaming speech-to-text transcription capability. The features are intended for customer service, education, media, events and creator platforms. Translate and Whisper are billed by the minute while GPT‑Realtime‑2 is billed by token consumption. OpenAI said it has embedded safety guardrails and active classifiers to halt conversations that violate harmful-content policies. The announcement follows OpenAI’s published evaluations showing gains versus prior realtime models and details pricing and safety controls in the Realtime API documentation.
OpenAI Releases GPT‑Realtime‑2, Translate, Whisper
OpenAI introduced a Chrome extension for Codex that lets the coding agent run in the browser background across tabs (extension currently in the Codex app but not yet available in the UK/EU) and reported that Codex sees over four million weekly users. Separately, OpenAI published three Realtime API voice models — GPT‑Realtime‑2, GPT‑Realtime‑Translate and GPT‑Realtime‑Whisper — bringing GPT‑5‑class reasoning and low‑latency streaming to voice agents. GPT‑Realtime‑2 supports interruption recovery and tool use; Translate covers 70+ input to 13 output languages; Whisper delivers streaming transcription. Pricing published: GPT‑Realtime‑2 — $32 per 1M audio‑input tokens and $64 per 1M audio‑output tokens; GPT‑Realtime‑Translate ~$0.00034/min; GPT‑Realtime‑Whisper ~$0.00017/min. OpenAI says ChatGPT will receive related voice updates in future releases.
OpenAI launches GPT‑Live voice models
OpenAI introduced GPT‑Live, a new full‑duplex family of voice models powering ChatGPT Voice that can listen and speak simultaneously and delegate complex tasks to other models in the background. GPT‑Live (including GPT‑Live‑1 and GPT‑Live‑1 mini) adds live translation, visual answer cards and improved noise suppression; the mini version will be available to free ChatGPT users. OpenAI says GPT‑Live can hand off deeper analyses to larger models such as GPT‑5.5 (and an upcoming GPT‑5.6) and then return results into the ongoing conversation. The models are rolling out on ChatGPT for web, iOS and Android, with API access planned and developer registration open. The article references hands‑on testing by AI expert Jens Polomski and notes more than 150 million people already use ChatGPT’s voice/dictation features.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
