Observed Signal · Sep 1, 2026 · Product Launch · Source: https://martechseries.com/feed/ · Impact: 3/5 · Sentiment: Positive

Phonely Launches Alma Voice LLM, Faster and Cheaper Than OpenAI

Executive Signal Summary

Phonely, an AI-powered voice agent platform, announced the launch of Alma, a large language model (LLM) built specifically for voice agents. Alma was trained on over 10 million real phone conversations, enabling it to handle interruptions, background noise, and transcription errors. It claims sub-185ms time-to-first-token, 63% faster than OpenAI's GPT-4.1, and costs $0.55 per million blended tokens, 84% cheaper. Alma already powers 100% of Phonely's agent conversations, handling millions of calls monthly. The model integrates with any transcriber and text-to-speech provider and offers self-improving capabilities via live call feedback. Phonely reports that one customer saw a 3% increase in sales conversion after switching from OpenAI to Alma, adding $300,000 in monthly revenue.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

A new voice-specific LLM with significant cost and latency advantages over GPT-4.1 could accelerate adoption of AI voice agents in customer engagement and marketing, though Phonely is not a major platform.

SIGNAL RADAR

Track OpenAI Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Alma is trained on over 10 million real phone conversations.
  • Alma delivers sub-185ms time-to-first-token vs ~500ms for OpenAI's GPT-4.1.
  • Alma costs $0.55 per blended million tokens, 84% cheaper than GPT-4.1 at $3.50.
  • Alma already powers 100% of Phonely's agent conversations.
  • One customer switching from OpenAI to Alma saw a 3% sales conversion increase, generating $300,000 extra monthly revenue.

Connected Companies & Entities

1 Entity mapped

“...compared with roughly 500 milliseconds for OpenAI’s GPT-4.1......”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: https://martechseries.com/feed/•Published: Sep 1, 2026
Original Coverage Title: “Phonely Launches Alma, a Voice LLM That’s 63% Faster and 84% Cheaper Than OpenAI”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

AISep 10, 2026

OpenAI Launches GPT-Live-1 Voice Model in API

OpenAI has launched GPT-Live-1 in its API, a voice-native model designed for building voice-enabled applications and business workflows. The model handles listening and speaking simultaneously, reducing latency by eliminating the need for chained speech-to-text, reasoning, and text-to-speech components. It features improved interruption handling, background noise resilience, and the ability to delegate reasoning and tool calls to backend models like GPT-6 Astra or third-party models. The API includes telephony support for full-duplex voice agents. Pricing is set at $0.05 per minute for the voice layer, with backend models billed separately. Early evaluations show an 80% reduction in interruptions for language tutoring compared to turn-based systems.

Read assessment
InfrastructureSep 6, 2026

Microsoft Launches MAI-Transcribe-2, Cheapest and Most Accurate in Market

Microsoft AI has launched MAI-Transcribe-2, a speech-to-text model that supports 60 languages and achieves top accuracy on the FLEURS benchmark with a 5.2% average word error rate. The model is 10x faster than OpenAI's GPT-Transcribe, 7x faster than ElevenLabs' Scribe v2, and 5x faster than Google's Gemini 3.5 Transcribe. It offers features like speaker diarization, word-level timestamps, keyword biasing, and code-switching. Priced at $0.10 per audio hour (promotional until end of year), it undercuts competitors significantly. For Thai, it achieves a 3.4% word error rate, besting Gemini 3.5 Transcribe (3.8%) and Whisper v3-large (8.7%). The model is available via Microsoft Foundry, MAI Playground, and OpenRouter, and is part of Microsoft's strategy to replace OpenAI technologies with in-house models.

Read assessment
Conversational AI & ChatbotsMay 7, 2026

OpenAI launches three Realtime voice models

OpenAI announced the addition of three realtime voice-intelligence models to its Realtime API on May 7, 2026: GPT‑Realtime‑2, a GPT‑5‑class reasoning voice model for realistic conversational and agentic workflows; GPT‑Realtime‑Translate, which provides live translation with support for more than 70 input languages and 13 output languages; and GPT‑Realtime‑Whisper, a low-latency streaming speech-to-text transcription capability. The features are intended for customer service, education, media, events and creator platforms. Translate and Whisper are billed by the minute while GPT‑Realtime‑2 is billed by token consumption. OpenAI said it has embedded safety guardrails and active classifiers to halt conversations that violate harmful-content policies. The announcement follows OpenAI’s published evaluations showing gains versus prior realtime models and details pricing and safety controls in the Realtime API documentation.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.