Observed Signal · May 20, 2026 · Product Launch · Source: https://martechseries.com/feed/ · Impact: 2/5 · Sentiment: Neutral
Corti launches Symphony clinical Speech-to-Text
Corti announced Symphony for Speech-to-Text, a new generation of clinical-grade speech-to-text models and a production-grade API for real-time dictation, conversational transcription, and batch audio processing. Corti reports large accuracy gains over generalist models and clinical dictation baselines: Symphony achieved 1.4% WER on English medical terminology (vs. 17–19% for tested generalist models), 98.3% recall on formatted clinical entities (dosages, measurements, dates), 4.6% WER on real-world English medical dictation versus Dragon Medical One’s 5.7%, and consistent multilingual improvements (German 2.4% WER, French 3.9% WER). Corti positions Symphony as a speech layer tailored for the emerging "agentic" generation of clinical AI tools; early adopters such as Voicepoint report using it in multilingual clinical environments.
A technical product launch that materially improves clinical speech accuracy and structured clinical output, which matters for healthcare AI and voice workflows; limited direct impact on core AdTech/MarTech ecosystems.
Track ElevenLabs Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Corti launched Symphony for Speech-to-Text: clinical-grade speech models and a production-grade API for real-time, conversational and batch transcription.
- Corti reports up to 93% lower word error rate (WER) versus leading speech models on English, German and French medical terminology, achieving 1.4% WER in English.
- Symphony reached 98.3% recall on formatted clinical entities (dosages, measurements, dates) compared with 44.3% for the strongest baseline.
- On real-world English medical dictation, Symphony achieved 4.6% WER versus Dragon Medical One’s 5.7%, and showed multilingual gains (German 2.4% WER; French 3.9% WER).
- Voicepoint is an early adopter using Symphony in multilingual clinical environments in Switzerland.
Connected Companies & Entities
3 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Cohere open-sources Transcribe voice model
Cohere has released Transcribe, its first open-source automatic speech recognition (ASR) model designed for transcription tasks like note-taking and speech analysis. The 2-billion-parameter model is optimized for consumer-grade GPUs and supports 14 languages. Cohere says Transcribe achieved a 5.42 average word error rate (WER) on the Hugging Face Open ASR leaderboard, outperforming several comparator models, and processed audio at an estimated 525 minutes per minute. Human evaluators favored its outputs in head-to-head comparisons (61% average win rate), though it lagged on Portuguese, German and Spanish. Cohere will offer Transcribe free via its API, make it available on Model Vault, and plans integration into its enterprise orchestration platform North.
Microsoft Launches MAI-Transcribe-2, Cheapest and Most Accurate in Market
Microsoft AI has launched MAI-Transcribe-2, a speech-to-text model that supports 60 languages and achieves top accuracy on the FLEURS benchmark with a 5.2% average word error rate. The model is 10x faster than OpenAI's GPT-Transcribe, 7x faster than ElevenLabs' Scribe v2, and 5x faster than Google's Gemini 3.5 Transcribe. It offers features like speaker diarization, word-level timestamps, keyword biasing, and code-switching. Priced at $0.10 per audio hour (promotional until end of year), it undercuts competitors significantly. For Thai, it achieves a 3.4% word error rate, besting Gemini 3.5 Transcribe (3.8%) and Whisper v3-large (8.7%). The model is available via Microsoft Foundry, MAI Playground, and OpenRouter, and is part of Microsoft's strategy to replace OpenAI technologies with in-house models.
Devnagri AI Unveils Multilingual Speech Tech for Enterprises
Devnagri AI announced Speech AI, a new capability within its Sovereign Language Infrastructure that adds enterprise-grade speech recognition and voice generation to its platform. Speech AI combines a Conversational Speech Layer (automatic speech recognition, ASR) and Enterprise Voice Generation (text‑to‑speech, TTS) to enable real-time speech-to-text and natural voice responses across 15+ major Indian languages. The capability is designed to integrate into enterprise systems (CRMs, contact center platforms, mobile apps) to support multilingual customer journeys such as contact centers, voice-based KYC/onboarding, collections automation for financial services, and government helplines. Devnagri positions Speech AI as part of its strategy to bridge the gap between English-first digital systems and regional language user preferences in India.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
