Observed Signal · Mar 26, 2026 · Product Launch · Source: techcrunch · Impact: 3/5 · Sentiment: Positive

Cohere open-sources Transcribe voice model

Executive Signal Summary

Cohere has released Transcribe, its first open-source automatic speech recognition (ASR) model designed for transcription tasks like note-taking and speech analysis. The 2-billion-parameter model is optimized for consumer-grade GPUs and supports 14 languages. Cohere says Transcribe achieved a 5.42 average word error rate (WER) on the Hugging Face Open ASR leaderboard, outperforming several comparator models, and processed audio at an estimated 525 minutes per minute. Human evaluators favored its outputs in head-to-head comparisons (61% average win rate), though it lagged on Portuguese, German and Spanish. Cohere will offer Transcribe free via its API, make it available on Model Vault, and plans integration into its enterprise orchestration platform North.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

An open-source, efficient ASR model that outperforms peers and is self-hostable can accelerate transcription, captioning and audio analytics tools across media and marketing workflows; its free API and integration into Cohere’s enterprise stack broaden adoption.

SIGNAL RADAR

Track Cohere Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Cohere launched Transcribe, an open-source automatic speech recognition (ASR) model.
  • Transcribe is 2 billion parameters and supports 14 languages: English, French, German, Italian, Spanish, Portuguese, Greek, Dutch, Polish, Chinese, Japanese, Korean, Vietnamese, and Arabic.
  • Cohere reports Transcribe achieved an average WER of 5.42 on the Hugging Face Open ASR leaderboard, ranking above Zoom Scribe v1, IBM Granite 4.0 1B, ElevenLabs Scribe v2, and Qwen3-ASR-1.7B Speech.
  • Cohere says Transcribe can process 525 minutes of audio in one minute and had an average 61% win rate in human evaluations for accuracy, coherence and usability; it performed worse on Portuguese, German and Spanish.
  • Cohere will make Transcribe available free via its API, on its Model Vault managed inference platform, and plans to integrate it into its North enterprise agent orchestration platform.

Ontology Mapping & Concepts

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: techcrunch•Published: Mar 26, 2026
Original Coverage Title: “Cohere launches an open source voice model specifically for transcription | TechCrunch”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

InfrastructureSep 6, 2026

Microsoft Launches MAI-Transcribe-2, Cheapest and Most Accurate in Market

Microsoft AI has launched MAI-Transcribe-2, a speech-to-text model that supports 60 languages and achieves top accuracy on the FLEURS benchmark with a 5.2% average word error rate. The model is 10x faster than OpenAI's GPT-Transcribe, 7x faster than ElevenLabs' Scribe v2, and 5x faster than Google's Gemini 3.5 Transcribe. It offers features like speaker diarization, word-level timestamps, keyword biasing, and code-switching. Priced at $0.10 per audio hour (promotional until end of year), it undercuts competitors significantly. For Thai, it achieves a 3.4% word error rate, besting Gemini 3.5 Transcribe (3.8%) and Whisper v3-large (8.7%). The model is available via Microsoft Foundry, MAI Playground, and OpenRouter, and is part of Microsoft's strategy to replace OpenAI technologies with in-house models.

Read assessment
Speech Recognition / Conversational AIMay 20, 2026

Corti launches Symphony clinical Speech-to-Text

Corti announced Symphony for Speech-to-Text, a new generation of clinical-grade speech-to-text models and a production-grade API for real-time dictation, conversational transcription, and batch audio processing. Corti reports large accuracy gains over generalist models and clinical dictation baselines: Symphony achieved 1.4% WER on English medical terminology (vs. 17–19% for tested generalist models), 98.3% recall on formatted clinical entities (dosages, measurements, dates), 4.6% WER on real-world English medical dictation versus Dragon Medical One’s 5.7%, and consistent multilingual improvements (German 2.4% WER, French 3.9% WER). Corti positions Symphony as a speech layer tailored for the emerging "agentic" generation of clinical AI tools; early adopters such as Voicepoint report using it in multilingual clinical environments.

Read assessment
Large Language Models (LLM) & AIJun 1, 2026

Gemini 3.1 Pro Enables Audio-to-Text via One API

This technical guide explains how to transcribe audio to text using multimodal LLMs, highlighting that Google’s Gemini 3.1 Pro Preview accepts audio input and can return a transcription (and semantic outputs) in a single request. The article separates transcription into two distinct steps: ASR (audio-to-text) and LLM postprocessing (cleaning, summaries, action-item extraction). Promptra — a Russian OpenAI‑compatible API aggregator — exposes flagship models (including Gemini) via a single endpoint (https://api.promptra.ru/v1) with ruble billing at the Central Bank rate and a 5% top-up service fee. The piece lists per‑model token pricing (e.g., Gemini 3.1 Pro: $2/$12 per 1M tokens → 140/860 ₽) and gives an example cost of roughly 30–40 ₽ to transcribe and produce a protocol for a one‑hour meeting. It notes the catalog lacks dedicated STT endpoints like Whisper.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.