Observed Signal · Aug 28, 2026 · Technical Release · Source: DEV Community · Impact: 4/5 · Sentiment: Neutral
Gemini 3.5 Transcribe: Real-time Transcription & Diarization
A developer post documents adding real-time transcription and offline speaker diarization support to a macOS meeting-translation app using Google's Gemini 3.5 transcription models. The author explains the critical differences between gemini-3.5-transcribe-live (Live API, streaming, no diarization, 10-minute sessions) and gemini-3.5-transcribe (Interactions API, batch, speaker diarization up to 8 speakers, word-level timestamps, 30-minute diarization limit). The article details required request fields (e.g., timestamp_granularities: ["word"]) to receive word annotations, common pitfalls (chunking, speaker ID continuity, CJK spacing), privacy/cleanup practices, and test-driven engineering lessons. Code is published on GitHub and the post includes links to Google's official transcription docs. Publication date: 2026-08-28.
Google's Gemini transcription model capabilities and limitations (streaming vs batch, diarization, timestamps, API differences) materially affect how developers design real-time captioning, diarization workflows, and privacy handling—relevant to audio captioning, podcast indexing and voice analytics used in media and ad tech.
Track Google Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Google published two similar-named models: gemini-3.5-transcribe-live (Live API / WebSocket streaming) and gemini-3.5-transcribe (Interactions API / batch HTTP).
- gemini-3.5-transcribe-live (streaming) does not support speaker diarization, has a 10-minute-per-session limit, and provides interimInputTranscription for tentative subtitles.
- gemini-3.5-transcribe (batch) supports speaker diarization (up to 8 speakers), word-level timestamps, and diarization is limited to 30 minutes per chunk; audio must be uploaded to the Files API before calling the Interactions API.
- Speaker diarization requires explicit request of word-level annotations (e.g., timestamp_granularities: ["word"]); omitting this field can return a full transcript with zero annotations and no error.
- Implementation code and examples are available at the GitHub repository kkdai/gemini-live-translate-macos; the article was published on 2026-08-28.
Connected Companies & Entities
2 Entities mapped“However, after checking the documentation, I realized that Google released **two models with very similar names but very different capabilit...”
“The code is at [kkdai/gemini-live-translate-macos](https://github.com/kkdai/gemini-live-translate-macos)....”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Gemini 3.1 Pro Enables Audio-to-Text via One API
This technical guide explains how to transcribe audio to text using multimodal LLMs, highlighting that Google’s Gemini 3.1 Pro Preview accepts audio input and can return a transcription (and semantic outputs) in a single request. The article separates transcription into two distinct steps: ASR (audio-to-text) and LLM postprocessing (cleaning, summaries, action-item extraction). Promptra — a Russian OpenAI‑compatible API aggregator — exposes flagship models (including Gemini) via a single endpoint (https://api.promptra.ru/v1) with ruble billing at the Central Bank rate and a 5% top-up service fee. The piece lists per‑model token pricing (e.g., Gemini 3.1 Pro: $2/$12 per 1M tokens → 140/860 ₽) and gives an example cost of roughly 30–40 ₽ to transcribe and produce a protocol for a one‑hour meeting. It notes the catalog lacks dedicated STT endpoints like Whisper.
Gemini adds in-person note-taking to Google Meet
Google has expanded Gemini's Meet capabilities to transcribe and summarise in-person meetings. Announced via a Google Workspace blog post on 2026-08-14, the feature adds a "Take Notes" button in Google Meet on laptop and smartphone that records what is said, produces a structured summary, task list and full transcript, and saves the documents to Google Docs/Drive with an emailed link. Rollout: Android access began on August 11, web rollout started August 14, and iOS will receive the feature after August 31. The feature is targeted at business customers (Workspace Business Standard/Pro, Enterprise Standard/Plus); private users need Google AI Pro or Ultra subscriptions.
Google Launches Gemini 3.8 Live and Extended Thinking Audio Models
Google has introduced two new AI audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, designed for near-zero-latency, real-time conversations and complex multi-step tasks. The models support parallel processing of speech, visual inputs, and function calls, ensuring uninterrupted interactions. They are available via the Gemini Live API and Google AI Studio, with pricing at $0.005 per minute for audio input and $0.018 for output. Gemini 3.8 Live is integrated into Search Live for camera-based interaction, while Extended Thinking is incorporated into Gemini Live and Workspace (Docs, Gmail, Keep) for Google AI subscribers. All generated audio is watermarked with SynthID for transparency. Independent tests show strong performance, with Extended Thinking scoring 82.6 on the Speech-to-Speech Index. Additionally, Google highlighted Gemini 3.5 Transcribe, a streaming speech-to-text model with a 4.0% word error rate.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
