Observed Signal · Aug 28, 2026 · Technical Release · Source: DEV Community · Impact: 4/5 · Sentiment: Neutral

Gemini 3.5 Transcribe: Real-time Transcription & Diarization

Executive Signal Summary

A developer post documents adding real-time transcription and offline speaker diarization support to a macOS meeting-translation app using Google's Gemini 3.5 transcription models. The author explains the critical differences between gemini-3.5-transcribe-live (Live API, streaming, no diarization, 10-minute sessions) and gemini-3.5-transcribe (Interactions API, batch, speaker diarization up to 8 speakers, word-level timestamps, 30-minute diarization limit). The article details required request fields (e.g., timestamp_granularities: ["word"]) to receive word annotations, common pitfalls (chunking, speaker ID continuity, CJK spacing), privacy/cleanup practices, and test-driven engineering lessons. Code is published on GitHub and the post includes links to Google's official transcription docs. Publication date: 2026-08-28.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Google's Gemini transcription model capabilities and limitations (streaming vs batch, diarization, timestamps, API differences) materially affect how developers design real-time captioning, diarization workflows, and privacy handling—relevant to audio captioning, podcast indexing and voice analytics used in media and ad tech.

SIGNAL RADAR

Track Google Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Google published two similar-named models: gemini-3.5-transcribe-live (Live API / WebSocket streaming) and gemini-3.5-transcribe (Interactions API / batch HTTP).
  • gemini-3.5-transcribe-live (streaming) does not support speaker diarization, has a 10-minute-per-session limit, and provides interimInputTranscription for tentative subtitles.
  • gemini-3.5-transcribe (batch) supports speaker diarization (up to 8 speakers), word-level timestamps, and diarization is limited to 30 minutes per chunk; audio must be uploaded to the Files API before calling the Interactions API.
  • Speaker diarization requires explicit request of word-level annotations (e.g., timestamp_granularities: ["word"]); omitting this field can return a full transcript with zero annotations and no error.
  • Implementation code and examples are available at the GitHub repository kkdai/gemini-live-translate-macos; the article was published on 2026-08-28.

Connected Companies & Entities

2 Entities mapped

“However, after checking the documentation, I realized that Google released **two models with very similar names but very different capabilit...”

“The code is at [kkdai/gemini-live-translate-macos](https://github.com/kkdai/gemini-live-translate-macos)....”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: Aug 28, 2026
Original Coverage Title: “[AI in Practice] Gemini 3.5 Transcribe: Real-time Transcription and Speaker Diarization in a macOS Meeting Translation App”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIJun 1, 2026

Gemini 3.1 Pro Enables Audio-to-Text via One API

This technical guide explains how to transcribe audio to text using multimodal LLMs, highlighting that Google’s Gemini 3.1 Pro Preview accepts audio input and can return a transcription (and semantic outputs) in a single request. The article separates transcription into two distinct steps: ASR (audio-to-text) and LLM postprocessing (cleaning, summaries, action-item extraction). Promptra — a Russian OpenAI‑compatible API aggregator — exposes flagship models (including Gemini) via a single endpoint (https://api.promptra.ru/v1) with ruble billing at the Central Bank rate and a 5% top-up service fee. The piece lists per‑model token pricing (e.g., Gemini 3.1 Pro: $2/$12 per 1M tokens → 140/860 ₽) and gives an example cost of roughly 30–40 ₽ to transcribe and produce a protocol for a one‑hour meeting. It notes the catalog lacks dedicated STT endpoints like Whisper.

Read assessment
Productivity & CollaborationAug 14, 2026

Gemini adds in-person note-taking to Google Meet

Google has expanded Gemini's Meet capabilities to transcribe and summarise in-person meetings. Announced via a Google Workspace blog post on 2026-08-14, the feature adds a "Take Notes" button in Google Meet on laptop and smartphone that records what is said, produces a structured summary, task list and full transcript, and saves the documents to Google Docs/Drive with an emailed link. Rollout: Android access began on August 11, web rollout started August 14, and iOS will receive the feature after August 31. The feature is targeted at business customers (Workspace Business Standard/Pro, Enterprise Standard/Plus); private users need Google AI Pro or Ultra subscriptions.

Read assessment
AISep 16, 2026

Google Launches Gemini 3.8 Live and Extended Thinking Audio Models

Google has introduced two new AI audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, designed for near-zero-latency, real-time conversations and complex multi-step tasks. The models support parallel processing of speech, visual inputs, and function calls, ensuring uninterrupted interactions. They are available via the Gemini Live API and Google AI Studio, with pricing at $0.005 per minute for audio input and $0.018 for output. Gemini 3.8 Live is integrated into Search Live for camera-based interaction, while Extended Thinking is incorporated into Gemini Live and Workspace (Docs, Gmail, Keep) for Google AI subscribers. All generated audio is watermarked with SynthID for transparency. Independent tests show strong performance, with Extended Thinking scoring 82.6 on the Speech-to-Speech Index. Additionally, Google highlighted Gemini 3.5 Transcribe, a streaming speech-to-text model with a 4.0% word error rate.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.