Observed Signal · Jun 1, 2026 · Technical Release · Source: DEV Community · Impact: 4/5 · Sentiment: Positive
Gemini 3.1 Pro Enables Audio-to-Text via One API
This technical guide explains how to transcribe audio to text using multimodal LLMs, highlighting that Google’s Gemini 3.1 Pro Preview accepts audio input and can return a transcription (and semantic outputs) in a single request. The article separates transcription into two distinct steps: ASR (audio-to-text) and LLM postprocessing (cleaning, summaries, action-item extraction). Promptra — a Russian OpenAI‑compatible API aggregator — exposes flagship models (including Gemini) via a single endpoint (https://api.promptra.ru/v1) with ruble billing at the Central Bank rate and a 5% top-up service fee. The piece lists per‑model token pricing (e.g., Gemini 3.1 Pro: $2/$12 per 1M tokens → 140/860 ₽) and gives an example cost of roughly 30–40 ₽ to transcribe and produce a protocol for a one‑hour meeting. It notes the catalog lacks dedicated STT endpoints like Whisper.
A major platform model (Google Gemini 3.1 Pro) supports direct audio input and long-context multimodal transcription; availability via a local aggregator with ruble billing and low per‑hour costs materially affects ASR workflows and automation choices for enterprises.
Track Google Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Google's Gemini 3.1 Pro Preview accepts audio as an input modality and can transcribe audio in a single request.
- Promptra provides an OpenAI‑compatible endpoint (https://api.promptra.ru/v1) aggregating flagship models and bills in rubles using the Central Bank exchange rate.
- Catalog pricing example for Gemini 3.1 Pro: $2 / $12 per 1M tokens (input / output), quoted as 140 / 860 ₽ at the stated exchange rate.
- Estimated token consumption: ~32 tokens per second of audio; one hour ≈ 115,000 input tokens; rough cost ≈ 25–30 ₽ for Gemini transcription and ~30–40 ₽ total including postprocessing.
- Promptra charges a 5% service fee on balance top-ups; the catalog (as of 2026-05-29) does not include a specialized STT endpoint like Whisper.
Connected Companies & Entities
4 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Gemini 3.5 Transcribe: Real-time Transcription & Diarization
A developer post documents adding real-time transcription and offline speaker diarization support to a macOS meeting-translation app using Google's Gemini 3.5 transcription models. The author explains the critical differences between gemini-3.5-transcribe-live (Live API, streaming, no diarization, 10-minute sessions) and gemini-3.5-transcribe (Interactions API, batch, speaker diarization up to 8 speakers, word-level timestamps, 30-minute diarization limit). The article details required request fields (e.g., timestamp_granularities: ["word"]) to receive word annotations, common pitfalls (chunking, speaker ID continuity, CJK spacing), privacy/cleanup practices, and test-driven engineering lessons. Code is published on GitHub and the post includes links to Google's official transcription docs. Publication date: 2026-08-28.
Google Launches Gemini 3.8 Live and Extended Thinking Audio Models
Google has introduced two new AI audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, designed for near-zero-latency, real-time conversations and complex multi-step tasks. The models support parallel processing of speech, visual inputs, and function calls, ensuring uninterrupted interactions. They are available via the Gemini Live API and Google AI Studio, with pricing at $0.005 per minute for audio input and $0.018 for output. Gemini 3.8 Live is integrated into Search Live for camera-based interaction, while Extended Thinking is incorporated into Gemini Live and Workspace (Docs, Gmail, Keep) for Google AI subscribers. All generated audio is watermarked with SynthID for transparency. Independent tests show strong performance, with Extended Thinking scoring 82.6 on the Speech-to-Speech Index. Additionally, Google highlighted Gemini 3.5 Transcribe, a streaming speech-to-text model with a 4.0% word error rate.
GPT Proto Integrates Google’s Gemini 3.1 for Developers
GPT Proto, an API platform operated by Talent Tech Global Limited and based in Hong Kong, has integrated Google’s Gemini 3.1 Pro Preview into its unified developer API gateway. The integration gives developers access to Google’s latest multimodal model — released by Google in early 2026 — through GPT Proto’s RESTful API (compatible with OpenAI’s SDK) alongside models from OpenAI, Anthropic, Meta and others. Gemini 3.1 Pro Preview supports extended context windows, multi-step logical reasoning, and processing of text, code and structured data in a single call. GPT Proto’s gateway supports streaming, function calling, batch processing and model routing, and targets use cases such as RAG pipelines, code generation, automated content and analytics. Founder & CEO Sammi Cen said the addition removes separate onboarding or billing requirements for accessing Google’s model.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
