Observed Signal · May 12, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Positive
Open-source Local Meeting Transcription App for macOS
A developer built Scripta, an open-source macOS app that records dual-channel meetings (microphone + system audio), transcribes both streams entirely on-device, and generates AI summaries without cloud requests. The app uses whisper.cpp with Metal GPU acceleration to transcribe microphone audio and Apple’s SFSpeechRecognizer for system/remote audio captured via ScreenCaptureKit. Scripta integrates with a local Ollama instance for streaming summary generation (default model qwen2.5:3b). The post documents engineering details: building a static whisper.cpp library, Swift bridging, a 5-second sliding-window transcription strategy, handling Voice Processing IO quirks (unexpected 9-channel mic format and audio ducking), and distribution via GitHub Releases with a curl-based installer rather than the App Store.
Demonstrates a practical, privacy-preserving on-device transcription + local LLM summary workflow relevant to engineers and organizations concerned with data exfiltration, but is a developer project rather than an industry-wide product or platform change.
Track Apple Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Scripta is an open-source macOS app that records microphone and system audio separately, transcribes both in real time, and generates summaries entirely on-device.
- Microphone audio is transcribed with whisper.cpp (Metal GPU acceleration); the author reports the 'base' model (142 MB) runs >15x real-time on Apple Silicon (≈5s audio in ~0.3s).
- System/remote audio is captured using ScreenCaptureKit (audio-only capture) and transcribed with Apple’s on-device SFSpeechRecognizer.
- Scripta uses a local Ollama server (POST to localhost:11434) for streaming AI summary generation; the default model is qwen2.5:3b.
- Voice Processing IO causes undocumented behavior: it can switch mic output to a 9-channel format and enable system audio ducking; the author implements manual channel extraction and resampling to avoid crashes and silent recordings.
Connected Companies & Entities
4 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Meetergo Log: Free Offline Meeting Transcription
t3n reviewed Meetergo Log, a free meeting-transcription tool from German meetings specialist Meetergo that performs transcription and optional summarization entirely on the user’s machine. The application captures audio from the local microphone and speakers and does not send data to a cloud interface; it leverages the open-source Whisper transcription model and can run a large language model locally to generate post-meeting summaries. The article situates Meetergo Log in a broader market for offline privacy-preserving tools, cites a 2025 Bitkom survey on widespread use of video-conferencing software, and notes reporting (Huffington Post) about potential legal risks from incorrect AI summaries. t3n tests remaining usability and accuracy trade-offs in its Tool Time episode and highlights the privacy rationale for local transcription.
Voice2Sub: Local AI Desktop Subtitle Generator
Voice2Sub is a desktop AI subtitle and transcription app built to generate subtitles and transcripts from local video and audio files without uploading media to browser tools. Published May 21, 2026, the app uses Whisper AI recognition for speech-to-text, runs on Windows x64, macOS Apple Silicon and Linux x64, and supports hardware acceleration (CUDA on compatible Windows/Linux, Metal on Apple Silicon). Voice2Sub exports common subtitle and transcript formats (SRT, VTT, TXT, LRC, CSV) and provides users control over model selection and transcription settings. The project is available via a website, downloadable releases, and a GitHub repository.
Gemini 3.5 Transcribe: Real-time Transcription & Diarization
A developer post documents adding real-time transcription and offline speaker diarization support to a macOS meeting-translation app using Google's Gemini 3.5 transcription models. The author explains the critical differences between gemini-3.5-transcribe-live (Live API, streaming, no diarization, 10-minute sessions) and gemini-3.5-transcribe (Interactions API, batch, speaker diarization up to 8 speakers, word-level timestamps, 30-minute diarization limit). The article details required request fields (e.g., timestamp_granularities: ["word"]) to receive word annotations, common pitfalls (chunking, speaker ID continuity, CJK spacing), privacy/cleanup practices, and test-driven engineering lessons. Code is published on GitHub and the post includes links to Google's official transcription docs. Publication date: 2026-08-28.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
