Observed Signal · May 12, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Positive

Open-source Local Meeting Transcription App for macOS

Executive Signal Summary

A developer built Scripta, an open-source macOS app that records dual-channel meetings (microphone + system audio), transcribes both streams entirely on-device, and generates AI summaries without cloud requests. The app uses whisper.cpp with Metal GPU acceleration to transcribe microphone audio and Apple’s SFSpeechRecognizer for system/remote audio captured via ScreenCaptureKit. Scripta integrates with a local Ollama instance for streaming summary generation (default model qwen2.5:3b). The post documents engineering details: building a static whisper.cpp library, Swift bridging, a 5-second sliding-window transcription strategy, handling Voice Processing IO quirks (unexpected 9-channel mic format and audio ducking), and distribution via GitHub Releases with a curl-based installer rather than the App Store.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Demonstrates a practical, privacy-preserving on-device transcription + local LLM summary workflow relevant to engineers and organizations concerned with data exfiltration, but is a developer project rather than an industry-wide product or platform change.

SIGNAL RADAR

Track Apple Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Scripta is an open-source macOS app that records microphone and system audio separately, transcribes both in real time, and generates summaries entirely on-device.
  • Microphone audio is transcribed with whisper.cpp (Metal GPU acceleration); the author reports the 'base' model (142 MB) runs >15x real-time on Apple Silicon (≈5s audio in ~0.3s).
  • System/remote audio is captured using ScreenCaptureKit (audio-only capture) and transcribed with Apple’s on-device SFSpeechRecognizer.
  • Scripta uses a local Ollama server (POST to localhost:11434) for streaming AI summary generation; the default model is qwen2.5:3b.
  • Voice Processing IO causes undocumented behavior: it can switch mic output to a 9-channel format and enable system audio ducking; the author implements manual channel extraction and resampling to avoid crashes and silent recordings.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: May 12, 2026
Original Coverage Title: “Building a 100% Local Meeting Transcription App for macOS with whisper.cpp and ScreenCaptureKit”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Productivity & Collaboration SaaSMay 23, 2026

Meetergo Log: Free Offline Meeting Transcription

t3n reviewed Meetergo Log, a free meeting-transcription tool from German meetings specialist Meetergo that performs transcription and optional summarization entirely on the user’s machine. The application captures audio from the local microphone and speakers and does not send data to a cloud interface; it leverages the open-source Whisper transcription model and can run a large language model locally to generate post-meeting summaries. The article situates Meetergo Log in a broader market for offline privacy-preserving tools, cites a 2025 Bitkom survey on widespread use of video-conferencing software, and notes reporting (Huffington Post) about potential legal risks from incorrect AI summaries. t3n tests remaining usability and accuracy trade-offs in its Tool Time episode and highlights the privacy rationale for local transcription.

Read assessment
Creation & Asset ManagementMay 21, 2026

Voice2Sub: Local AI Desktop Subtitle Generator

Voice2Sub is a desktop AI subtitle and transcription app built to generate subtitles and transcripts from local video and audio files without uploading media to browser tools. Published May 21, 2026, the app uses Whisper AI recognition for speech-to-text, runs on Windows x64, macOS Apple Silicon and Linux x64, and supports hardware acceleration (CUDA on compatible Windows/Linux, Metal on Apple Silicon). Voice2Sub exports common subtitle and transcript formats (SRT, VTT, TXT, LRC, CSV) and provides users control over model selection and transcription settings. The project is available via a website, downloadable releases, and a GitHub repository.

Read assessment
Large Language Models (LLM) & AIAug 28, 2026

Gemini 3.5 Transcribe: Real-time Transcription & Diarization

A developer post documents adding real-time transcription and offline speaker diarization support to a macOS meeting-translation app using Google's Gemini 3.5 transcription models. The author explains the critical differences between gemini-3.5-transcribe-live (Live API, streaming, no diarization, 10-minute sessions) and gemini-3.5-transcribe (Interactions API, batch, speaker diarization up to 8 speakers, word-level timestamps, 30-minute diarization limit). The article details required request fields (e.g., timestamp_granularities: ["word"]) to receive word annotations, common pitfalls (chunking, speaker ID continuity, CJK spacing), privacy/cleanup practices, and test-driven engineering lessons. Code is published on GitHub and the post includes links to Google's official transcription docs. Publication date: 2026-08-28.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.