Observed Signal · Jul 23, 2026 · Technical Release · Source: DEV Community · Impact: 3/5 · Sentiment: Positive
Stateful Video Editing Skill Using Gemini Interactions
A third-party open-source project (omni-skill-claude) packages Google’s gemini-omni-flash-preview model (Omni Flash) behind a small FastMCP server and a Claude Code skill to enable stateful, iterative video generation and editing. The setup exposes eight MCP tools (generate, edit, animate, interpolate, subject-driven generation, restyle user videos, upload to YouTube, and help) and relies on Gemini’s Interactions API to persist visual context with interaction IDs so subsequent edits preserve continuity. The repo is on GitHub (Apache-2.0) and requires Python 3.10+, a Gemini API key, and Claude Code; video generation is synchronous, billable, and supports inline or File-API delivery modes depending on size.
Demonstrates a practical, stateful integration between generative video (Gemini Omni Flash), MCP tooling, and Claude Code that can streamline iterative creative workflows and asset pipelines, but is a third-party technical release rather than a platform policy or market-shifting announcement.
Track Google Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- The omni-skill-claude project wraps Google’s gemini-omni-flash-preview model (Omni Flash) in a FastMCP server and packages it as a Claude Code skill.
- The server exposes eight MCP tools, including generate_video, edit_video (stateful edits), animate_image, interpolate_images, generate_with_subjects, edit_user_video, upload_to_youtube, and get_help.
- Gemini’s Interactions API provides stateful video editing via stored interaction IDs (store=True) so the model preserves visual context across iterative edits.
- Video generation calls are synchronous and billable; outputs can be returned inline (base64) or via the Google File API (uri) — use uri for larger outputs (above ~4 MB).
- The project is open-source on GitHub (xbill9/omni-skill-claude) under Apache-2.0 and supports publishing results via the YouTube Data API v3 with OAuth.
Connected Companies & Entities
5 Entities mapped“wraps Google's `gemini-omni-flash-preview` model (Omni Flash) in a tiny FastMCP server and packages it as a Claude Code skill....”
“wraps Google's `gemini-omni-flash-preview` model (Omni Flash) in a tiny FastMCP server and packages it as a Claude Code skill....”
“This is a third-party community project, not affiliated with or endorsed by Anthropic or Google....”
“Publishes a finished `.mp4` via the YouTube Data API v3 (one-time OAuth setup; defaults to `private`)....”
“Repo: github.com/xbill9/omni-skill-claude (Apache-2.0)...”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Stateful Image-Editing Claude Code Skill Using Gemini
A third-party open-source project, nb2lite-skill-claude, packages Google’s gemini-3.1-flash-lite-image model behind a small FastMCP server and a Claude Code skill to enable stateful image generation and iterative editing. The setup uses Google’s Interactions API so each generation returns an interaction_id that preserves visual context on the server, allowing incremental edits (e.g., “add a neon sign”) without re-describing the entire scene. The repo provides an MCP server exposing four tools (generate_image, edit_image, edit_local_image, get_help), installation options (Claude plugin marketplace, repo clone, project install, Docker), and ships under an Apache-2.0 license. The project dogfoods itself: the article’s cover was generated by the skill. Publication date: 2026-07-22.
Google launches Gemini Omni: In‑chat AI video editing
Google announced Gemini Omni, a multimodal generative-AI video model, at I/O 2026. Gemini Omni enables users to edit, extend and remix videos directly in the Gemini Chat using voice prompts and combines text, images, audio and video understanding in a single system. Google is rolling out the first Omni model (Gemini Omni Flash) globally to Google AI Plus, Pro and Ultra subscribers via the Gemini app and Google Flow, and is integrating the model into YouTube Shorts and YouTube Create for free. Developers and enterprise customers are slated to receive API access in the coming weeks. Early discoveries and limited tests reported via Reddit and TestingCatalog indicate accurate prompt execution and improved audio quality compared with earlier Veo models.
Agent-built generative video pipeline using Claude Code
A developer describes building a two-minute video entirely via an agentic Claude Code session (named “Simona”) that created and composed image generation, text-to-speech, AI-video, and ffmpeg editing skills. The post is a technical walkthrough showing how the agent iteratively built reusable "skills" (with SKILL.md docs and CLI wrappers), tracked costs in a WORKLOG.md ledger, and recovered after a git mishap that deleted assets. The author lists the models and services used (OpenAI gpt-image-2, Google Gemini/Nano Banana, Seedance 2.0, Kling, LTX, ElevenLabs, Google TTS, local Kokoro), provides a cost breakdown ($27.76 for the final locked cut; $45.26 total project spend), and documents engineering patterns and guardrails for safe agent-driven media production.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
