Observed Signal · Jul 23, 2026 · Technical Release · Source: DEV Community · Impact: 3/5 · Sentiment: Positive

Stateful Video Editing Skill Using Gemini Interactions

Executive Signal Summary

A third-party open-source project (omni-skill-claude) packages Google’s gemini-omni-flash-preview model (Omni Flash) behind a small FastMCP server and a Claude Code skill to enable stateful, iterative video generation and editing. The setup exposes eight MCP tools (generate, edit, animate, interpolate, subject-driven generation, restyle user videos, upload to YouTube, and help) and relies on Gemini’s Interactions API to persist visual context with interaction IDs so subsequent edits preserve continuity. The repo is on GitHub (Apache-2.0) and requires Python 3.10+, a Gemini API key, and Claude Code; video generation is synchronous, billable, and supports inline or File-API delivery modes depending on size.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Demonstrates a practical, stateful integration between generative video (Gemini Omni Flash), MCP tooling, and Claude Code that can streamline iterative creative workflows and asset pipelines, but is a third-party technical release rather than a platform policy or market-shifting announcement.

SIGNAL RADAR

Track Google Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • The omni-skill-claude project wraps Google’s gemini-omni-flash-preview model (Omni Flash) in a FastMCP server and packages it as a Claude Code skill.
  • The server exposes eight MCP tools, including generate_video, edit_video (stateful edits), animate_image, interpolate_images, generate_with_subjects, edit_user_video, upload_to_youtube, and get_help.
  • Gemini’s Interactions API provides stateful video editing via stored interaction IDs (store=True) so the model preserves visual context across iterative edits.
  • Video generation calls are synchronous and billable; outputs can be returned inline (base64) or via the Google File API (uri) — use uri for larger outputs (above ~4 MB).
  • The project is open-source on GitHub (xbill9/omni-skill-claude) under Apache-2.0 and supports publishing results via the YouTube Data API v3 with OAuth.

Connected Companies & Entities

5 Entities mapped

“wraps Google's `gemini-omni-flash-preview` model (Omni Flash) in a tiny FastMCP server and packages it as a Claude Code skill....”

“wraps Google's `gemini-omni-flash-preview` model (Omni Flash) in a tiny FastMCP server and packages it as a Claude Code skill....”

“This is a third-party community project, not affiliated with or endorsed by Anthropic or Google....”

“Publishes a finished `.mp4` via the YouTube Data API v3 (one-time OAuth setup; defaults to `private`)....”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: Jul 23, 2026
Original Coverage Title: “Teaching Claude Code to Direct: A Stateful Video-Editing Skill Built on Gemini's Interactions API and MCP”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Layer 3: Creation & Asset ManagementJul 22, 2026

Stateful Image-Editing Claude Code Skill Using Gemini

A third-party open-source project, nb2lite-skill-claude, packages Google’s gemini-3.1-flash-lite-image model behind a small FastMCP server and a Claude Code skill to enable stateful image generation and iterative editing. The setup uses Google’s Interactions API so each generation returns an interaction_id that preserves visual context on the server, allowing incremental edits (e.g., “add a neon sign”) without re-describing the entire scene. The repo provides an MCP server exposing four tools (generate_image, edit_image, edit_local_image, get_help), installation options (Claude plugin marketplace, repo clone, project install, Docker), and ships under an Apache-2.0 license. The project dogfoods itself: the article’s cover was generated by the skill. Publication date: 2026-07-22.

Read assessment
Large Language Models (LLM) & AIMay 20, 2026

Google launches Gemini Omni: In‑chat AI video editing

Google announced Gemini Omni, a multimodal generative-AI video model, at I/O 2026. Gemini Omni enables users to edit, extend and remix videos directly in the Gemini Chat using voice prompts and combines text, images, audio and video understanding in a single system. Google is rolling out the first Omni model (Gemini Omni Flash) globally to Google AI Plus, Pro and Ultra subscribers via the Gemini app and Google Flow, and is integrating the model into YouTube Shorts and YouTube Create for free. Developers and enterprise customers are slated to receive API access in the coming weeks. Early discoveries and limited tests reported via Reddit and TestingCatalog indicate accurate prompt execution and improved audio quality compared with earlier Veo models.

Read assessment
Generative AI / Creative ProductionJun 13, 2026

Agent-built generative video pipeline using Claude Code

A developer describes building a two-minute video entirely via an agentic Claude Code session (named “Simona”) that created and composed image generation, text-to-speech, AI-video, and ffmpeg editing skills. The post is a technical walkthrough showing how the agent iteratively built reusable "skills" (with SKILL.md docs and CLI wrappers), tracked costs in a WORKLOG.md ledger, and recovered after a git mishap that deleted assets. The author lists the models and services used (OpenAI gpt-image-2, Google Gemini/Nano Banana, Seedance 2.0, Kling, LTX, ElevenLabs, Google TTS, local Kokoro), provides a cost breakdown ($27.76 for the final locked cut; $45.26 total project spend), and documents engineering patterns and guardrails for safe agent-driven media production.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.