Source: Google Blog: Gemini Omni & Google Blog: Build with Gemini Omni 1.1 Flash
Google has expanded the Gemini model family with Gemini Omni and Gemini Omni 1.1 Flash, shifting the architecture from pure multimodal comprehension and coding reasoning toward direct generative video creation. While image generation was previously added through Nano Banana, Gemini Omni integrates real-world physics understanding and historical context into generative video workflows.

Conversational video editing and scene transformation
A key difference in Gemini Omni is its support for conversational, multi-turn video editing. Instead of regenerating an entire video from scratch when a detail needs adjustment, creators can supply plain-language instructions to modify specific elements while keeping character appearance, lighting, and environmental continuity intact.
Users can transform existing real-world video clips, alter character actions, add or remove scene elements, and apply material shaders:
Transforming physical objects in a video clip into structured bubble foam while preserving underlying geometry.
Dynamic surface reaction: A mirror rippling like liquid upon touch with reflective material transfer.
Physics simulation and scientific reasoning
Standard generative video tools often struggle with consistent physical dynamics, producing unnatural motion or warped objects. Gemini Omni incorporates physical constraints, allowing the model to simulate gravity, kinetic momentum, inertia, and fluid dynamics more accurately.
This allows the system to generate technical demonstrations, educational animations, and scientific explainers from short text descriptions:
Continuous single-shot tracking of a marble navigating a complex chain-reaction track with consistent gravity and inertia.
Stop-motion claymation explainer illustrating molecular protein folding mechanisms without manual 3D modeling.
Developer features in Gemini Omni 1.1 Flash
Alongside the consumer launch, Google released Gemini Omni 1.1 Flash through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. This release adds several production controls for software engineers and creative tool builders:
- Scene extension up to 40 seconds: The model can analyze up to 10 seconds of prior video context (compared to only the final second in earlier versions) and extend footage in 10-second increments up to a total length of 40 seconds with consistent narrative flow.
- First and last frame interpolation: Developers can define specific starting and ending keyframes to generate continuous camera moves, complex whip-pan transitions, orbital rotations, and seamless looping clips.
- Rapid 360p draft previews: Low-resolution 360p drafts generate up to 60% faster at one-third the standard token cost, enabling fast storyboarding and cheap parameter iteration before final rendering.
- Native 4K upscaling: Final sequences can be upscaled directly to 1080p or 4K resolution with enhanced texture detail and edge sharpness.
- Multimodal video references: Up to three seconds of reference video can be attached to input prompts to guide character movement and scene pacing.
Scene extension demo: Executing cinematic dolly-zoom and 360-degree orbital rotation shots across consecutive generations.
First and last frame control: Seamless transition between two distinct keyframe compositions in a single continuous shot.
Rapid 360p draft generation: Exploring biological microscopy textures with reduced latency and compute costs.
4K upscaling: Macro close-ups with fine foliage details and natural depth-of-field blur.
Safety measures and digital watermarking
To prevent malicious misuse and deepfakes, all videos produced by Gemini Omni include Google’s imperceptible SynthID digital watermark embedded directly into the video frames. The origin of generated content can be verified using the Gemini application, Google Search, and Chrome.
For custom avatar generation, users can create digital representations using their own verified voice and likeness. Voice-altering tools for third-party audio remain restricted while safety testing continues.
Partner integrations and ecosystem adoption
Several creative platforms have integrated Gemini Omni Flash into their production stacks:
- Adobe Firefly: Integrates Omni Flash for prompt-driven video editing and texture replacement.
- Figma Weave: Uses scene extensions and reference branching to enable teams to direct storyboard sequences collaboratively.
- Runway: Employs Omni Flash to let creators transition between video drafts, prompt iterations, and camera modifications.
- GMI Cloud: Provides centralized access to Omni Flash for educational and scientific visualization pipelines.
API pricing and availability
Gemini Omni 1.1 Flash is accessible through multiple channels:
- Developers: Available via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.
- Subscribers: Rolling out globally to Google AI Plus, Pro, and Ultra members inside the Gemini app and Google Flow.
- Content creators: Available at no cost inside YouTube Shorts and the YouTube Create application.

| Feature / Tier | Standard Video Generation (720p) | Rapid Draft Preview (360p) | High Resolution (1080p / 4K) |
|---|---|---|---|
| Max Duration per Call | 10 seconds | 10 seconds | 10 seconds (up to 40s extension) |
| Context Memory Window | 10 seconds prior footage | 10 seconds prior footage | 10 seconds prior footage |
| Speed & Latency | Standard generation throughput | Up to 60% faster generation | Production render pipeline |
| Relative Compute Cost | Base rate | 1/3 of standard cost | High-resolution tier |
| Watermarking & Provenance | SynthID embedded | SynthID embedded | SynthID embedded |







It’s important to ensure that video content generation incorporates robust protective features to prevent misuse.