The Rundown AI homepage

Independent tool overview

Gemini Omni Flash at a glance

Gemini Omni Flash 1.1 is Google's generally available developer model for generating and conversationally editing short videos with native audio through the Gemini API.

Visit the official Gemini Omni Flash site ↗
Gemini Omni Flash product preview
Current model
Gemini Omni Flash 1.1
API model ID
gemini-omni-1.1-flash
Availability
Generally available on the paid Gemini API tier
Output length
3–10 seconds at 24 FPS
Resolution
360p, 720p, upscaled 1080p or upscaled 4K
Aspect ratios
16:9 or 9:16
Context window
1,048,576 tokens
720p video price
Approximately $0.10 per output second

Overview

What Gemini Omni Flash is

The product introduced as Gemini Omni is now available as Gemini Omni Flash 1.1, with the stable API model ID gemini-omni-1.1-flash. It accepts text, images and short videos, generates video with audio, and supports iterative natural-language edits through Google's Interactions API.

Its clearest fit is short, API-driven creative work rather than long-form production. Outputs run from 3 to 10 seconds at 24 FPS in landscape or portrait formats. The model can return 360p or 720p video and upscale to 1080p or 4K, but those higher resolutions are explicitly upscaled rather than native.

Use cases

Who Gemini Omni Flash is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Short marketing clips

Developers creating 3–10 second product, social or campaign video variants with generated sound.

Conversational video editing

Teams that want to generate a clip, then refine specific elements through follow-up natural-language instructions.

Image animation

Creators bringing product shots, illustrations or photographs to life from an image plus a motion prompt.

Transitions and extensions

Workflows that need first-to-last-frame interpolation or a short continuation appended to an existing clip.

Video features inside an app

Product teams embedding video generation and editing through Google's SDKs or REST API.

Capabilities

Core Gemini Omni Flash features

1

Multimodal video generation

Processes text, image, audio and video context while producing video with an automatically generated audio track.

2

Conversational editing

The Interactions API can carry a generated video's context into later turns so the user can request focused changes while preserving the rest of the scene.

3

Image-to-video

Animates a supplied reference image using a text description of motion, camera behavior, mood and sound.

4

First and last frame interpolation

Accepts two images as boundary frames and generates the transition between them.

5

Video extension

Appends a 3–10 second continuation to a model-generated or uploaded video, with additional image references supported.

6

Resolution and orientation control

Supports 360p and 720p output, upscaled 1080p or 4K, and either 16:9 landscape or 9:16 portrait framing.

7

Native audio

Generates sound with the video and accepts prompting for dialogue, ambience, music and timed audio events.

8

SynthID provenance

Every generated video includes Google's invisible SynthID watermark for programmatic provenance verification.

Process

How the Gemini Omni Flash workflow works

  1. Step 1

    Choose the generation mode

    Start from text, an image, an uploaded video or a previous Omni interaction depending on whether the goal is creation, editing, interpolation or extension.

  2. Step 2

    Specify the shot

    Describe scene, subject, camera movement, lighting, mood, timing and audio. Explicitly request a continuous shot when unwanted cuts would be a problem.

  3. Step 3

    Set delivery properties

    Choose 16:9 or 9:16, a 3–10 second duration and the required resolution. Use URI delivery for larger outputs instead of moving large base64 payloads.

  4. Step 4

    Generate through the Interactions API

    Call gemini-omni-1.1-flash and retain the interaction ID when later conversational edits or model-generated extensions are planned.

  5. Step 5

    Refine with narrow instructions

    Ask for one clear change at a time and add 'Keep everything else the same' when visual continuity matters.

  6. Step 6

    Review before publishing

    Inspect every frame, audio cue, spoken line, logo and text element; disclose synthetic media where required and keep human approval in the release path.

Cost

Gemini Omni Flash pricing and free plan

Gemini Omni Flash is available to developers on the paid Gemini API tier. At Standard rates, Google bills input at $1.50 per million tokens, text and thinking output at $9 per million, and video output at $17.50 per million. A 720p output uses 5,792 tokens per second, or approximately $0.10 per second.

Google AI Studio prototyping

$0 to try where available

Google links to AI Studio for testing, but Gemini Omni Flash does not have a free production API tier.

  • Availability and rate limits can vary by region and account
  • Do not treat AI Studio access as free production API capacity
  • Production use requires a paid Gemini API project

Gemini API Standard

$1.50 input / $9 text output / $17.50 video output per 1M tokens

Usage-based paid access to the stable Gemini Omni Flash 1.1 model.

  • Input covers text, image, video and audio tokens
  • 720p video is billed at 5,792 tokens per second
  • Effective 720p price is approximately $0.10 per output second
  • Paid-tier data is not used to improve Google products

Modeled 720p examples

About $0.30 for 3s or $1.01 for 10s

Approximate output-only math based on Google's published $0.10-per-second effective 720p rate; input and text or thinking tokens add cost.

  • These examples are calculations, not quoted package prices
  • Generation failures, retries and extra turns can change total spend
  • Verify token usage for 360p, 1080p and 4K rather than assuming the 720p rate

Pricing checked . Check current pricing at the source ↗

Assessment

Gemini Omni Flash strengths and limitations

Where it stands out

  • Combines generation, editing, interpolation and extension in one multimodal API model.
  • Follow-up editing through an interaction ID is more natural than rebuilding every prompt from scratch.
  • Generated audio can reduce the number of separate tools needed for a finished short clip.
  • Portrait and landscape output cover the two most common social and presentation formats.
  • The stable model offers a 1,048,576-token context window and generally available paid API access.
  • SynthID provides a built-in machine-detectable provenance signal.

What to consider

  • Output is limited to 3–10 seconds, so longer narratives require multiple generations and careful editing.
  • Uploaded videos for editing or extension must be 10 seconds or less, except for certain model-generated multi-turn extensions.
  • 1080p and 4K are upscaled outputs rather than native-generation resolutions.
  • Only 16:9 and 9:16 aspect ratios are documented.
  • There is no free Gemini API tier for this model.
  • Voice editing and uploaded audio references are not supported in the current API.
  • Extension only appends to the end of a clip; it cannot prepend footage or extend the middle.
  • Uploaded video editing and extension have regional restrictions in the EEA, Switzerland and the United Kingdom.
  • The model cannot reason reliably across multiple video references, and reference-video audio is ignored.
  • System instructions, sampling controls, stop sequences and dedicated negative-prompt fields are not supported.
  • English is fully supported, while Google's documentation says other languages have not been evaluated.
  • Generated motion, identity, text, speech, sound continuity and physical behavior still require frame-by-frame human review.
  • SynthID is invisible to viewers, so it does not replace visible disclosure or compliance with platform and advertising rules.

Compare

Gemini Omni Flash alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Content Creator

Grok Imagine

Choose Grok Imagine for image and short-video generation within xAI's consumer and API ecosystem.

Explore Grok Imagine

Content Creator

Runway Gen-4.5

Choose Runway Gen-4.5 for a creator-facing video workflow with editing tools and a broader production interface.

Explore Runway Gen-4.5

Content Creator

Kling 3.0

Choose Kling 3.0 to compare another short-form video model with image, motion and audio capabilities.

Explore Kling 3.0

Content Creator

Hailuo AI

Choose Hailuo AI for a web-based generative-video workflow aimed at creators rather than API-first implementation.

Explore Hailuo AI

Questions

Gemini Omni Flash FAQs

What happened to Gemini Omni?

Google's original Gemini Omni announcement evolved into Gemini Omni Flash. The current stable developer model is Gemini Omni Flash 1.1 with the API ID gemini-omni-1.1-flash.

Is Gemini Omni Flash generally available?

Yes. Google lists Gemini Omni Flash 1.1 as generally available to developers on the paid tier of the Gemini API.

How long can Gemini Omni Flash videos be?

A generated output can be 3 to 10 seconds long at 24 FPS. Uploaded source videos for editing and extension are generally limited to 10 seconds.

Can Gemini Omni Flash generate audio?

Yes. It generates video with an audio track and accepts prompts for sound, music and dialogue. It does not currently support voice editing or uploading a separate audio reference.

Does Gemini Omni Flash support 4K?

Yes, as an upscaled output option. Google's documentation labels both 1080p and 4K as upscaled, while 720p is the default output resolution.

How much does Gemini Omni Flash cost?

Standard pricing is $1.50 per million input tokens, $9 per million text and thinking output tokens, and $17.50 per million video output tokens. Google estimates 720p video at about $0.10 per output second.

Is there a free Gemini Omni Flash API tier?

No. Google's pricing table lists the free API tier as unavailable. AI Studio may provide a way to test the model where supported, but that is not free production API capacity.

Can Gemini Omni Flash edit an existing video?

Yes. It supports uploaded-video editing and conversational edits to prior generations, subject to source-length, content and regional restrictions.

Can it extend a video?

Yes. It can append a 3–10 second continuation. Extension works only at the end of a clip, and uploaded-video extension is unavailable in some regions.

Are Gemini Omni Flash videos watermarked?

Yes. Google says every generated video contains an invisible SynthID watermark that can be detected programmatically for provenance verification.

Bottom line

Our Gemini Omni Flash verdict

Gemini Omni Flash 1.1 is a strong fit for developers who need short video generation, native audio and iterative edits inside one Google API workflow. Its main constraints are equally clear: 3–10 second outputs, paid-only API use, narrow aspect-ratio choices, upscaled high-resolution modes and several editing and regional limitations. It is most useful as a controllable clip engine with human review, not an unattended replacement for long-form video production.

Visit Gemini Omni Flash website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.