The Rundown AI homepage

Independent tool overview

Veo 3.1 at a glance

Veo 3.1 is Google's current Veo video-generation family for short cinematic clips with native audio, portrait or landscape output, image guidance, frame control, and video extension.

Visit the official Veo 3.1 site ↗
Veo 3.1 product preview
Current model
Veo 3.1
Original Veo 3 status
Gemini API models retired June 30, 2026
Clip length
4, 6, or 8 seconds per generation
Resolution
720p, 1080p, and supported 4K output
Aspect ratios
16:9 landscape and 9:16 portrait
Audio
Native sound, ambience, effects, and dialogue cues
Access
Google Flow, Gemini API, and Vertex AI
Reviewed
August 30, 2026

Overview

What Veo 3.1 is

Veo 3.1 is the current successor to Veo 3, Google's first broadly available video model with native synchronized audio. It creates 4-, 6-, or 8-second clips from text or images in landscape or portrait format, with 720p, 1080p, and—on supported variants—4K output. It can use first and last frames, reference images for a person, character, product, or style, and extensions of prior Veo clips, making it useful for concept shots, ads, social creative, storyboards, product visuals, and cinematic experiments.

The original Gemini API models `veo-3.0-generate-001` and `veo-3.0-fast-generate-001` were retired on June 30, 2026, so new integrations should use Veo 3.1 or Google's newer default video model where appropriate. Google now recommends Gemini Omni Flash for general conversational video generation and editing, while Veo 3.1 remains relevant for scene extension, explicit last-frame control, reference-image direction, and existing Veo workflows. Access and capabilities differ across Google Flow, the Gemini API, and Vertex AI, so users should confirm the active model, resolution, credit cost, and launch stage before each production workflow.

Use cases

Who Veo 3.1 is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Short cinematic concepts

Turn a detailed shot description into a polished visual-and-audio clip for previsualization, mood exploration, and creative pitching.

Product and character consistency

Use reference images to guide the appearance of a product, character, scene, or style across a generated shot.

Social and advertising variations

Generate portrait or landscape creative, test several prompts and variants, and use Fast or Lite models when iteration volume matters more than maximum quality.

Controlled transitions

Define both a starting and ending frame when a shot must move between two planned visual states.

Extending Veo scenes

Continue compatible Veo-generated clips in seven-second increments for longer sequences while preserving the preceding action and visual context.

Capabilities

Core Veo 3.1 features

1

Native video and audio generation

Creates visuals together with ambient sound, effects, music cues, and spoken dialogue described in the prompt.

2

Text-to-video

Generates a shot from a detailed scene description including subject, action, setting, camera, lighting, style, and audio direction.

3

Image-to-video

Animates a supplied starting image while following prompt guidance for motion, camera behavior, atmosphere, and sound.

4

First-and-last-frame control

Interpolates a shot between defined opening and closing images for more deliberate composition and transitions.

5

Reference images

Uses up to three images to guide the identity or appearance of a person, character, product, or visual ingredient in supported Veo 3.1 variants.

6

Video extension

Extends eligible Veo-generated videos by seven seconds at a time, up to 20 extensions and a combined maximum described as 148 seconds.

7

Portrait and landscape output

Supports 16:9 and 9:16 aspect ratios for widescreen, mobile, and social-video use cases.

8

Multiple quality and speed variants

Offers Lite, Fast, and Standard or Quality choices across access surfaces, trading cost and generation speed against resolution and output quality.

9

720p, 1080p, and 4K

Supports higher-resolution output on applicable models; 1080p and 4K generations require eight-second duration, while Lite tops out at 1080p.

10

Google Flow workflow

Provides a creative interface for prompting, ingredients, frames, extensions, scene building, asset management, and direct YouTube handoff.

11

Programmatic API access

Runs asynchronously through the Gemini API and is also available in Vertex AI with access, quota, model IDs, and launch stages that differ by platform.

12

SynthID watermarking

Embeds an invisible watermark in Veo-generated video so Google tools can help identify AI-generated content.

13

Safety and memorization checks

Screens prompts and outputs and applies checks intended to reduce harmful content, privacy, copyright, and memorization risks.

Process

How the Veo 3.1 workflow works

  1. Step 1

    Choose the access surface

    Use Flow for hands-on creative work, the Gemini API for application workflows, or Vertex AI when Google Cloud governance, quota, and deployment controls are required.

  2. Step 2

    Pick the current model

    Select Veo 3.1 Lite, Fast, or Quality based on feature compatibility, resolution, iteration speed, and cost; do not start new work on retired Veo 3.0 API IDs.

  3. Step 3

    Design one shot

    Write a prompt with subject, action, environment, camera framing and movement, lighting, style, timing, dialogue, sound effects, and ambience rather than trying to direct an entire film at once.

  4. Step 4

    Add visual constraints

    Supply a starting image, first and last frames, or up to three reference images when composition or subject identity must be more controlled.

  5. Step 5

    Generate and select

    Create several candidates, inspect frame consistency, physics, text, faces, dialogue, lip sync, audio, and brand accuracy, then keep only outputs that pass review.

  6. Step 6

    Extend or edit the sequence

    Use compatible extensions or Flow scene tools to build continuity, then finish timing, transitions, color, audio, captions, disclosure, and rights review in an editor.

Cost

Veo 3.1 pricing and free plan

Veo pricing depends on access method. Google Flow uses per-generation credits, while the Gemini API bills by successful output second with separate Lite, Fast, and Standard rates. Flow gives non-subscribers 50 daily credits; paid Google AI plans include monthly credits, and extra credits can be purchased in supported regions. Rates and credit costs can change, so verify the selected model in Flow or the official API pricing page before a large run.

Google Flow free credits

50 credits/day

Daily credits for non-subscribers to try Veo 3.1 Lite, Fast, and Quality generations in Flow.

  • Unused free credits do not roll over
  • Lite costs 10 credits per generation for non-Ultra users
  • Fast costs 20 credits per generation for non-Ultra users
  • Quality costs 100 credits per generation
  • One request may create multiple separately charged generations

Veo 3.1 Lite API

$0.05/sec 720p · $0.08/sec 1080p

Lowest-cost Gemini API Veo variant for short video with audio.

  • Approximately $0.40 for an 8-second 720p clip
  • Approximately $0.64 for an 8-second 1080p clip
  • 4K is not supported
  • No Gemini API free tier

Veo 3.1 Fast API

$0.10–$0.30/sec

Faster paid API generation for higher-volume creative and application workflows.

  • $0.10/sec at 720p
  • $0.12/sec at 1080p
  • $0.30/sec at 4K
  • An 8-second clip costs about $0.80, $0.96, or $2.40 respectively

Veo 3.1 Standard API

$0.40/sec · $0.60/sec 4K

Highest-priced Gemini API Veo quality tier.

  • $0.40/sec at 720p or 1080p
  • $0.60/sec at 4K
  • An 8-second clip costs about $3.20 or $4.80 at 4K
  • Google charges only for successfully generated output when audio processing or safety blocks prevent completion

Pricing checked . Check current pricing at the source ↗

Assessment

Veo 3.1 strengths and limitations

Where it stands out

  • Generates visuals and audio together instead of requiring a separate sound-design pass for every concept.
  • First-and-last-frame control and reference images provide more direct composition guidance than text prompting alone.
  • Portrait and landscape support makes the same model family useful for social, advertising, and cinematic formats.
  • Lite, Fast, and Quality choices let teams separate cheap iteration from higher-cost final candidates.
  • Veo extension can turn a successful short shot into a longer compatible sequence.
  • Flow provides a creative interface while Gemini API and Vertex AI support programmatic and governed deployments.
  • Published per-second API pricing makes cost modeling more transparent than credit-only systems.
  • SynthID and safety checks provide a built-in provenance and risk-reduction layer.
  • Google's broader image, video, Gemini, and Flow ecosystem supports an end-to-end concept workflow.

What to consider

  • The original Veo 3.0 Gemini API models are retired; integrations that did not migrate by June 30, 2026 will fail and need a current model ID.
  • Google now recommends Gemini Omni Flash as the default for general video generation, leaving Veo 3.1 positioned around specific controls and legacy workflows.
  • Gemini API Veo 3.1 model IDs are documented as preview, so availability, behavior, quotas, and compatibility can change before a stable release.
  • Each base generation is only four, six, or eight seconds, and higher resolutions or reference-image workflows generally require eight-second output.
  • Video extension is limited to compatible Veo-generated clips, works at 720p, and depends on source assets retained or referenced within the platform's time window.
  • Generated API videos are stored for two days, so applications must download and manage files promptly.
  • Native dialogue and short spoken audio can still be incoherent or poorly synchronized; Google describes this as an active area of development.
  • Safety filters and audio processing can block a generation even when a prompt appears acceptable, which affects predictable batch completion.
  • Generation latency can range from roughly 11 seconds to six minutes during peak periods according to the API documentation.
  • English is the fully supported prompt language; results from other languages may be inconsistent or unevaluated.
  • Person-generation controls and allowed inputs vary by region, particularly in the EU, UK, Switzerland, and MENA.
  • Visual continuity, object permanence, physics, hands, typography, logos, dialogue, and character identity still require inspection across every frame.
  • Repeated generation can become expensive because many candidates may be needed for one usable shot, and Flow requests can produce multiple charged generations.
  • Reference images, uploads, music, voices, likenesses, brands, and commercial use still require appropriate rights and consent; safety filters do not provide legal clearance.
  • SynthID helps identify Google-generated media but is not a complete substitute for visible disclosure, provenance records, and publisher policies.

Compare

Veo 3.1 alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Content Creator

Runway Gen-4.5

Runway Gen-4.5 is a strong creative-production alternative with its own motion, consistency, editing, and filmmaking workflow; compare controllability, audio, clip economics, and editing needs.

Explore Runway Gen-4.5

Content Creator

Hailuo AI

Hailuo AI is worth comparing for fast text- and image-to-video iteration, especially when broad consumer access and lower-cost experimentation matter.

Explore Hailuo AI

Content Creator

LTX Studio

LTX Studio is better suited when the priority is planning and assembling a larger story with scripts, shots, characters, and timelines rather than calling one video model directly.

Explore LTX Studio

Questions

Veo 3.1 FAQs

Is Veo 3 still available?

The original `veo-3.0-generate-001` and Fast Gemini API models were retired on June 30, 2026. The Veo product continues as Veo 3.1 through Flow, Gemini API preview models, and supported Vertex AI model versions.

How much does Veo 3.1 cost?

Gemini API prices range from $0.05 per second for 720p Lite to $0.60 per second for 4K Standard. Google Flow uses credits, including 50 free daily credits for non-subscribers and monthly allocations with Google AI plans.

Does Veo 3.1 generate audio?

Yes. Audio is always on for the documented Veo 3.1 API variants and can include dialogue, ambience, sound effects, and other cues, though speech quality and synchronization can still fail.

How long are Veo 3.1 videos?

A base generation can be four, six, or eight seconds. Certain features and higher resolutions require eight seconds. Compatible Veo clips can be extended in seven-second increments under documented limits.

Can Veo 3.1 make vertical video?

Yes. It supports both 16:9 landscape and 9:16 portrait output, although feature compatibility can vary by model and access surface.

Can I keep a character or product consistent?

Veo 3.1 can use up to three reference images to guide a person, character, product, or style. This improves direction but does not guarantee perfect identity, logo, or detail consistency.

What is the difference between Veo 3.1 Lite, Fast, and Quality?

Lite is the cheapest, Fast prioritizes speed and volume, and Standard or Quality targets stronger output at a much higher cost. Resolution, extension, reference, and Flow feature support differ, so check the active model before generating.

Should developers use Veo 3.1 or Gemini Omni Flash?

Google recommends Gemini Omni Flash as the default for general video generation and conversational editing. Veo 3.1 remains useful when you need Veo scene extension, explicit first/last-frame control, reference-image direction, or compatibility with an existing Veo pipeline.

Are Veo videos watermarked?

Yes. Google says Veo-generated videos contain an invisible SynthID watermark. Flow may also apply or offer visible watermarking depending on plan and region.

Bottom line

Our Veo 3.1 verdict

Veo 3.1 remains one of Google's most capable shot-generation options when native audio and explicit creative controls matter. Reference images, first-and-last-frame interpolation, portrait output, extensions, and clear API rates make it practical for disciplined production experiments. The key is to treat the model as a shot generator, not a finished-film button: plan for multiple attempts, manual editing, rights review, audio cleanup, provenance, and model migration. New API work should use current Veo 3.1 IDs—or evaluate Gemini Omni Flash first—rather than the retired Veo 3.0 models this page originally covered.

Visit Veo 3.1 website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.