The Rundown AI homepage

Independent tool overview

Wan2.2-S2V at a glance

Wan2.2-S2V-14B is Alibaba Tongyi Lab's open-weight speech-to-video model. It animates a reference image from audio and an optional prompt, supports half- and full-body performance at 480p or 720p, and can follow a pose video, but local inference requires at least 80GB of GPU memory.

Visit the official Wan2.2-S2V site ↗
Wan2.2-S2V product preview
Best for
Technical creators and research teams making consented dialogue, singing, performance, or character animation
Inputs
Reference image, audio, optional text prompt, and optional pose video
Outputs
Audio-synchronized 480p or 720p character video
Model size
14B parameters
Local hardware
At least 80GB of GPU VRAM for the documented single-GPU workflow
Length
Automatically follows input audio unless the number of clips is limited
License
Apache 2.0 for the official model and repository, subject to published use restrictions and third-party rights
Last reviewed
September 1, 2026

Overview

What Wan2.2-S2V is

Wan2.2-S2V takes a still character image, an audio track, and an optional scene prompt and generates a synchronized performance video. Unlike a simple talking-head lip-sync tool, the research release targets facial expression, body movement, camera motion, environment, dialogue, singing, and broader cinematic staging.

The official implementation supports 480p and 720p output. Video length follows the input audio unless the user limits the number of generated clips, and an optional pose video can guide body motion. The repository also integrates CosyVoice for generating source speech, although that adds dependencies and a second identity-and-rights review.

This is a 14B-parameter self-hosted model rather than a polished editing subscription. The first-party command is documented for a GPU with at least 80GB of VRAM, while multi-GPU inference uses FSDP and DeepSpeed Ulysses. An official Hugging Face Space is running for experimentation, but shared demos are not a production service or privacy guarantee.

The code and model are published under Apache 2.0, and the model card says the authors claim no rights over generated content. That does not grant rights to a person's image, voice, performance, music, character, trademark, or source footage. Because the model can make a real person appear to speak, sing, or act, explicit consent, disclosure, provenance, and human review are essential.

Use cases

Who Wan2.2-S2V is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Cinematic character performances

Teams with licensed character art and audio that want more body and camera motion than basic portrait lip sync.

Dialogue and singing prototypes

Filmmakers and music-video creators testing a consented performance before committing to final production.

Pose-guided animation

Technical artists who can provide a pose video to shape body movement alongside the audio.

Research and custom pipelines

Developers with high-memory GPU infrastructure who need open weights and control over the local generation stack.

Capabilities

Core Wan2.2-S2V features

1

Image-plus-audio generation

Animates one reference character image in sync with an input speech, song, or performance track.

2

Half- and full-body motion

The model is designed to generate facial expression, gestures, posture, and broader body movement.

3

Prompted scene control

An optional text prompt describes setting, action, emotion, weather, camera movement, and visual style.

4

Pose-video guidance

An optional motion reference can drive a more specific pose sequence during generation.

5

Audio-length adaptation

The pipeline sets output length from the supplied audio unless a shorter clip count is requested.

6

480p and 720p output

The official S2V model supports both resolutions while following the source image's aspect ratio.

7

Optional CosyVoice synthesis

The repository can generate the audio track through CosyVoice before animating the image.

8

Open inference stack

Official code, weights, single-GPU commands, and distributed inference are available for local deployment.

Process

How the Wan2.2-S2V workflow works

  1. Step 1

    Secure all rights and consent

    Document permission for the image, likeness, voice, performance, music, character, pose footage, script, commercial use, distribution, and synthetic transformation.

  2. Step 2

    Choose the safest identity strategy

    Prefer fictional, owned, or clearly licensed characters. Avoid real people when a designed identity meets the creative need.

  3. Step 3

    Prepare the reference image

    Use a clear, appropriately framed half- or full-body image with no private background details, bystanders, or unauthorized marks.

  4. Step 4

    Clean and review the audio

    Confirm speaker and music rights, remove private conversation, and listen for words that could create a false endorsement or harmful claim.

  5. Step 5

    Write a bounded scene prompt

    Describe action, emotion, setting, camera, and lighting without changing identity or inventing consequential conduct.

  6. Step 6

    Preview a short clip

    Use the clip-count control to test sync, identity, hands, movement, clothing, and scene consistency before a long run.

  7. Step 7

    Add pose guidance only when licensed

    Use an owned or consented motion reference and check that the generated body movement remains appropriate.

  8. Step 8

    Review frame by frame

    Check lip sync, face, hands, anatomy, camera motion, text, logos, cultural details, and unintended sexualization or violence.

  9. Step 9

    Disclose and preserve provenance

    Label the video as AI-generated or altered, retain consent records and source provenance, and remove it quickly if rights are withdrawn or misuse appears.

Cost

Wan2.2-S2V pricing and free plan

Wan2.2-S2V is an open-weight model with no per-generation software fee in the official repository. Real cost comes from an 80GB-or-larger GPU, storage, engineering, retries, and moderation. The official Hugging Face Space can be used as a shared demo while available, not as guaranteed free production capacity.

Official Hugging Face Space

Shared demo

A running public interface for testing the model without a local installation.

  • Availability, queues, limits, retention, and hardware can change
  • Do not upload sensitive or unlicensed identity media
  • Not a service-level-backed production endpoint

Self-hosted model

No model usage fee

Download the official code and weights and operate the model on your own infrastructure.

  • Apache 2.0 license obligations and published use restrictions apply
  • The documented single-GPU S2V command requires at least 80GB VRAM
  • GPU time, storage, bandwidth, engineering, security, and review are separate costs

Distributed production deployment

Infrastructure-dependent

Scale generation with multiple GPUs and the repository's distributed inference approach.

  • Requires PyTorch FSDP and DeepSpeed Ulysses expertise
  • Longer audio, 720p output, previews, and retries increase compute
  • Budget for moderation, provenance, access control, monitoring, and incident response

Pricing checked . Check current pricing at the source ↗

Assessment

Wan2.2-S2V strengths and limitations

Where it stands out

  • Generates more than facial lip sync by modeling full-body action, environment, emotion, and camera movement
  • Accepts speech, singing, and other performance audio
  • Optional pose guidance gives creators a stronger motion reference
  • Output length can follow the source audio automatically
  • Supports both 480p and 720p in the official S2V pipeline
  • Open weights and inference code enable local control and customization
  • Apache 2.0 is more permissive for commercial engineering than many noncommercial research-model licenses
  • An official running demo lowers the barrier to an initial quality test

What to consider

  • The official single-GPU workflow requires at least 80GB of VRAM, well beyond most consumer graphics cards
  • Multi-GPU deployment adds distributed-systems complexity, cost, debugging, and security work
  • This is a technical model release rather than a polished timeline editor, managed API, or supported creative suite
  • The official repository's S2V ComfyUI and Diffusers integrations are still shown as unchecked items in its public task list
  • Shared Hugging Face and ModelScope demos can queue, change, log uploads, impose limits, or become unavailable
  • Apache 2.0 covers the released software and model; it does not grant rights to input images, people, voices, music, scripts, characters, brands, or pose footage
  • The model card's statement that the authors claim no rights over outputs does not guarantee that a generated video is lawful, unique, noninfringing, or commercially safe
  • Making a real person appear to say, sing, endorse, confess, or perform something can cause fraud, defamation, harassment, privacy, publicity, labor, election, or consumer-protection harm
  • Use only real-person likenesses and voices covered by explicit, use-specific, revocable consent; parental or legal authority is required for minors
  • AI-generated or materially altered people should be clearly disclosed wherever viewers could believe the footage is authentic
  • Do not use generated video as evidence, identity verification, news footage, political testimony, medical communication, or proof of consent
  • Audio may contain copyrighted music, a cloned voice, private conversation, or third-party performance rights that the model cannot validate
  • Optional CosyVoice generation creates a second synthetic-media and voice-consent risk
  • Pose footage has its own performer, choreography, privacy, and recording rights
  • Identity, lip sync, hands, teeth, eyes, clothing, props, and anatomy can drift across frames
  • Longer output can accumulate visual errors and motion inconsistency even when the duration follows the audio
  • Generated camera work, environmental motion, and body movement may not match the exact creative direction
  • 720p generation is not equivalent to final broadcast quality and may still require editing, cleanup, grading, sound work, and upscaling
  • The project's benchmark is author-reported and selected metrics do not capture every production defect, rights issue, or audience preference
  • The research page says training used millions of human-centric videos from large datasets and manual curation but does not provide a per-item rights audit for downstream buyers
  • Local deployment shifts moderation, storage, access control, provenance, data deletion, and abuse response entirely to the operator
  • Prompt extension through cloud services can send image or text context to an additional provider and add cost and privacy exposure

Compare

Wan2.2-S2V alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Content Creator

Hedra

Choose Hedra for a more accessible hosted character-video workflow with integrated voice and creative tools instead of an 80GB self-hosted model.

Explore Hedra

Content Creator

Hailuo AI

Choose Hailuo AI for a managed image-to-video product when ease of generation matters more than open weights and pose-guided control.

Explore Hailuo AI

Content Creator

Vidu

Choose Vidu for a hosted video platform and API with reference-based generation, current commercial plans, and less infrastructure work.

Explore Vidu

Questions

Wan2.2-S2V FAQs

What is Wan2.2-S2V?

Wan2.2-S2V-14B is an open-weight speech-to-video model from Alibaba Tongyi Lab that animates a reference image from audio, an optional prompt, and optional pose guidance.

Is Wan2.2-S2V still available?

Yes. The official project page, GitHub repository, model weights, and running Hugging Face Space were available when checked on September 1, 2026.

What does S2V mean?

It means speech-to-video. The model turns an image and audio performance into a synchronized character video rather than generating speech itself by default.

What inputs does Wan2.2-S2V need?

A reference image and audio track are the core inputs. A text prompt can direct the scene, and an optional pose video can guide body movement.

What resolution does it generate?

The official pipeline supports 480p and 720p, with output aspect ratio following the reference image.

How much VRAM does Wan2.2-S2V require?

The documented single-GPU S2V command requires at least 80GB of VRAM. Multiple GPUs are supported through a more complex distributed setup.

How long can the video be?

Without a clip-count limit, the repository says video length automatically adapts to the input audio. Longer runs require more compute and can accumulate visual drift.

Can it animate a full body?

Yes. The research release targets both half- and full-body characters, including facial expressions, gestures, movement, and scene-level camera behavior.

Can it generate singing videos?

Yes. Singing and performance are highlighted use cases, but the creator must hold rights to the voice, recording, music, lyrics, image, and depicted identity.

Does Wan2.2-S2V include text to speech?

The repository can optionally use CosyVoice to synthesize audio before video generation. That requires extra dependencies and separate voice-cloning consent and disclosure controls.

Is Wan2.2-S2V free?

There is no official per-generation model fee for self-hosting, and a shared Hugging Face demo is running. Hardware, cloud GPU time, storage, engineering, and review can be substantial.

Can I use Wan2.2-S2V commercially?

The official model card lists Apache 2.0 and says generated content belongs to the user, but commercial use still requires compliance with the full license, published restrictions, and every third-party image, voice, music, character, trademark, privacy, and publicity right.

Can I animate a real person?

Only with explicit, specific, revocable authorization for the exact synthetic use. Do not create deceptive speech, fake endorsements, fabricated evidence, sexual content, political manipulation, or other harmful impersonation.

Should Wan2.2-S2V videos be labeled?

Yes whenever a reasonable viewer could think the depicted event or performance is real. Preserve source and consent records and use platform provenance tools where available.

Bottom line

Our Wan2.2-S2V verdict

Wan2.2-S2V is a technically impressive open speech-to-video model for teams that need broad character motion, promptable scenes, audio synchronization, and pose control. It is not an instant replacement for a hosted creator app: the 80GB VRAM requirement, engineering burden, frame-level defects, and deepfake risk are substantial. Use it for fictional or explicitly consented identities, budget for human post-production, and treat rights and disclosure as core pipeline requirements.

Visit Wan2.2-S2V website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.