Short videos with synchronized audio
Create compact ads, social concepts, music-led clips, and dialogue scenes with generated sound or an uploaded audio track.
Independent tool overview
Wan 2.6 is Alibaba's still-available family of image and video generation models for text-to-video, image-to-video, reference-led video, synchronized audio, multi-shot narratives, and image creation. It is useful for existing workflows, but Wan 2.7 now succeeds it for image and segmented video tasks, while the preview Wan 3.0 Video is Alibaba's newest unified video model.
Visit the official Wan 2.6 site ↗
Overview
Wan 2.6 is a model family rather than one universal endpoint. Alibaba exposes separate model IDs for text-to-video, image-to-video, reference-to-video, text-to-image, and image generation or editing. The creator site packages these capabilities into a friendlier interface, while Model Studio provides regional APIs for production use.
The 2.6 video models generate 720p or 1080p MP4 video at 30 fps, with clips up to 15 seconds for text- and image-led generation. They can create synchronized sound automatically or use supplied audio, and their multi-shot mode can switch camera framing while trying to preserve the main subject.
Wan 2.6 remains listed and priced in Alibaba Cloud documentation, so this is not a historical archive. However, buyers should treat it as an older generation: Wan 2.7 expands control and video continuation, and Wan 3.0 consolidates multiple video tasks into one preview model with up to 30-second output and richer reference inputs.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Create compact ads, social concepts, music-led clips, and dialogue scenes with generated sound or an uploaded audio track.
Produce a sequence of changing shots in one clip while asking the model to preserve the central subject across transitions.
Turn a first-frame image into a 720p or 1080p clip while controlling action, camera movement, audio, and duration.
Maintain a known Wan 2.6 endpoint while evaluating whether Wan 2.7 or the preview Wan 3.0 model improves the workflow.
Capabilities
Generates a video from a written scene description, with optional custom audio and automatic dubbing.
Uses an image as the opening frame and animates it into a short clip with optional synchronized audio.
Uses image or video references to preserve one or more subjects while placing them into a newly directed scene.
Supports shot changes within one generation. In Wan 2.6, developers enable multi-shot mode together with intelligent prompt rewriting.
Can automatically produce a soundtrack or accept MP3 or WAV input for dialogue, voiceover, music, and sound synchronized to the visuals.
The Wan 2.6 image endpoints support text-to-image, image-led creation, multiple-reference fusion, style transfer, and controls for framing and lighting.
Provides resolution, duration, random seed, prompt rewriting, and optional visible watermark controls, with endpoints scoped by deployment region.
Process
Step 1
Decide whether the job is text-to-video, first-frame image-to-video, reference-to-video, text-to-image, or image editing before selecting a model ID.
Step 2
For a new video integration, test Wan 3.0 alongside 2.6; for image work, compare Wan 2.7 Image. Keep 2.6 when it offers the right quality, price, or migration risk.
Step 3
Describe the subject, action, scene, lighting, camera movement, dialogue, and shot timing. For 2.6 multi-shot video, set shot_type to multi and enable prompt rewriting.
Step 4
Choose 720p or 1080p, a 2-to-15-second duration where supported, and whether to use generated audio, uploaded audio, or a silent Flash output.
Step 5
Model Studio video calls run asynchronously. Poll the task until completion and copy the output to durable storage because the generated download URL expires after 24 hours.
Step 6
Check continuity, faces, hands, dialogue timing, brand assets, safety, and usage rights before publishing or scaling a campaign.
Cost
Alibaba Cloud Model Studio bills Wan 2.6 video by each successfully generated second and image output per generated image. International standard video is $0.10 per second at 720p or $0.15 per second at 1080p; image generation is $0.03 per image. Flash video modes can cut the per-second rate, especially for silent output.
$0.10–$0.15 per second
International text-to-video, image-to-video, and reference-to-video output.
$0.025–$0.075 per second
Lower-cost image-to-video and reference-to-video variants with audio or silent output.
$0.03 per image
International text-to-image and image generation or editing endpoints.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Content Creator
Google's video model is a strong comparison for cinematic generation, prompt adherence, and synchronized audio.
Explore Veo 3.1 →Content Creator
A direct short-form image and video generation alternative with its own motion and audio workflow.
Explore Kling 2.6 →Content Creator
A creator-oriented video platform with an integrated editing workflow and production-facing tools.
Explore Runway Gen-4.5 →Content Creator
Another accessible image-to-video and text-to-video option for rapid social and creative experiments.
Explore Hailuo AI →Questions
Yes. Alibaba Cloud still documents and prices Wan 2.6 image and video endpoints. It is active, although Wan 2.7 and Wan 3.0 are newer choices for several workflows.
Wan 2.6 uses separate endpoints for text-to-video, image-to-video, and reference-to-video and generally tops out at 15 seconds. Wan 3.0 is a unified preview model that covers those tasks, supports richer reference inputs, adaptive aspect ratios, editing and extension, and clips up to 30 seconds.
Yes. The supported 2.6 video models can automatically dub a clip or synchronize an uploaded audio file. Flash image-to-video and reference-to-video variants can also generate silent output at a lower API rate.
As of August 28, 2026, international standard output is $0.10 per second at 720p and $0.15 per second at 1080p. Flash variants range from $0.025 per second for silent 720p to $0.075 per second for 1080p with audio.
The international text-to-video and first-frame image-to-video endpoints support integer durations from 2 to 15 seconds. Reference-to-video output supports up to 10 seconds.
Yes. Alibaba provides wan2.6-t2i for text-to-image and wan2.6-image for generation and editing with text and image inputs.
New users in the international Model Studio scope receive time-limited quotas: 50 generated video seconds for listed models and 50 images for the 2.6 image endpoints, valid for 90 days after activation. The Wan creator site's consumer credits are separate.
Bottom line
Wan 2.6 is still practical for short, audio-enabled image and video generation, particularly when a team already has a working Model Studio integration or benefits from its Flash pricing. New projects should benchmark it against Wan 2.7 and Wan 3.0 before committing: 2.6 remains active, but it is no longer Alibaba's most capable video architecture.
Visit Wan 2.6 website ↗
Kling AI: Enables you to generate imaginative videos and images with cutting-edge generative models.

ChatGPT Images - OpenAI's new AI image generation system with upgraded editing, text rendering, quality, and dedicated Images creation space

Suno AI: Turns text prompts into studio-quality songs using cutting-edge music generation models.

Ray3 Modify - Edit and reimagine videos with precise keyframe and character reference controls

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.