Connected multi-shot scenes
Generate several related shots in one clip while asking the model to preserve characters, setting, lighting, style, and sound.
Independent tool overview
LTX-2.5 is Lightricks' 22B open-weight audio-video model for multi-shot clips, synchronized sound, and controllable local or API generation. It is unusually flexible, but its license is not unrestricted and the hosted 2.5 API omits several editing endpoints.
Visit the official LTX-2.5 site ↗
Overview
LTX-2.5 is a downloadable world model focused on synchronized video and audio generation from text, images, video, and audio inputs. Its major upgrade is native multi-shot generation: one run can create connected shots while attempting to preserve characters, environments, lighting, visual style, and voice across cuts.
Creators can use LTX-2.5 through LTX's hosted API and playground or download the weights for Python, ComfyUI, and Diffusers workflows. The open-weight path is the more differentiated option because it supports local execution, fine-tuning, and deployment on infrastructure you control. The model card provides both distilled and trainable 22B transformer checkpoints plus separate text encoder, video decoder, audio, duration, and upscaling components.
The model is free for eligible organizations under LTX's Community License, but it should not be described as unrestricted open source. Entities with at least $10 million in annual revenue need a paid agreement for commercial use, and transferring fine-tunes can carry additional conditions. Teams should review the binding license before building it into a product.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Generate several related shots in one clip while asking the model to preserve characters, setting, lighting, style, and sound.
Run the downloadable model on controlled infrastructure when customization, data locality, or avoiding per-generation API fees matters.
Fine-tune the trainable checkpoint and build LoRA or IC-LoRA workflows for a specific visual domain.
Use the hosted API for text-to-video, image-to-video, or audio-to-video without managing model infrastructure.
Capabilities
A single generation can contain multiple connected shots instead of only one continuous camera take.
The model generates visual and audio streams together and can also drive video from an input audio track.
LTX replaced the previous reconstruction stage with a diffusion decoder intended to improve faces, texture, text, motion, and pixel-level detail.
The model can choose an appropriate clip length from the prompt instead of requiring a fixed duration.
Fast supports more resolutions and up to 4K, while Pro targets higher fidelity but currently tops out at 1080p.
Downloadable distilled and trainable checkpoints support local inference, fine-tuning, ComfyUI, Python pipelines, and an early Diffusers integration.
Process
Step 1
Use the API for quick integration; choose the weights when you need local control, fine-tuning, or predictable infrastructure ownership.
Step 2
Start from text or an image, or use audio-to-video when an existing voice or soundtrack should drive the result.
Step 3
Choose Fast or Pro, select a supported resolution and frame rate, and either set duration or allow automatic duration.
Step 4
Check character continuity, prompt adherence, text, hands, motion artifacts, and audio synchronization before using the clip.
Step 5
Confirm the Community License covers your organization, revenue, distribution, and fine-tune plans before production deployment.
Cost
The hosted LTX-2.5 API is billed per generated second. Text-to-video and image-to-video use the same rates; audio-to-video is billed from input-audio duration. Self-hosting has no LTX usage fee for qualifying users, but infrastructure costs apply and organizations at or above $10M annual revenue need a commercial agreement.
$0.09–$0.30 per second
The speed-oriented hosted model with the broadest resolution support.
$0.12–$0.17 per second
The higher-fidelity hosted model with a lower maximum resolution.
No LTX usage fee for eligible users
Local and production use under the LTX-2.x Community License for entities below the revenue threshold.
Custom
Required for commercial use by entities with at least $10M in annual revenue.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Content Creator
Choose LTX Studio when you want a visual filmmaking workspace instead of integrating or self-hosting the underlying model.
Explore LTX Studio →Content Creator
Choose Runway for a polished hosted creative suite with editing and production workflows.
Explore Runway Gen-4.5 →Content Creator
Consider Sora when you prefer OpenAI's consumer video experience over an open-weight deployment.
Explore Sora 2 →Content Creator
Compare Wan's current generation for another modern image-and-video model ecosystem.
Explore Wan 3.0 →Questions
LTX-2.5 has downloadable open weights and code, but the model uses the LTX-2.x Community License rather than an unrestricted license. Commercial entities with at least $10M in annual revenue need a paid agreement.
Fast costs $0.09 per second at 720p, $0.13 at 1080p, $0.19 at 1440p, and $0.30 at 4K. Pro costs $0.12 at 720p and $0.17 at 1080p.
Yes. LTX-2.5 jointly generates synchronized video and audio, and its API supports audio-to-video as well as text-to-video and image-to-video.
Yes. LTX publishes weights and workflows for its Python pipelines, ComfyUI, and Diffusers. The 22B model and its supporting components still require capable GPU hardware, storage, and technical setup.
Fast is optimized for speed and supports up to 4K in the API. Pro targets higher fidelity but currently supports only 720p and 1080p and costs more per second at those resolutions.
The model launch describes editing capabilities, but the current LTX-2.5 hosted API supports text-to-video, image-to-video, and audio-to-video only. Retake, Extend, and Reframe remain on LTX-2.3 Pro.
Bottom line
LTX-2.5 is one of the more practical choices for teams that need both a hosted video API and downloadable weights. Multi-shot continuity, synchronized audio, fine-tuning, and local deployment are meaningful advantages. The tradeoffs are a demanding self-hosted stack, incomplete 2.5 editing endpoints, and a revenue-based license that larger businesses must price before committing.
Visit LTX-2.5 website ↗
Seedance 2.5 - ByteDance’s new SOTA video model with 30-second generations

Pika Audio - Pika’s four-model AI audio family spanning music, speech, SFX, and soundtracks

FLUX 3 - BFL’s new multimodal visual intelligence model with 20-second generations, now in early access

Black Forest Labs' video upscaler that regenerates clips at native 4K

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.