AI video researchers
Study counterfactual video editing and whether generative models can preserve causal and physical consistency.
Independent tool overview
VOID is an open-source research model from Netflix and INSAIT that removes an object from video and attempts to rewrite the physical interactions caused by that object.
Visit the official VOID site ↗
Overview
VOID, short for Video Object and Interaction Deletion, tackles a harder problem than ordinary video inpainting. When an object is removed, it also tries to regenerate consequences such as another object falling, a collision not occurring, or a trajectory changing.
The workflow combines user-selected points, SAM-based segmentation, vision-language reasoning, interaction-aware quadmasks, and a CogVideoX-based diffusion model. An optional second pass uses warped noise to improve temporal consistency when the first output contains object-morphing artifacts.
This is a research implementation rather than a polished commercial editor. Netflix provides code and two model checkpoints under Apache 2.0, plus a community-hosted browser demo, but local use requires technical setup, a Gemini API key for the supplied mask pipeline, and a GPU with at least 40GB of VRAM.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Study counterfactual video editing and whether generative models can preserve causal and physical consistency.
Prototype object removal where shadows, contact, collisions, or dependent motion must also be regenerated.
Inspect, adapt, or benchmark a fully released two-pass video-inpainting pipeline and checkpoints.
Test difficult removal shots that go beyond filling the pixels directly behind an object.
Capabilities
Attempts to remove both the selected object and its causal effects on the rest of the scene.
Uses vision-language reasoning and segmentation to separate the removed object, affected regions, overlaps, and protected background.
Runs a base inpainting pass, with an optional warped-noise refinement pass for improved shape and temporal stability.
Provides separate Pass 1 and Pass 2 checkpoint files through the official Hugging Face repository.
Includes point selection, automatic mask generation, and a manual quadmask editor for correcting problem frames.
Process
Step 1
Load a source video and click points on the object that should be removed.
Step 2
Run SAM segmentation and the Gemini-assisted reasoning pipeline to identify both the object and interaction-affected regions.
Step 3
Provide the source, quadmask, and a prompt describing the clean remaining background to generate the counterfactual video.
Step 4
Check whether the object, downstream interactions, and unaffected portions of the video remain coherent.
Step 5
Correct the mask manually or run Pass 2 when morphing and temporal instability remain.
Cost
VOID has no subscription price. The official code and checkpoints are released under Apache 2.0, while users supply their own compute and any third-party API usage.
Free
Download and run the official implementation under the Apache 2.0 license.
Usage-based costs
Pay for the GPU, storage, hosting, and external model APIs used by your deployment.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Content Creator
Choose Runway Aleph for a commercially hosted, prompt-driven workflow for editing existing footage.
Explore Runway Aleph →Content Creator
Choose Runway Creative for a broader web-based video generation and editing suite with less technical setup.
Explore Runway Creative →Agents
Consider MiniMax's hosted AI products when managed access matters more than running an open research pipeline locally.
Explore MiniMax Agent →Questions
VOID is a video object and interaction deletion framework from Netflix and INSAIT researchers. It removes a selected object and attempts to regenerate the scene as if that object and its physical influence had not been present.
The official code and model checkpoints are available under Apache 2.0 at no license fee. You still pay for the GPU infrastructure, storage, engineering work, and any third-party API use needed to run it.
A community-hosted Hugging Face Space is linked from the official repository. The main release, however, is designed for technical users who can run the model and its mask pipeline themselves.
Ordinary inpainting usually fills the region behind an object and may remove effects such as shadows. VOID also targets downstream interactions, such as making a supported object fall or preventing a collision after the selected object is removed.
The official quick start recommends a GPU with at least 40GB of VRAM, such as an NVIDIA A100. The authors report training on eight A100 80GB GPUs, but that training requirement is separate from inference.
It is better viewed as a research starting point. Teams should test output consistency, security, third-party dependencies, compute cost, and failure handling before using it in production workflows.
Bottom line
VOID is a compelling open research option for physics-aware object removal, especially when deleting an object should also change what happens around it. Its value is in experimentation and specialized R&D; creators wanting a fast, supported editor will be better served by a hosted commercial tool.
Visit VOID website ↗
Veo 3.1 Lite - Google's new budget-friendly video generation mode

MAI Image 2 - Microsoft AI's image model with upgraded photorealism and creativity

Phota Studio - Phota's personalized photo editing and generation model that preserves your identity across edits

Avatar V - HeyGen's new AI avatar model that generates studio-quality, hyperrealistic videos from a 15-second clip

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.