3D world-model researchers
Study long-horizon, camera-controlled generation, spatial memory, drift correction, and generative reconstruction.
Independent tool overview
NVIDIA Lyra 2.0 is an open research implementation that extends a starting image along user-defined camera paths, then reconstructs the generated views as a 3D Gaussian Splatting scene.
Visit the official Lyra 2.0 site ↗
Overview
Lyra 2.0 is a research model from NVIDIA's Spatial Intelligence Lab for creating persistent, explorable 3D worlds. It starts with one image, a camera trajectory, and optional text captions, generates a long camera-controlled video, and lifts the result into a Gaussian Splatting scene for real-time rendering.
The core research problem is consistency over long exploration paths. Lyra stores per-frame 3D geometry to retrieve relevant earlier views when the camera revisits an area, and its training exposes the model to degraded histories so it learns to correct drift rather than compound it.
NVIDIA released the paper, model weights, inference code, training code, and an interactive GUI in 2026. This is not a hosted consumer product or pay-per-use API: users install and run the stack on their own Linux and NVIDIA GPU infrastructure.
The source code uses Apache 2.0, but the model weights use NVIDIA's Internal Scientific Research and Development Model License. Commercial deployment is therefore not automatically granted; organizations needing other rights must request a custom license from NVIDIA.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Study long-horizon, camera-controlled generation, spatial memory, drift correction, and generative reconstruction.
Explore research workflows that turn generated camera paths into scenes suitable for real-time rendering and simulated navigation.
Generate a view sequence and reconstruct it as a Gaussian point cloud for downstream rendering research.
Run or fine-tune the full pipeline when Linux, CUDA, large-memory NVIDIA GPUs, and ML engineering support are already available.
Capabilities
Extends a scene across many autoregressive video chunks while trying to preserve appearance and structure over large viewpoint changes.
Uses stored per-frame geometry to retrieve earlier views and establish correspondences when previously seen areas become relevant again.
Supports preset trajectories, custom pose sequences, and interactive trajectory editing through the released GUI.
Per-chunk captions can describe content for newly revealed regions, while the GUI can also use a vision-language model to rewrite a short prompt from scene context.
A second pipeline stage estimates pose and depth from the generated video and writes a reconstructed Gaussian point cloud.
An optional four-step distilled LoRA cuts the documented 80-frame generation time substantially, with a tradeoff in prompt following and repetition.
NVIDIA provides weights, inference code, training code, sample inputs, and a local client-server GUI.
Process
Step 1
Install the Linux, CUDA, Conda, PyTorch, Flash Attention, and compiled-extension stack documented by NVIDIA.
Step 2
Obtain the Lyra 2.0 model files from NVIDIA's official Hugging Face repository and review the model license.
Step 3
Provide a starting image, define a camera trajectory, and add captions for content that should appear as the view expands.
Step 4
Run the full-quality sampler or the faster DMD option, then inspect the sequence for hallucinations, drift, repetition, and path errors.
Step 5
Lift the accepted video into a Gaussian Splatting scene, review the PLY and rendered trajectory, and validate it before any simulation use.
Cost
NVIDIA provides the Lyra 2.0 code and research weights without a subscription fee, but users pay for their own high-end GPU compute, storage, and engineering time. The source and model have different licenses, and commercial rights may require a custom agreement.
No software fee
Downloadable code and model weights for eligible research use.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Content Creator
Choose World Labs Marble for a more productized workflow for creating persistent 3D worlds from images, video, or text.
Explore World Labs Marble →Miscellaneous
Choose HY-World 2.0 for another open-source world-model stack that accepts text, image, or video inputs.
Explore HY-World 2.0 →Miscellaneous
Choose HunyuanWorld-Mirror for Tencent's open approach to generating 3D worlds from text, image, or video.
Explore HunyuanWorld-Mirror →Questions
Lyra 2.0 is a research framework that turns a starting image, camera trajectory, and optional captions into an explorable view sequence and a reconstructed 3D Gaussian Splatting scene.
The code and research weights can be downloaded without a subscription fee, but users provide their own GPU compute and must follow the separate source and model licenses.
NVIDIA officially tested the stack on H100 GPUs with Ubuntu and CUDA. Other configurations may work, but the dependency stack and memory requirements make this a high-end research workload rather than a typical consumer installation.
The pipeline writes generated videos and, after reconstruction, a PLY Gaussian point cloud plus a rendered flythrough video.
Yes. NVIDIA released a Linux GUI for seeding a scene from an image, authoring camera paths, prompting new content, extending the world, and reverting an unsatisfactory generation.
Do not assume the research weights are cleared for commercial use. The code is Apache 2.0, but the models use NVIDIA's Internal Scientific Research and Development Model License; contact NVIDIA for custom licensing.
Bottom line
Lyra 2.0 is a substantial open research release for teams studying persistent generative 3D worlds and Gaussian Splatting reconstruction. Its released GUI and training stack make it unusually complete for research, but the hardware burden, model-license restrictions, and generation artifacts keep it far from a plug-and-play commercial tool.
Visit Lyra 2.0 website ↗
Harrier - Microsoft Bing’s SOTA, open-source embedding model for search and RAG grounding

Gemini 3.1 Flash TTS - Google's new speech model with inline tags for voice direction across 70+ languages

Google Edge Eloquent - Google AI Edge Eloquent — Free voice dictation app that turns messy speech into polished text, runs fully offline, no subscription, no usage caps

Deep Max - Exa's new SOTA agentic search tool

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.