The Rundown AI homepage

Independent tool overview

Nemotron 3.5 Lightning at a glance

NVIDIA Nemotron 3.5 Lightning is an open-weight 30B mixture-of-experts model with 3B active parameters, built for fast, high-volume execution inside long-running AI agents.

Visit the official Nemotron 3.5 Lightning site ↗
Nemotron 3.5 Lightning product preview
Model size
30B total, 3B active
Context window
Up to 1M tokens
License
OpenMDW 1.1
Modalities
Text input and output
Best for
High-volume agent execution
Release date
August 11, 2026

Overview

What Nemotron 3.5 Lightning is

Nemotron 3.5 Lightning is the execution-focused member of NVIDIA's Nemotron family. Its hybrid Mamba-2, mixture-of-experts, and attention architecture activates about 3B of its 30B parameters per token, targeting lower latency for tool calls, validation, coding work, and delegated sub-agent tasks.

NVIDIA releases BF16 and NVFP4 checkpoints, speculative-decoding options, training data, and recipes under the OpenMDW 1.1 license. The model supports up to a 1 million-token context and can be served locally or in a data center using tools such as vLLM, SGLang, TensorRT-LLM, Ollama, llama.cpp, and LM Studio.

This is an infrastructure component, not a consumer chatbot. Teams need suitable NVIDIA hardware or a hosted endpoint, an agent harness, evaluation datasets, security controls, and routing logic that sends harder planning work to a more capable model when needed.

Use cases

Who Nemotron 3.5 Lightning is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Long-running agents

Handle repeated tool calls, result validation, formatting, and other execution-heavy steps.

Sub-agent workers

Use a smaller local or hosted model for delegated tasks while reserving frontier models for planning.

Private deployment

Run open weights on controlled infrastructure when data locality and operational control matter.

Specialized fine-tuning

Adapt the model to a narrow workflow with LoRA, supervised fine-tuning, or reinforcement learning.

Capabilities

Core Nemotron 3.5 Lightning features

1

Sparse MoE architecture

Uses 30B total parameters while activating about 3B per token to improve serving efficiency.

2

One-million-token context

Supports long prompts and agent histories on NVIDIA-validated configurations.

3

Speculative decoding

Ships with multi-token prediction plus DSpark and DFlash draft-model options for different serving loads.

4

BF16 and NVFP4 checkpoints

Offers a standard checkpoint and an NVIDIA-optimized quantized version for supported hardware.

5

Open customization stack

Includes weights, data, recipes, and compatibility with NVIDIA's fine-tuning and evaluation tools.

6

Tool-use orientation

Training and benchmarks emphasize coding, terminal work, instruction following, and agentic execution.

7

Flexible serving

Can run through NVIDIA's endpoint, partner providers, or self-hosted inference frameworks.

Process

How the Nemotron 3.5 Lightning workflow works

  1. Step 1

    Define the task tier

    Identify repetitive agent steps that need speed and do not require the strongest available reasoning model.

  2. Step 2

    Choose a checkpoint

    Select BF16 or NVFP4 based on hardware support, memory, quality requirements, and serving software.

  3. Step 3

    Deploy and integrate

    Serve the model through a compatible runtime and connect it to an agent harness with explicit tool schemas.

  4. Step 4

    Evaluate real workflows

    Measure task success, tool-call correctness, latency, throughput, and cost on representative production traces.

  5. Step 5

    Route and govern

    Escalate uncertain or complex tasks to a stronger model and constrain tools, credentials, data, and action permissions.

Cost

Nemotron 3.5 Lightning pricing and free plan

The model weights are downloadable under OpenMDW 1.1, and NVIDIA offers a free prototype endpoint. Production cost depends on self-hosted GPU infrastructure or the selected partner endpoint.

Model download

No model fee listed

Download checkpoints from NVIDIA's official Hugging Face repository under the OpenMDW 1.1 license.

  • Self-hosting required
  • Infrastructure and operations are separate
  • Review license terms before deployment

NVIDIA prototype endpoint

Free

Evaluate the model interactively or through NVIDIA's free API endpoint.

  • API key required
  • Intended for prototyping
  • Production limits may differ

Production deployment

Varies

Run on owned NVIDIA hardware or a compatible hosted inference provider.

  • GPU, storage, and network costs apply
  • Provider pricing varies
  • Throughput depends on checkpoint and runtime

Pricing checked . Check current pricing at the source ↗

Assessment

Nemotron 3.5 Lightning strengths and limitations

Where it stands out

  • Efficient 3B-active architecture for high-volume execution
  • Very long supported context window
  • Open weights, data, and training recipes
  • Multiple checkpoints and speculative-decoding strategies
  • Can run from a single supported high-memory GPU configuration
  • Broad compatibility with local and data-center serving tools

What to consider

  • Requires technical deployment and suitable GPU infrastructure
  • Text-only model with six listed natural-language families
  • A small execution model will not match frontier models on every planning or reasoning task
  • One-million-token operation can demand substantial memory and careful serving configuration
  • NVIDIA's benchmark results should be reproduced on the buyer's own workload
  • Agents still need tool permission controls, validation, monitoring, and escalation
  • Training data freshness is limited by the published 2025-2026 cutoffs

Compare

Nemotron 3.5 Lightning alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Consumer

Nemotron 3 Super

Choose Nemotron 3 Super when stronger reasoning is more important than Lightning's execution efficiency.

Explore Nemotron 3 Super

Consumer

Qwen3.5 Small

Consider Qwen3.5 Small for another compact open-model family with different deployment tradeoffs.

Explore Qwen3.5 Small

Questions

Nemotron 3.5 Lightning FAQs

What is Nemotron 3.5 Lightning?

It is NVIDIA's open 30B mixture-of-experts text model with about 3B active parameters, optimized for fast execution steps in autonomous agents.

Is Nemotron 3.5 Lightning open source?

NVIDIA describes it as an open model and releases weights, data, and recipes under the OpenMDW 1.1 license. Teams should review that license for their intended use.

How much does Nemotron 3.5 Lightning cost?

NVIDIA lists a free prototype endpoint and downloadable model weights. Production cost comes from GPU infrastructure or a hosted inference provider rather than a universal per-token price.

Can Nemotron 3.5 Lightning run locally?

Yes, on supported systems. NVIDIA lists single-GPU deployment on DGX Spark or H100 and also documents local workflows for hardware such as GeForce RTX 5090, with configuration-dependent limits.

What is the context window?

The official model card lists up to a 1 million-token context on validated configurations.

Which languages are supported?

The model card lists English and coding languages, plus Spanish, French, German, Italian, and Japanese.

Should it replace a frontier model in an agent?

Usually not for every step. A practical design routes routine execution to Lightning and escalates complex planning, ambiguous decisions, or high-risk actions to a stronger model or a human.

Bottom line

Our Nemotron 3.5 Lightning verdict

Nemotron 3.5 Lightning is compelling for teams that can operate open-model infrastructure and want a fast worker model inside an agent stack. Its value should be proven with task-level evaluations and routing, not assumed from token speed or vendor benchmarks alone.

Visit Nemotron 3.5 Lightning website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.