The Rundown AI homepage

Independent tool overview

Nemotron 3 at a glance

NVIDIA Nemotron 3 is a family of downloadable models and hosted NIM endpoints for agentic reasoning, coding, tool use, long-context work, multimodal understanding, and retrieval.

Visit the official Nemotron 3 site ↗
Nemotron 3 product preview
Developer
NVIDIA
Core sizes
Nano, Super, and Ultra
Largest context
Up to 1 million tokens
Deployment
Download, NIM, or partner endpoint

Overview

What Nemotron 3 is

Nemotron 3 is NVIDIA's open-model family for developers building specialized AI agents and enterprise inference systems. Its core Nano, Super, and Ultra language models use hybrid Mamba-Transformer mixture-of-experts architectures that activate only part of their total parameter count, aiming to balance reasoning quality, long context, and serving efficiency.

The family has expanded beyond the original three sizes. Nemotron 3 Nano Omni adds image, video, audio, and text understanding, while Nemotron 3 Embed supports multilingual semantic search and retrieval. Weights, model cards, training resources, and deployment instructions are available through NVIDIA and Hugging Face, with free prototype endpoints and production options through NIM, partner endpoints, or self-hosted infrastructure.

Use cases

Who Nemotron 3 is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Agentic enterprise systems

Build reasoning, planning, coding, tool-calling, and retrieval agents under an organization-controlled deployment.

High-throughput inference

Use mixture-of-experts models that activate a smaller parameter subset per token than their total model size.

Multimodal document and media analysis

Use Nano Omni for text output from images, video, audio, documents, OCR, and graphical interfaces.

Capabilities

Core Nemotron 3 features

1

Three core reasoning sizes

Nano has 30B total and about 3.5B active parameters, Super has 120B total and 12B active, and Ultra has 550B total and 55B active.

2

Configurable reasoning

The chat template can enable or disable thinking mode to trade additional reasoning work for speed and cost.

3

Long context

Super and Ultra support context windows up to one million tokens for large repositories, documents, and multi-step agent histories.

4

Tool use and coding

The core models are trained for instruction following, code, planning, retrieval, and tool-calling workflows.

5

Nano Omni

A roughly 31B-total, 3B-active model that understands text, images, audio, and video and returns text with up to 256K context.

6

Open deployment paths

Provides downloadable checkpoints plus integrations for vLLM, SGLang, TensorRT-LLM, Ollama, llama.cpp, NVIDIA NIM, and partner services.

Process

How the Nemotron 3 workflow works

  1. Step 1

    Match the model to the workload

    Choose Nano for efficiency, Super for stronger long-context agents, Ultra for demanding reasoning, Omni for multimodal input, or Embed for retrieval.

  2. Step 2

    Check license and hardware

    Review the exact model card, governing license, supported precision, memory requirement, and deployment geography before committing.

  3. Step 3

    Prototype with an endpoint

    Test prompts, tool schemas, context behavior, and output quality through NVIDIA's free trial endpoint or another provider.

  4. Step 4

    Evaluate on private tasks

    Measure accuracy, latency, throughput, safety, and total cost against the organization's real documents and agent workflows.

  5. Step 5

    Deploy and monitor

    Select NIM, a partner endpoint, or self-hosted inference, then add application-level permissions, observability, evaluation, and fallback controls.

Cost

Nemotron 3 pricing and free plan

Nemotron 3 model weights can be downloaded under the license attached to each checkpoint, and NVIDIA offers free prototype endpoints. Production cost depends on GPU infrastructure, NVIDIA NIM licensing or entitlement, cloud and partner endpoint rates, model size, precision, context length, and traffic.

Downloadable weights

No model download fee

Run eligible Nemotron 3 checkpoints on infrastructure you operate.

  • Model-specific NVIDIA or OpenMDW license
  • Compute and operations paid separately
  • Several precision variants available

NVIDIA prototype endpoint

Free trial endpoint

Prototype through an OpenAI-compatible hosted API subject to NVIDIA's trial terms and limits.

  • API key required
  • Not a universal production pricing commitment
  • Availability varies by model

Production deployment

Infrastructure or provider pricing varies

Use NVIDIA NIM, partner endpoints, or self-managed GPUs for production inference.

  • Hardware requirements rise sharply by model size
  • Cloud GPU and endpoint rates vary
  • Long context increases memory and compute demand

Pricing checked . Check current pricing at the source ↗

Assessment

Nemotron 3 strengths and limitations

Where it stands out

  • Provides multiple efficiency and quality points instead of forcing one model size on every workload
  • Publishes model weights, detailed cards, training resources, and deployment integrations
  • Supports configurable reasoning, tool use, coding, retrieval, and long-context agent workflows
  • Fits NVIDIA's optimized serving stack while remaining available to several open inference frameworks

What to consider

  • The name covers several materially different models, context limits, modalities, licenses, and hardware profiles
  • Super and Ultra require substantial GPU capacity; NVIDIA lists eight H100 80GB GPUs as the minimum BF16 configuration for each
  • A free prototype endpoint does not make production inference free, and NVIDIA does not publish one universal hosted price for the family
  • Benchmarks and vendor evaluations do not replace testing on private data, tool schemas, safety requirements, latency targets, and total cost

Compare

Nemotron 3 alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Coding

Meta Llama Models

Choose Meta Llama when ecosystem breadth and broad third-party deployment support are the main priorities.

Explore Meta Llama Models

Business Operations

DeepSeek

Choose DeepSeek for another family of downloadable reasoning and coding models with separate hosted access.

Explore DeepSeek

Consumer

Qwen3.5 Small

Choose Qwen3.5 Small when compact open models and more modest local hardware targets are more important than NVIDIA-stack optimization.

Explore Qwen3.5 Small

Questions

Nemotron 3 FAQs

Is Nemotron 3 one model?

No. It is a family that includes Nano, Super, Ultra, Nano Omni, embedding models, multiple precisions, and base or post-trained variants.

What are the main Nemotron 3 sizes?

Nano is 30B total with about 3.5B active parameters, Super is 120B total with 12B active, and Ultra is 550B total with 55B active.

Can Nemotron 3 process images or video?

The core Nano, Super, and Ultra chat models are text models. Nemotron 3 Nano Omni accepts video, audio, image, and text inputs and returns text.

Is Nemotron 3 free?

Weights can be downloaded without a per-model purchase fee and NVIDIA offers free trial endpoints, but self-hosting, production endpoints, NIM, cloud GPUs, and operations can create substantial costs.

Can Nemotron 3 be used commercially?

NVIDIA model cards mark major Nemotron 3 checkpoints as ready for commercial use, but the exact governing license differs by checkpoint and must be reviewed before deployment.

Bottom line

Our Nemotron 3 verdict

Nemotron 3 is particularly attractive to teams that want inspectable, deployable agent models and already operate NVIDIA infrastructure. Nano is the practical starting point, while Super and Ultra should be selected only after workload-specific testing justifies their much larger hardware footprint; teams should treat each checkpoint as a separate technical and licensing decision.

Visit Nemotron 3 website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.