Multi-agent systems
Handle complex subtasks, long histories, and large tool libraries inside orchestrated agent workflows.
Independent tool overview
NVIDIA Nemotron 3 Super is an open-weight 120B hybrid mixture-of-experts model built for long-context reasoning, tool use, coding, RAG, and agentic workloads.
Visit the official Nemotron 3 Super site ↗
Overview
NVIDIA Nemotron 3 Super is a 120-billion-parameter language model designed for production agent systems rather than a consumer chatbot. Its sparse architecture activates about 12 billion parameters per token, combining Mamba-2, mixture-of-experts, attention, and multi-token prediction to reduce the cost of long reasoning and high-volume inference.
The model stands out for a native context window of up to one million tokens, configurable reasoning, tool calling, and openly released weights, data, and training recipes. Those advantages come with meaningful engineering requirements: the standard BF16 checkpoint needs substantial accelerator capacity, maximum context raises memory demands sharply, and teams remain responsible for evaluation, safeguards, orchestration, and operating costs.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Handle complex subtasks, long histories, and large tool libraries inside orchestrated agent workflows.
Load large codebases or long technical context for code generation, debugging, and software-agent tasks.
Reason across extensive reports, knowledge bases, tickets, and retrieved evidence without aggressively fragmenting context.
Use sparse activation and optimized precision to improve throughput for workloads such as IT ticket handling.
Run open weights in controlled infrastructure and fine-tune or adapt the model for specialized domains.
Capabilities
Nemotron 3 Super contains 120B total parameters while activating roughly 12B for each forward pass.
The model supports up to 1M tokens for very large documents, repositories, histories, and retrieval sets.
Developers can enable or disable thinking through the chat template to trade extra reasoning for speed and cost.
The post-trained model is designed for function calling and agent workflows that interact with external systems.
Mamba-2 layers improve long-sequence efficiency while selected attention layers preserve global reasoning capacity.
A sparse expert design activates specialized capacity without paying the full dense-model inference cost per token.
The architecture predicts multiple future tokens and includes a speculative-decoding head intended to increase generation throughput.
NVIDIA publishes BF16, FP8, and Blackwell-oriented NVFP4 variants for different deployment hardware.
Weights can be served with supported frameworks or packaged through NVIDIA NIM for on-premises and cloud deployments.
NVIDIA publishes training data, post-training methods, reinforcement-learning environments, evaluation recipes, and technical reports alongside the model.
The model card lists English, French, German, Italian, Japanese, Spanish, and Chinese as supported languages.
Process
Step 1
Choose BF16, FP8, or NVFP4 based on target hardware, accuracy requirements, and memory constraints.
Step 2
Confirm the Nemotron Open Model License and any hosted-service terms fit the intended commercial or research use.
Step 3
Test prompts, reasoning modes, tool schemas, and long-context behavior with NVIDIA's trial endpoint or another supported provider.
Step 4
Measure task accuracy, tool-call validity, hallucinations, latency, throughput, and cost using real workload samples.
Step 5
Run with a supported framework or NIM container, including NVIDIA's required reasoning parser and recommended generation settings.
Step 6
Wrap the model with access controls, input and output validation, observability, retries, human review, and safe tool permissions.
Step 7
Use the smallest useful context window and enable thinking only where quality gains justify the added latency and tokens.
Cost
The downloadable weights do not carry a per-token model fee, but every practical path has infrastructure or provider costs. NVIDIA does not publish one universal Nemotron 3 Super price because the model is available through self-hosting, an NVIDIA API trial, and multiple inference providers.
No model download fee
Download and deploy under the NVIDIA Nemotron Open Model License.
Trial access
Prototype against NVIDIA's hosted, OpenAI-compatible endpoint under API trial terms.
Provider-specific
Use Nemotron 3 Super through participating clouds and inference platforms.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Coding
Meta's Llama family offers a broader ecosystem of open-weight sizes and deployments for teams that do not need Nemotron's specific long-context agent focus.
Explore Meta Llama Models →Business Operations
DeepSeek provides open reasoning models and hosted access that may be easier to evaluate for general coding and reasoning workloads.
Explore DeepSeek →Coding
Ollama is a practical local model runner for teams exploring smaller open models before committing to Nemotron 3 Super's infrastructure footprint.
Explore Ollama →Questions
Nemotron 3 Super is NVIDIA's open-weight 120B language model for agentic reasoning, tool use, coding, RAG, and high-volume automation. Its sparse architecture activates about 12B parameters per token.
NVIDIA describes the release as open and publishes weights, datasets, and recipes. Usage is governed by the NVIDIA Nemotron Open Model License, so it is more precise to call it an open-weight model and review that license for your use case.
The model supports up to one million tokens. The default Hugging Face configuration is 256K because longer contexts require substantially more memory.
Yes, but not on an ordinary laptop in its standard form. NVIDIA lists eight H100 80GB GPUs for the BF16 checkpoint, while the NVFP4 variant can run on a single B200 or DGX Spark according to the official model card.
Yes. It is post-trained for agentic workflows and tool use, but each integration still needs schema validation, permission controls, retries, and application-level safety checks.
Yes. The official chat template exposes an enable_thinking flag so developers can disable extended reasoning for simpler or more latency-sensitive requests.
The weights have no per-token download fee under the model license. Actual cost depends on self-hosted GPU infrastructure or the pricing of the selected cloud or inference provider; NVIDIA Build also offers trial API access.
The model card lists English, French, German, Italian, Japanese, Spanish, and Chinese as supported post-training languages.
No model should receive broad autonomous permissions solely because it scores well on tool-use benchmarks. Use least-privilege credentials, sandboxing, allowlisted tools, output validation, monitoring, and human approval for consequential actions.
Bottom line
Nemotron 3 Super is a compelling option for engineering teams that need open weights, very long context, and efficient reasoning inside sophisticated agent systems. It is best evaluated as infrastructure, not a chatbot: the model can lower inference cost relative to dense systems, but production value depends on hardware, serving expertise, task-specific evaluation, and strong controls around tools and data.
Visit Nemotron 3 Super website ↗
Copilot Cowork - Microsoft's Anthropic-powered AI for running multi-step tasks across M365 apps

Mistral 4 Small - Mistral's open-source 119B MoE model combining reasoning, coding, and vision in one package

GPT 5.4 - OpenAI's new flagship reasoning model with native computer use and 1M context

GPT-5.4 mini & nano - OpenAI's fast, cheap small models built for coding agents and subagent workflows

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.