Model Summary
Total Parameters 30B (3B active) Architecture MoE - Mamba-2 + MoE + Attention hybrid Context Length Up to 1M tokens Single-GPU Deployment 1× DGX Spark (GB10) or 1× H100 Supported Hardware NVIDIA Blackwell (DGX Spark / GB10, GB200, GeForce RTX 5090); NVIDIA Hopper (H100, H200); NVIDIA Ampere via W4A16 Supported Languages English (and coding languages), Spanish, French, German, Italian, Japanese Speculative Decoding DSpark for low-concurrency Data Centre and DGX Spark Workflows — Read more below, also provided are MTP (Multi-Token Prediction) and DFlash Recommended Sampling Temperature 1.0, Top_P 0.95 Best For Long-running autonomous agents, sub-agent workhorse deployments, and efficient local inference on personal hardware License OpenMDW License Agreement, version 1.1 Release Date August 11, 2026
Model Overview
Model Developer: NVIDIA Corporation
Model Dates: December 2025 - May 2026
Data Freshness:
The pre-training data has a cutoff date of September 2025.
The post-training data has a cutoff date of May 2026.
What is Nemotron?
NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.
... continue reading