Skip to content
Tech News
← Back to articles

Nvidia Nemotron 3.5 Lightning

read original more articles
Why This Matters

NVIDIA Nemotron 3.5 Lightning is a cutting-edge large language model designed for high efficiency and accuracy in AI applications, supporting long context lengths and optimized for deployment on NVIDIA hardware. Its open architecture and specialized features make it a significant advancement for autonomous agents, data centers, and personalized AI solutions, offering both performance and flexibility for the tech industry and consumers.

Key Takeaways

Model Summary

Total Parameters 30B (3B active) Architecture MoE - Mamba-2 + MoE + Attention hybrid Context Length Up to 1M tokens Single-GPU Deployment 1× DGX Spark (GB10) or 1× H100 Supported Hardware NVIDIA Blackwell (DGX Spark / GB10, GB200, GeForce RTX 5090); NVIDIA Hopper (H100, H200); NVIDIA Ampere via W4A16 Supported Languages English (and coding languages), Spanish, French, German, Italian, Japanese Speculative Decoding DSpark for low-concurrency Data Centre and DGX Spark Workflows — Read more below, also provided are MTP (Multi-Token Prediction) and DFlash Recommended Sampling Temperature 1.0, Top_P 0.95 Best For Long-running autonomous agents, sub-agent workhorse deployments, and efficient local inference on personal hardware License OpenMDW License Agreement, version 1.1 Release Date August 11, 2026

Model Overview

Model Developer: NVIDIA Corporation

Model Dates: December 2025 - May 2026

Data Freshness:

The pre-training data has a cutoff date of September 2025.

The post-training data has a cutoff date of May 2026.

What is Nemotron?

NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.

... continue reading