NVIDIA Nemotron
NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.
NVIDIA Nemotron Models
Nemotron models are transparent—the training data used for these models, as well as their weights, are open and available on Hugging Face for you to evaluate before deploying them in production. The technical reports outlining the steps necessary to recreate these models are also freely available.
The new Nemotron 3 family provides the most efficient multimodal models, powered by hybrid Mamba‑Transformer MoE with 1M-token context, delivering top accuracy for complex, high-throughput agentic AI applications.
Easily deploy models using open frameworks like vLLM, SGLang, Ollama and llama.cpp on any NVIDIA GPUs—from the edge and cloud to the data center. Endpoints are also available as NVIDIA NIM™ microservices for easy deployment on any GPU-accelerated system.
Nemotron reasoning models are optimized for various platforms:
Nano provides cost efficiency with high accuracy specialized sub-agents—now with multimodal capabilities with Nano Omni..
Super delivers highest efficiency with leading accuracy for reasoning and tool calling for multi-agent applications.
-
Ultra is designed for applications demanding the highest reasoning accuracy for complex agentic tasks.
Additionally, these models provide the highest throughput, enabling agents to think faster and generate higher-accuracy responses while lowering inference cost.
Nemotron models are also available for visual understanding, information retrieval, speech, and safety.
Nemotron 3.5 Lightning
- Fastest 30B MoE model for completing specialized tasks for always-on agents
- Easily post-trained to achieve leading domain-specific accuracy and efficiency
- Deploy anywhere, from local infrastructure to the cloud
Nemotron 3 Ultra 550B A55B
- Ideal for multi-agent enterprise workflows requiring highest accuracy, such as customer service automation, supply chain management, and IT security
- Frontier-level reasoning across multi-step planning, tool use, synthesis, verification, and recovery
- Handles the hardest agent workflow calls, including planning, code generation, and deep research
Nemotron 3 Nano Omni 30B A3B
- Single model for video, audio, image, and text understanding for a simplified agent workflow
- Multimodal reasoning for sub-agents within agentic use cases such as computer use agent, document intelligence, and video/audio understanding
- Highest in-class efficiency and with low costs
Nemotron 3 Super 120B A12B
- Highest in-class efficiency and leading accuracy
- Great for addressing complex tasks in multi-agent environment
- Suitable for single data center GPU deployments
Nemotron 3 Nano 30B A3B
- Nemotron 3 Nano offers 4x faster throughput compared to Nemotron 2 Nano
- Leading accuracy for coding, reasoning, math and long context tasks
- Perfect for agents that need to deliver highest accuracy and efficiency for targeted tasks
Nemotron Retriever
- Industry-leading extraction, embed, and rerank models
- Best-in-class accuracy for multimodal document intelligence, question answering, and passage retrieval
Nemotron Parse
- Understands document semantics and extract text and tables elements with spatial grounding
- Overcomes traditional OCR limitations with support for multi-column layouts, LaTeX table extraction, markdown formatting, and reading-order reconstruction
- Designed to accelerate document intelligence pipelines for RAG, LLM training data curation, and agentic document workflows
Nemotron Speech
- A family of open models optimized for high-throughput, ultra-low latency automatic speech recognition (ASR), text-to-speech (TTS), speech-to-speech (S2S), full-duplex, and neural machine translation (NMT) for agentic AI applications
- Nemotron Speech models with the NVIDIA Riva GPU-accelerated speech AI library deliver state-of-the-art ASR and TTS capabilities for seamless production deployment
Nemotron Safety
- Advanced multilingual, multimodal safety models that deliver high accuracy jailbreak detection, content moderation with cultural nuance, fine-grained PII detection, reasoning-based custom policy enforcement, and topic control for more secure and more compliant LLMs across global domains and use cases.
- NeMo Guardrails, a flexible, open library for defining and enforcing enterprise AI policies in real time—covering dialogue control, topic guidance, RAG grounding, tool‑call governance, safety filtering, and more—with parallel, low‑latency execution across custom, community, and NVIDIA safety rails.
NVIDIA Nemotron Datasets
Improve reasoning capabilities of large language models (LLMs) with one of the broadest commercially usable open data collections for agentic AI — spanning pre-training, post-training, personas, safety, RL, and RAG. Includes 10T+ tokens and 40M+ post-training samples, covering the full training lifecycle from foundation models to agent workflows.
Built with large-scale synthetic data generation, filtering, and curation — and released under permissive licenses. Developers can train, fine-tune, and evaluate models with full visibility into the data, accelerating development and reducing reliance on opaque datasets.
Nemotron Pre- and Post-Training Datasets
NVIDIA provides over 10T tokens of multilingual reasoning, coding, and safety data to help the community build their custom models.
Nemotron Personas Datasets
Fully synthetic, privacy-safe personas are grounded in real-world demographic, geographic, and cultural distributions. Part of NVIDIA’s growing global collection for Sovereign AI development, featuring datasets for USA, Japan, India, Singapore, Brazil, France, and South Korea.
Nemotron Omni Datasets
Multimodal data extending the Nemotron training pipeline beyond text to image, video, and speech. ~127B tokens of cross-modal pretraining data and ~124M curated post-training examples for document reasoning, computer use, and long-horizon workflows.
Nemotron Safety Datasets
High-quality, curated datasets built to power multilingual content safety, advanced policy reasoning, and threat-aware AI—spanning moderation data and audio-based safety signals for modern AI assistants.
Nemotron RL Datasets
Train models with the same reinforcement learning (RL) data powering Nemotron, including multi-turn trajectories, tool calls, and preference signals across coding, math, reasoning, and agentic tasks to build adaptive, reliable real-world AI.
Nemotron Retriever Datasets
Unlock the foundation behind our leaderboard-topping model with the release of 15 meticulously curated datasets—spanning instruction-following, reasoning, coding, and evaluation data—to accelerate open research and transparent model development.
Developer Tools
NeMo Switchyard
NVIDIA NeMo Switchyard is an open source model routing library for AI agents. It routes prompts to the most capable and efficient model for each step of an agent workflow automatically, based on configurable routing profiles, making agents smarter, faster, and more efficient.
NVIDIA TensorRT-LLM
TensorRT™-LLM is an open source library built to deliver high-performance, real-time inference optimization for large language models like Nemotron on NVIDIA GPUs. This open source library is available on the TensorRT-LLM GitHub repo and includes a modular Python runtime, PyTorch-native model authoring, and a stable production API.
Open Source Frameworks
Deploy Nemotron models using open source frameworks such as Hugging Face transformers for development or vLLM for deployment and production use cases on all supported platforms.
Introductory Resources
NVIDIA Nemotron 3.5 Lightning for Fast, Specialized Agent Tasks
NVIDIA Nemotron 3.5 Lightning is an open 30B MoE model with 3B active parameters for specialized tasks in always-on agents. With NeMo Switchyard, developers can route high-volume execution to Lightning when it's the right fit, while reserving frontier models for complex planning.
Switchyard LLM Model Routing: Smarter, Faster, More Efficient AI
Learn how to build and evaluate a model router with NeMo Switchyard, using request signals and routing policies to send each agent task to the right model so your agent runs smarter, faster, and more efficiently.
NVIDIA Nemotron 3 Ultra Powers Faster, More Efficient Reasoning for Long-Running Agents
NVIDIA Nemotron 3 Ultra, a 550B MoE, is the frontier-intelligence open model for long-running agents, built for orchestration and complex reasoning across coding, deep research, and enterprise workflows. Powered by the highest throughput in its class, Nemotron 3 Ultra delivers fastest time to task completion without sacrificing accuracy and can be customized for any domain.
Starter Kits
Start solving AI challenges by developing custom agents with NVIDIA Nemotron models for various use cases. Explore implementation scripts, explainer blogs, and more how-to documentation for various stages of AI development.
Build a Report Generation Agent With Nemotron
The workshop guides developers in building a report generation agent using NVIDIA Nemotron and LangGraph, focusing on four core considerations of AI agents: model, tools, memory and state, and routing.
Tutorial Video: Building a Report Generation Agent With NVIDIA Nemotron Nano v2
NVIDIA Launchable: Build an Agent Workshop
Learning Path: How to Build an AI Agent
Build a RAG Agent With Nemotron
In this self-paced workshop, gain a deep understanding of agentic retrieval-augmented generation (RAG) core principles, including the NVIDIA Nemotron model family, and learn how to build your own customized, shareable agentic RAG system using LangGraph within a turnkey, portable development environment.
Tutorial Video: Build a RAG Agent With NVIDIA Nemotron
On-Demand Livestream: Build a RAG Agent With NVIDIA Nemotron | Nemotron Labs
NVIDIA Launchable: Build an Agent Workshop
Learning Path: How To Build an Agent RAG Application
Build a Bash Computer Use Agent With Nemotron
In this self-paced workshop, gain a deep understanding of agentic retrieval-augmented generation (RAG) core principles, including the NVIDIA Nemotron model family, and learn how to build your own customized, shareable agentic RAG system using LangGraph within a turnkey, portable development environment.
Tutorial Video: Create a Bash Agent in One Hour
On-Demand Livestream: Build a Bash Computer Operator Agent | Nemotron Labs
Nemotron 3 Ultra 550B A55B
Below are the resources to use the Nemotron 3 Ultra model.
Models: Nemotron 3 Model Collection
Datasets: Pretraining and Post training
Notebook/Docs: Fine-Tune for Your Use Case with Unsloth Notebook and Docs
Nemotron 3 Super 120B A3B
Below is a set of resources that outline the process NVIDIA used to produce the Nemotron 3 Super model.
Datasets: Pretraining, Post training, and RL Dataset
Models: Nemotron 3 Model Collection
Build a Voice Agent With RAG and Safety Guardrails With Nemotron
In this tutorial, you’ll learn how to build a voice-powered RAG agent with safety guardrails using Nemotron models. By the end, your agent will listen to spoken input, ground itself in your data, reason over long context, apply guardrails, and return safe answers as audio.
Tutorial Video: How to Build a Voice Agent With RAG and Safety Guardrails
Models: Speech
Models: RAG
Models: Safety
Models: Reasoning
Run Production AI Agents With Nemotron Models Across Hosted and Self-Managed Infrastructure
Build and deploy production-ready AI agents using NVIDIA Nemotron across hosted inference providers, AI clouds, or your own infrastructure.
Managed inference providers handle optimized runtimes, elastic scaling, and production deployment so you can focus on building agentic AI applications—not infrastructure. With NeMo Switchyard, you can also build intelligent multi-model AI systems that automatically route every task to the most capable and efficient model based on accuracy, latency, and cost, based on configurable routing profiles.
For self-managed deployments, platforms like Canonical (Ubuntu, Kubernetes, and MLOps tooling) enable running Nemotron models across private cloud, on-premises, or hybrid environments with full control over infrastructure.
Available Providers
Run Nemotorn Models Locally
For local development and on-device workflows, Nemotorn models can also be run locally. These tools enable you to prototype quickly, experiment privately, and build offline without relying on hosted endpoints:
- LM Studio —Built-in interface and OpenAI-compatible API
- Ollama—CLI and developer-friendly local API
- Llama.cpp—Lightweight, high-performance inference engine (GGUF models available via Hugging Face)
- Unsloth—Efficient local fine-tuning inference with optimized memory usage and performance
Optimize Your Inference Stack
Need additional performance or deployment flexibility? Optimize your inference stack with deployment guides and cookbooks for vLLM, SGLang, or TensorRT-LLM.
Discover and Access Nemotron
Explore Nemotron models, documentation, deployment options, and access paths through:
- Anaconda (Nemotron available through Anaconda AI Orchestrator)
More Resources
Ethical Considerations
NVIDIA believes Trustworthy AI is a shared responsibility, and we have established policies and practices to enable development for a wide array of AI applications. When downloading or using this model in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
NVIDIA has collaborated with Google DeepMind to watermark generated videos from the NVIDIA API catalog.
For more detailed information on ethical considerations for this model, please see the System Card, Model Card++ Explainability, Bias, Safety & Security, and Privacy Subcards. Please report security vulnerabilities or NVIDIA AI concerns here.