[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:model-optimizer":3},"\u003Cdiv align=\"center\">\u003Cp>\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FNVIDIA\u002Fmodel-optimizer\u002FHEAD\u002Fdocs\u002Fsource\u002Fassets\u002Fmodel-optimizer-banner.png\" alt=\"Banner image\" \u002F>\u003C\u002Fp>\n\u003Ch1>NVIDIA Model Optimizer\u003C\u002Fh1>\n\u003Cp>\u003Ca href=\"https:\u002F\u002Fnvidia.github.io\u002FModel-Optimizer\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FDocumentation-latest-brightgreen.svg?style=flat\" alt=\"Documentation\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fpypi.org\u002Fproject\u002Fnvidia-modelopt\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fv\u002Fnvidia-modelopt?label=Release\" alt=\"version\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FNVIDIA\u002Fmodel-optimizer\u002Fblob\u002FHEAD\u002FLICENSE\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-Apache%202.0-blue\" alt=\"license\" \u002F>\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>\u003Ca href=\"https:\u002F\u002Fnvidia.github.io\u002FModel-Optimizer\" rel=\"nofollow ugc noopener\">Documentation\u003C\u002Fa> |\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FNVIDIA\u002FModel-Optimizer\u002Fissues\u002F1699\" rel=\"nofollow ugc noopener\">Roadmap\u003C\u002Fa>\u003C\u002Fp>\n\u003C\u002Fdiv>\u003Chr \u002F>\n\u003Cp>\u003Cstrong>NVIDIA Model Optimizer\u003C\u002Fstrong> (referred to as \u003Cstrong>Model Optimizer\u003C\u002Fstrong>, or \u003Cstrong>ModelOpt\u003C\u002Fstrong>) is a library comprising state-of-the-art model optimization \u003Ca href=\"#techniques\" rel=\"nofollow ugc noopener\">techniques\u003C\u002Fa> including quantization, pruning, Neural Architecture Search (NAS), distillation, speculative decoding and sparsity to accelerate models.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>[Input]\u003C\u002Fstrong> Model Optimizer currently supports inputs of a \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002F\" rel=\"nofollow ugc noopener\">Hugging Face\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fpytorch\u002Fpytorch\" rel=\"nofollow ugc noopener\">PyTorch\u003C\u002Fa> or \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fonnx\u002Fonnx\" rel=\"nofollow ugc noopener\">ONNX\u003C\u002Fa> model.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>[Optimize]\u003C\u002Fstrong> Model Optimizer provides Python APIs for users to easily compose the above model optimization techniques and export an optimized quantized checkpoint.\nModel Optimizer is also integrated with \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FMegatron-Bridge\" rel=\"nofollow ugc noopener\">NVIDIA Megatron-Bridge\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FNVIDIA\u002FMegatron-LM\" rel=\"nofollow ugc noopener\">Megatron-LM\u003C\u002Fa> and \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fhuggingface\u002Faccelerate\" rel=\"nofollow ugc noopener\">Hugging Face Accelerate\u003C\u002Fa> for training required inference optimization techniques.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>[Export for deployment]\u003C\u002Fstrong> Seamlessly integrated within the NVIDIA AI software ecosystem, the quantized checkpoint generated from Model Optimizer is ready for deployment in downstream inference frameworks like \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fsgl-project\u002Fsglang\" rel=\"nofollow ugc noopener\">SGLang\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FNVIDIA\u002FTensorRT-LLM\u002Ftree\u002Fmain\u002Fexamples\u002Fquantization\" rel=\"nofollow ugc noopener\">TensorRT-LLM\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FNVIDIA\u002FTensorRT\" rel=\"nofollow ugc noopener\">TensorRT\u003C\u002Fa>, or \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fvllm\" rel=\"nofollow ugc noopener\">vLLM\u003C\u002Fa>. The unified Hugging Face export API now supports both transformers and diffusers models.\u003C\u002Fp>\n\u003Ch2>Latest News\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>[2026\u002F08\u002F17] \u003Ca href=\"https:\u002F\u002Fdeveloper.nvidia.com\u002Fblog\u002Fdeveloping-nemotron-3-5-lightning-nvfp4-with-qad-using-nvidia-model-optimizer\u002F\" rel=\"nofollow ugc noopener\">BLOG: Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer\u003C\u002Fa>: Learn how quantization-aware distillation recovers accuracy from aggressive NVFP4 quantization while reducing model size and increasing throughput.\u003C\u002Fli>\n\u003Cli>[2026\u002F06\u002F26] \u003Ca href=\"https:\u002F\u002Fdeveloper.nvidia.com\u002Fblog\u002Fcreating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\u002F\" rel=\"nofollow ugc noopener\">BLOG: Creating the NVIDIA Nemotron 3 Ultra NVFP4 Checkpoint with NVIDIA Model Optimizer\u003C\u002Fa>: How we quantized Nemotron 3 Ultra (550B) to NVFP4 with Model Optimizer — up to 5.9× higher decode-heavy inference throughput than GLM-5.1 754B FP4 while matching BF16 accuracy. \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Fnvidia\u002FNVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4\" rel=\"nofollow ugc noopener\">NVFP4 Checkpoint\u003C\u002Fa> on Hugging Face.\u003C\u002Fli>\n\u003Cli>[2026\u002F05\u002F27] \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FNVIDIA\u002Fmodel-optimizer\u002Fblob\u002FHEAD\u002Fexamples\u002Fmegatron_bridge\u002Ftutorials\u002FNVIDIA-Nemotron-3-Nano-30B-A3B-BF16\" rel=\"nofollow ugc noopener\">\u003Cstrong>End-to-end Optimization tutorial for Nemotron-3-Nano-30B-A3B\u003C\u002Fstrong>\u003C\u002Fa>: Pruning + two-phase distillation + FP8 quantization achieving 2.6× vLLM throughput and 2.6× memory reduction.\u003C\u002Fli>\n\u003Cli>[2026\u002F05\u002F13] \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FNVIDIA\u002Fmodel-optimizer\u002Fblob\u002FHEAD\u002Fexamples\u002Fpuzzletron\" rel=\"nofollow ugc noopener\">\u003Cstrong>Puzzletron\u003C\u002Fstrong>\u003C\u002Fa>: A new algorithm for heterogeneous pruning &amp; NAS of LLM and VLM models.\u003C\u002Fli>\n\u003Cli>[2026\u002F04\u002F15] Customer story: \u003Ca href=\"https:\u002F\u002Fwww.domyn.com\u002Fblog\u002Fdomyn-large-the-journey-of-a-european-sovereign-ai-model-for-regulated-industries\" rel=\"nofollow ugc noopener\">Domyn compresses Colosseum-355B → 260B using ModelOpt's Minitron pruning + distillation\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>[2026\u002F03\u002F17] Customer story: \u003Ca href=\"https:\u002F\u002Fbielik.ai\u002Fen\u002Fnvidia-gtc-bielik-minitron-premiere\u002F\" rel=\"nofollow ugc noopener\">Bielik.AI builds Bielik Minitron 7B (33% smaller, 50% faster, 90% quality retained) using ModelOpt's Minitron pruning + distillation\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>[2026\u002F03\u002F11] Model Optimizer quantized Nemotron-3-Super checkpoints are available on Hugging Face for downl\u003C\u002Fli>\n\u003C\u002Ful>\n",1788219549001]