[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:labs-molt":3},"\u003Cdiv align=\"center\">\u003Ch1>🦋 Molt\u003C\u002Fh1>\n\u003Cp>\u003Cstrong>An agentic-first RL framework for research.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp>Ray · vLLM · NVIDIA AutoModel — the smallest PyTorch-native stack for\n1T-class fully-async, multimodal, multi-turn agentic RL.\u003C\u002Fp>\n\u003Cbr \u002F>\u003Cp>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002Flabs-molt\u002Fblob\u002FHEAD\u002FLICENSE\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-Apache_2.0-2563eb?style=flat-square\" alt=\"License\" \u002F>\u003C\u002Fa>\n\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FPython-3.10+-3776AB?style=flat-square&amp;logo=python&amp;logoColor=white\" alt=\"Python\" \u002F>\n\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FPyTorch-native-EE4C2C?style=flat-square&amp;logo=pytorch&amp;logoColor=white\" alt=\"PyTorch\" \u002F>\n\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FTraining-NVIDIA_AutoModel-76B900?style=flat-square&amp;logo=nvidia&amp;logoColor=white\" alt=\"NVIDIA AutoModel\" \u002F>\n\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FRollout-vLLM-7c3aed?style=flat-square\" alt=\"vLLM\" \u002F>\n\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FRuntime-Ray-028CF0?style=flat-square\" alt=\"Ray\" \u002F>\n\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FRL_code-~9.2K_LOC-10b981?style=flat-square\" alt=\"RL code\" \u002F>\n\u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.21653\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FTech_Report-arXiv:2607.21653-b31b1b?style=flat-square\" alt=\"Tech Report\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fdeepwiki.com\u002FNVIDIA-NeMo\u002Flabs-molt\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fdeepwiki.com\u002Fbadge.svg\" alt=\"Ask DeepWiki\" \u002F>\u003C\u002Fa>\u003C\u002Fp>\n\u003Cbr \u002F>\u003Cp>\u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.21653\" rel=\"nofollow ugc noopener\">\u003Cstrong>Paper\u003C\u002Fstrong>\u003C\u002Fa> ·\n\u003Ca href=\"#-architecture\" rel=\"nofollow ugc noopener\">\u003Cstrong>Architecture\u003C\u002Fstrong>\u003C\u002Fa> ·\n\u003Ca href=\"#-why-molt\" rel=\"nofollow ugc noopener\">\u003Cstrong>Why Molt\u003C\u002Fstrong>\u003C\u002Fa> ·\n\u003Ca href=\"#-quick-start\" rel=\"nofollow ugc noopener\">\u003Cstrong>Quick Start\u003C\u002Fstrong>\u003C\u002Fa> ·\n\u003Ca href=\"#-agent-contract\" rel=\"nofollow ugc noopener\">\u003Cstrong>Agent Contract\u003C\u002Fstrong>\u003C\u002Fa> ·\n\u003Ca href=\"#-recipes\" rel=\"nofollow ugc noopener\">\u003Cstrong>Recipes\u003C\u002Fstrong>\u003C\u002Fa> ·\n\u003Ca href=\"#-scaling-knobs\" rel=\"nofollow ugc noopener\">\u003Cstrong>Scaling\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Package\u003C\u002Fth>\n\u003Cth>SFT\u003C\u002Fth>\n\u003Cth>RL\u003C\u002Fth>\n\u003Cth>Runtime\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>\u003Ccode>molt\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>molt.cli.train_sft\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>molt.cli.train_rl_ray\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>vLLM\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\u003C\u002Ftable>\n\u003C\u002Fdiv>\u003Cp>Molt is \u003Cstrong>agentic-first\u003C\u002Fstrong> and \u003Cstrong>PyTorch-native\u003C\u002Fstrong>. The agent is the program;\nthe trainer is a single actor; reward is any Python you write inside an \u003Ccode>Env\u003C\u002Fcode>\nor \u003Ccode>ChatAgent\u003C\u002Fcode> — graders, multi-turn tools, VLM environments, LLM-as-judge.\nThree components carry the rest — \u003Cstrong>Ray\u003C\u002Fstrong> for placement and async queues,\n\u003Cstrong>vLLM\u003C\u002Fstrong> for rollout, \u003Cstrong>NVIDIA AutoModel + FSDP2\u003C\u002Fstrong> for training in pure\nPyTorch. That is the whole stack: \u003Cstrong>~9.2K lines of RL code that scale to\n1T-class MoE\u003C\u002Fstrong> on vLLM with TP \u002F EP \u002F CP — think DeepSeek-V3 at\n\u003Ccode>--fsdp.ep_size 256\u003C\u002Fcode>, Adam CPU offload for the largest actors. One agent\nAPI, one trainable actor, clean enough to read end-to-end.\u003C\u002Fp>\n\u003Ch2>🧩 Architecture\u003C\u002Fh2>\n\u003Cp>Three boxes. One async loop.\u003C\u002Fp>\n\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FNVIDIA-NeMo\u002Flabs-molt\u002FHEAD\u002Fassets\u002Fmolt.jpg\" alt=\"Molt architecture: Agent · vLLM rollout · Ray async queue · single-actor AutoModel\u002FFSDP2 trainer, fully async\" width=\"920\" \u002F>\n\u003C\u002Fp>\u003Cp>\u003Cstrong>Ray\u003C\u002Fstrong> owns placement and the async queue between the three boxes — that\nis the entire runtime. The contract is \u003Cstrong>token-first\u003C\u002Fstrong>: token ids,\nlogprobs, action ranges, rewards, and multimodal tensors stay aligned from\nrollout to training. Anything you can compute in Python is a valid reward,\nincluding LLM-as-judge calls back through the same vLLM engines that drive\nrollout.\u003C\u002Fp>\n\u003Ch2>✨ Why Molt\u003C\u002Fh2>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>\u003C\u002Fth>\n\u003Cth>What you get\u003C\u002Fth>\n\u003Cth>Why it matters for research\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>🤖 \u003Cstrong>Agentic-first\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>One Gymnasium-aligned API — \u003Ccode>Env.step()\u003C\u002Fcode> or \u003Ccode>ChatAgent.run()\u003C\u002Fcode> — covers graders, multi-turn tools, VLM environments, and OpenAI\u002FAnthropic-compatible servers\u003C\u002Ftd>\n\u003Ctd>The agent \u003Cem>is\u003C\u002Fem> the program — iterate on environments in plain Python, the trainer stays untouched\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>⚙️ \u003Cstrong>Fully-async runtime\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>Ray placement, async rollout queues, vLLM engines, partial rollout, weight sync\u003C\u002Ftd>\n\u003Ctd>Rollout, training, and weight sync overlap — a DeepSeek-V3-class actor stays fed without bespoke infra\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>🔥 \u003Cstrong>PyTorch-native, AutoModel-first\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>FSDP2 + NVIDIA AutoModel, pure PyTorch end-to-end\u003C\u002Ftd>\n\u003Ctd>Hack the model in the language you already write; no backend ceremony\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>🎯 \u003Cstrong>Single-actor simplicity\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>One actor, optional KL reference — the whole RL graph fits on a page\u003C\u002Ftd>\n\u003Ctd>Every gradient is explicit; every loss term is one file away\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>🚀 \u003Cstrong>Frontier-scale MoE\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>AutoModel + FSDP2 + TP \u002F EP \u002F CP + Adam CPU offload, MoE-native — e.g. DeepSeek-V3 with \u003Ccode>--fsdp.ep_size 256\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>The same script that trains 8B scales to 1T-class MoE — no rewrite between scales\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>🔗 \u003Cstrong>Token-first contract\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>Aligned token ids, logprobs, action ranges, rewards, multimodal ten\u003C\u002Ftd>\n\u003Ctd>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\u003C\u002Ftable>\n",1788652555472]