[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:meshy":3},"\u003Cdiv align=\"center\">\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FOpenBMB\u002Fmeshy\u002FHEAD\u002Fassets\u002Fmeshy-logo.png\" alt=\"Meshy — Asynchronous RL Engine for LLMs\" width=\"400\" \u002F>\u003Cp>\u003Ca href=\"https:\u002F\u002Fmaydomain.notion.site\u002Fmeshy-blog-en\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FNotion-000000?style=for-the-badge&amp;logo=notion&amp;logoColor=white\" alt=\"Notion Blog\" \u002F>\u003C\u002Fa> \u003Ca href=\"https:\u002F\u002Fzhuanlan.zhihu.com\u002Fp\u002F2080612686585402867\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FZhihu-0084FF?style=for-the-badge&amp;logo=zhihu&amp;logoColor=white\" alt=\"Zhihu\" \u002F>\u003C\u002Fa> \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FOpenBMB\u002FMeshy\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FGitHub-181717?style=for-the-badge&amp;logo=github&amp;logoColor=white\" alt=\"GitHub\" \u002F>\u003C\u002Fa> \u003Ca href=\"https:\u002F\u002Fhub.docker.com\u002Fr\u002Fztonyzhao\u002Fmeshy\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FDocker-2496ED?style=for-the-badge&amp;logo=docker&amp;logoColor=white\" alt=\"Docker\" \u002F>\u003C\u002Fa> \u003Ca href=\"https:\u002F\u002Fwww.apache.org\u002Flicenses\u002FLICENSE-2.0\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-Apache_2.0-green?style=for-the-badge\" alt=\"License\" \u002F>\u003C\u002Fa>\u003C\u002Fp>\n\u003C\u002Fdiv>\u003Cp>\u003Cstrong>Meshy\u003C\u002Fstrong> models every role of an RL run as an\nindependent service. Samples flow between services through a single\nTransferQueue data plane, control flow is driven by data availability, and the\nwhole topology is derived locally by each process from one declarative recipe.\nBuilt on SGLang and torchtitan.\u003C\u002Fp>\n\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FOpenBMB\u002Fmeshy\u002FHEAD\u002Fassets\u002Farchitecture.png\" alt=\"Meshy architecture\" width=\"800\" \u002F>\n\u003C\u002Fp>\u003Ch2>Highlights\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cp>🧩 \u003Cstrong>Every role as a service.\u003C\u002Fstrong> Inference, training and rollout run as\nindependent processes that talk through queue columns and a handful of gate\nsignals. There is no driver that fans out RPCs or forwards every tensor.\u003C\u002Fp>\n\u003C\u002Fli>\n\u003Cli>\u003Cp>🗂️ \u003Cstrong>TransferQueue as both data and control plane.\u003C\u002Fstrong> All communication happens through queue columns; column readiness is the only control signal, so services never handshake directly. Gate pulses, GPU ownership, and tensors themselves travel in the same middleware.\u003C\u002Fp>\n\u003C\u002Fli>\n\u003Cli>\u003Cp>⚡ \u003Cstrong>Native async algorithm support.\u003C\u002Fstrong> Recipes for on-policy, bounded\noff-policy and fully asynchronous training use the same set of services, with\nonly the rollout pacing window as a knob.\u003C\u002Fp>\n\u003C\u002Fli>\n\u003Cli>\u003Cp>🧭 \u003Cstrong>Topology as a pure function.\u003C\u002Fstrong> Full placement is calculated SPMD-style on\neach machine, without the need for service discovery. A misplaced recipe is\nidentified at startup.\u003C\u002Fp>\n\u003C\u002Fli>\n\u003Cli>\u003Cp>🔄 \u003Cstrong>Colocation with any number of services.\u003C\u002Fstrong> GPU\nownership is a token passed over TransferQueue; developers can freely\narrange any number of services colocated on the same set of GPUs.\u003C\u002Fp>\n\u003C\u002Fli>\n\u003Cli>\u003Cp>🪶 \u003Cstrong>Lightweight and debuggable.\u003C\u002Fstrong> Logs are kept one file per service with full\ntracebacks. When something stalls, the queue shows it as piled-up unconsumed\ncolumns.\u003C\u002Fp>\n\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>News\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>[2026.09.07] 🎉 Meshy is now open-source! Visit our \u003Ca href=\"https:\u002F\u002Fmaydomain.notion.site\u002FMeshy-A-Role-Driven-RL-Training-Framework-under-SPMD-Paradigm-3d34e1dff05a80e489a6d9a406991bae\" rel=\"nofollow ugc noopener\">blog\u003C\u002Fa> for details.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Quick Start\u003C\u002Fh2>\n\u003Ch3>Prerequisites\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>NVIDIA GPU with CUDA 12.9 support\u003C\u002Fli>\n\u003Cli>Docker with NVIDIA Container Toolkit (or a native Ubuntu 24.04 environment)\u003C\u002Fli>\n\u003Cli>Python 3.12+ (if installing manually)\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>Option 1: Use the Prebuilt Docker Image (Recommended)\u003C\u002Fh3>\n\u003Cp>The easiest way to get started is to pull and run our \u003Ca href=\"https:\u002F\u002Fhub.docker.com\u002Fr\u002Fztonyzhao\u002Fmeshy\" rel=\"nofollow ugc noopener\">prebuilt image\u003C\u002Fa>:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">docker pull ztonyzhao\u002Fmeshy:0.1.0-alpha\ndocker run --gpus all -it --rm ztonyzhao\u002Fmeshy:0.1.0-alpha\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Ch3>Option 2: Use the Provided Dockerfile\u003C\u002Fh3>\n\u003Cp>You can also build the Docker image yourself.\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">docker build -t meshy .\ndocker run --gpus all -it --rm meshy\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>This will drop you into a shell with the virtual environment already activated at \u003Ccode>\u002Fopt\u002Fmeshy\u003C\u002Fcode>. All dependencies (PyTorch, SGLang, TorchTitan, TransferQueue) are pre-installed.\u003C\u002Fp>\n\u003Ch3>Option 3: Manual Installation\u003C\u002Fh3>\n\u003Cp>If you prefer to set up the environment without Docker, follow the step-by-step guide in \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FOpenBMB\u002Fmeshy\u002Fblob\u002FHEAD\u002Fdocs\u002Fmanual_install.md\" rel=\"nofollow ugc noopener\">docs\u002Fmanual_install.md\u003C\u002Fa>.\u003C\u002Fp>\n\u003Ch3>Run a recipe\u003C\u002Fh3>\n\u003Cp>From the repository root, launch any recipe with the same command. The\nlauncher starts TransferQueue, then runs \u003Ccode>torchrun\u003C\u002Fcode> with one ignitor per\nGPU:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">python scripts\u002Flaunch.py --recipe recipe.grpo_gsm8k\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>This is the smallest end-to-end run: Qwen3-1.7B on\u003C\u002Fp>\n",1791678837690]