StackMap
Subscribe
Explore / Meshy
OpenBMB

Meshy

OpenBMB's RL training framework where every role — inference, training, rollout — is an independent service talking over one TransferQueue data plane. Built on SGLang and torchtitan.

386 22 Pythonupdated yesterday
View on GitHubDispute this mapping →
Curator's take

For RL-infra people who want on-policy, bounded off-policy and fully async GRPO from the same services, colocation layout as a recipe knob, and stalls visible as piled-up queue columns. It's alpha (0.1.0, CUDA 12.9 + Docker) and the bundled recipes are math-reasoning GRPO on Qwen3/MiniCPM. Need a proven production stack today? verl or slime have far more mileage.

Mapped by ShipWithAI editors · links verified

Continue your stack

What teams reach for next — and why each earns a place beside Meshy. Ranked by curator confidence.

alternativealternativealternativeverlslimelabs-moltMeshy
pairs wellalternativebuilt withpick a node for the why · open it from the panel
Weekly digest
README.md2 min read
Meshy — Asynchronous RL Engine for LLMs

Notion Blog Zhihu GitHub Docker License

Meshy models every role of an RL run as an independent service. Samples flow between services through a single TransferQueue data plane, control flow is driven by data availability, and the whole topology is derived locally by each process from one declarative recipe. Built on SGLang and torchtitan.

Meshy architecture

Highlights

  • 🧩 Every role as a service. Inference, training and rollout run as independent processes that talk through queue columns and a handful of gate signals. There is no driver that fans out RPCs or forwards every tensor.

  • 🗂️ TransferQueue as both data and control plane. All communication happens through queue columns; column readiness is the only control signal, so services never handshake directly. Gate pulses, GPU ownership, and tensors themselves travel in the same middleware.

  • ⚡ Native async algorithm support. Recipes for on-policy, bounded off-policy and fully asynchronous training use the same set of services, with only the rollout pacing window as a knob.

  • 🧭 Topology as a pure function. Full placement is calculated SPMD-style on each machine, without the need for service discovery. A misplaced recipe is identified at startup.

  • 🔄 Colocation with any number of services. GPU ownership is a token passed over TransferQueue; developers can freely arrange any number of services colocated on the same set of GPUs.

  • 🪶 Lightweight and debuggable. Logs are kept one file per service with full tracebacks. When something stalls, the queue shows it as piled-up unconsumed columns.

News

  • [2026.09.07] 🎉 Meshy is now open-source! Visit our blog for details.

Quick Start

Prerequisites

  • NVIDIA GPU with CUDA 12.9 support
  • Docker with NVIDIA Container Toolkit (or a native Ubuntu 24.04 environment)
  • Python 3.12+ (if installing manually)

Option 1: Use the Prebuilt Docker Image (Recommended)

The easiest way to get started is to pull and run our prebuilt image:

docker pull ztonyzhao/meshy:0.1.0-alpha
docker run --gpus all -it --rm ztonyzhao/meshy:0.1.0-alpha

Option 2: Use the Provided Dockerfile

You can also build the Docker image yourself.

docker build -t meshy .
docker run --gpus all -it --rm meshy

This will drop you into a shell with the virtual environment already activated at /opt/meshy. All dependencies (PyTorch, SGLang, TorchTitan, TransferQueue) are pre-installed.

Option 3: Manual Installation

If you prefer to set up the environment without Docker, follow the step-by-step guide in docs/manual_install.md.

Run a recipe

From the repository root, launch any recipe with the same command. The launcher starts TransferQueue, then runs torchrun with one ignitor per GPU:

python scripts/launch.py --recipe recipe.grpo_gsm8k

This is the smallest end-to-end run: Qwen3-1.7B on