Meshy models every role of an RL run as an independent service. Samples flow between services through a single TransferQueue data plane, control flow is driven by data availability, and the whole topology is derived locally by each process from one declarative recipe. Built on SGLang and torchtitan.
Highlights
🧩 Every role as a service. Inference, training and rollout run as independent processes that talk through queue columns and a handful of gate signals. There is no driver that fans out RPCs or forwards every tensor.
🗂️ TransferQueue as both data and control plane. All communication happens through queue columns; column readiness is the only control signal, so services never handshake directly. Gate pulses, GPU ownership, and tensors themselves travel in the same middleware.
⚡ Native async algorithm support. Recipes for on-policy, bounded off-policy and fully asynchronous training use the same set of services, with only the rollout pacing window as a knob.
🧭 Topology as a pure function. Full placement is calculated SPMD-style on each machine, without the need for service discovery. A misplaced recipe is identified at startup.
🔄 Colocation with any number of services. GPU ownership is a token passed over TransferQueue; developers can freely arrange any number of services colocated on the same set of GPUs.
🪶 Lightweight and debuggable. Logs are kept one file per service with full tracebacks. When something stalls, the queue shows it as piled-up unconsumed columns.
News
- [2026.09.07] 🎉 Meshy is now open-source! Visit our blog for details.
Quick Start
Prerequisites
- NVIDIA GPU with CUDA 12.9 support
- Docker with NVIDIA Container Toolkit (or a native Ubuntu 24.04 environment)
- Python 3.12+ (if installing manually)
Option 1: Use the Prebuilt Docker Image (Recommended)
The easiest way to get started is to pull and run our prebuilt image:
docker pull ztonyzhao/meshy:0.1.0-alpha
docker run --gpus all -it --rm ztonyzhao/meshy:0.1.0-alpha
Option 2: Use the Provided Dockerfile
You can also build the Docker image yourself.
docker build -t meshy .
docker run --gpus all -it --rm meshy
This will drop you into a shell with the virtual environment already activated at /opt/meshy. All dependencies (PyTorch, SGLang, TorchTitan, TransferQueue) are pre-installed.
Option 3: Manual Installation
If you prefer to set up the environment without Docker, follow the step-by-step guide in docs/manual_install.md.
Run a recipe
From the repository root, launch any recipe with the same command. The
launcher starts TransferQueue, then runs torchrun with one ignitor per
GPU:
python scripts/launch.py --recipe recipe.grpo_gsm8k
This is the smallest end-to-end run: Qwen3-1.7B on
