Meshy vs verl
OpenBMB's RL training framework where every role — inference, training, rollout — is an independent service talking over one TransferQueue data plane. Built on SGLang and torchtitan. — versus — ByteDance's RL post-training library (HybridFlow): PPO/GRPO dataflows in a few lines, FSDP/Megatron training with vLLM/SGLang rollouts, production-proven at frontier scale.
Both run GRPO-style RL post-training with separate rollout and training engines; verl drives HybridFlow from a controller, Meshy drops the driver and coordinates services through queue readiness.
| Meshy | verl | |
|---|---|---|
| Stars | 386 | 24k |
| Forks | 22 | 4.7k |
| Language | Python | Python |
| License | — | Apache-2.0 |
| Last activity | yesterday | 3 days ago |
| Topics | training | training |
| Curated connections | 3 | 8 |
Meshy — the curator's take
For RL-infra people who want on-policy, bounded off-policy and fully async GRPO from the same services, colocation layout as a recipe knob, and stalls visible as piled-up queue columns. It's alpha (0.1.0, CUDA 12.9 + Docker) and the bundled recipes are math-reasoning GRPO on Qwen3/MiniCPM. Need a proven production stack today? verl or slime have far more mileage.
verl — the curator's take
The community default for open RL post-training: the hybrid-controller model expresses PPO/GRPO/DAPO dataflows in a few lines, and the backend matrix (FSDP or Megatron for training, vLLM or SGLang for rollouts) means it fits infrastructure you already have. Proven on real frontier runs and the most-forked codebase in the space. NOT an afternoon tool — multi-GPU distributed debugging is table stakes; for single-node SFT/LoRA use TRL instead, and know the tradeoff slime calls out: multi-backend abstraction can lag upstream engine features.