StackMap
Subscribe

Meshy vs verl

OpenBMB's RL training framework where every role — inference, training, rollout — is an independent service talking over one TransferQueue data plane. Built on SGLang and torchtitan. — versus — ByteDance's RL post-training library (HybridFlow): PPO/GRPO dataflows in a few lines, FSDP/Megatron training with vLLM/SGLang rollouts, production-proven at frontier scale.

The curated verdict

Both run GRPO-style RL post-training with separate rollout and training engines; verl drives HybridFlow from a controller, Meshy drops the driver and coordinates services through queue readiness.

Meshyverl
Stars38624k
Forks224.7k
LanguagePythonPython
License—Apache-2.0
Last activityyesterday3 days ago
Topicstrainingtraining
Curated connections38

Meshy — the curator's take

For RL-infra people who want on-policy, bounded off-policy and fully async GRPO from the same services, colocation layout as a recipe knob, and stalls visible as piled-up queue columns. It's alpha (0.1.0, CUDA 12.9 + Docker) and the bundled recipes are math-reasoning GRPO on Qwen3/MiniCPM. Need a proven production stack today? verl or slime have far more mileage.

verl — the curator's take

The community default for open RL post-training: the hybrid-controller model expresses PPO/GRPO/DAPO dataflows in a few lines, and the backend matrix (FSDP or Megatron for training, vLLM or SGLang for rollouts) means it fits infrastructure you already have. Proven on real frontier runs and the most-forked codebase in the space. NOT an afternoon tool — multi-GPU distributed debugging is table stakes; for single-node SFT/LoRA use TRL instead, and know the tradeoff slime calls out: multi-backend abstraction can lag upstream engine features.