Meshy vs slime
OpenBMB's RL training framework where every role — inference, training, rollout — is an independent service talking over one TransferQueue data plane. Built on SGLang and torchtitan. — versus — THUDM's RL post-training framework behind the GLM releases — Megatron training plus SGLang rollouts with native arg pass-through, and pluggable reward, verifier and agentic data-generation workflows.
Both pair SGLang rollouts with a separate trainer for RL post-training; slime trains on Megatron, Meshy on torchtitan with a service-per-role design.
| Meshy | slime | |
|---|---|---|
| Stars | 386 | 8.6k |
| Forks | 22 | 1.3k |
| Language | Python | Python |
| License | — | Apache-2.0 |
| Last activity | yesterday | 3 days ago |
| Topics | training | training |
| Curated connections | 3 | 7 |
Meshy — the curator's take
For RL-infra people who want on-policy, bounded off-policy and fully async GRPO from the same services, colocation layout as a recipe knob, and stalls visible as piled-up queue columns. It's alpha (0.1.0, CUDA 12.9 + Docker) and the bundled recipes are math-reasoning GRPO on Qwen3/MiniCPM. Need a proven production stack today? verl or slime have far more mileage.
slime — the curator's take
One of the few open RL stacks proven on frontier releases (GLM-4.5 through 5.2, with Qwen/DeepSeek/Llama support): the Megatron+SGLang-only bet keeps the dataflow explicit and upstream engine features usable instead of flattened behind a multi-backend abstraction, and rollout-only/train-only debug paths take RL's silent-bug problem seriously. NOT an afternoon tool — you need Megatron-scale GPU infrastructure and RL literacy; for single-node SFT or LoRA use a lighter trainer. And if your rollout engine must be vLLM, this is the wrong framework by design.