StackMap
Subscribe

Meshy vs slime

OpenBMB's RL training framework where every role — inference, training, rollout — is an independent service talking over one TransferQueue data plane. Built on SGLang and torchtitan. — versus — THUDM's RL post-training framework behind the GLM releases — Megatron training plus SGLang rollouts with native arg pass-through, and pluggable reward, verifier and agentic data-generation workflows.

The curated verdict

Both pair SGLang rollouts with a separate trainer for RL post-training; slime trains on Megatron, Meshy on torchtitan with a service-per-role design.

Meshyslime
Stars3868.6k
Forks221.3k
LanguagePythonPython
License—Apache-2.0
Last activityyesterday3 days ago
Topicstrainingtraining
Curated connections37

Meshy — the curator's take

For RL-infra people who want on-policy, bounded off-policy and fully async GRPO from the same services, colocation layout as a recipe knob, and stalls visible as piled-up queue columns. It's alpha (0.1.0, CUDA 12.9 + Docker) and the bundled recipes are math-reasoning GRPO on Qwen3/MiniCPM. Need a proven production stack today? verl or slime have far more mileage.

slime — the curator's take

One of the few open RL stacks proven on frontier releases (GLM-4.5 through 5.2, with Qwen/DeepSeek/Llama support): the Megatron+SGLang-only bet keeps the dataflow explicit and upstream engine features usable instead of flattened behind a multi-backend abstraction, and rollout-only/train-only debug paths take RL's silent-bug problem seriously. NOT an afternoon tool — you need Megatron-scale GPU infrastructure and RL literacy; for single-node SFT or LoRA use a lighter trainer. And if your rollout engine must be vLLM, this is the wrong framework by design.