StackMap
Subscribe

labs-molt vs Meshy

Molt: NVIDIA's agentic-first RL framework in ~9K lines — Ray for placement, vLLM for rollout, AutoModel + FSDP2 for training — fully async, multimodal, multi-turn, scaling to 1T-class MoE. — versus — OpenBMB's RL training framework where every role — inference, training, rollout — is an independent service talking over one TransferQueue data plane. Built on SGLang and torchtitan.

The curated verdict

Both are newer fully-async RL frameworks; Molt uses Ray + vLLM + FSDP2, Meshy SGLang + torchtitan over TransferQueue.

labs-moltMeshy
Stars1.2k386
Forks11322
LanguagePythonPython
LicenseApache-2.0—
Last activity4 days agoyesterday
Topicstrainingtraining
Curated connections63

labs-molt — the curator's take

The RL stack to read end-to-end: the agent is the program (one Gymnasium-aligned Env.step() or ChatAgent.run(), reward is any Python you write), the trainer is a single actor with an optional KL reference, and Ray owns the async queues between rollout, training and weight sync. Token-first contract keeps ids, logprobs, action ranges, rewards and multimodal tensors aligned so multi-turn tool use and VLM environments share one format. Same script trains 8B and DeepSeek-V3-class MoE with expert parallelism. NOT for breadth — it deliberately has fewer algorithms and integrations than verl or TRL — and NOT for SFT-only work; it's for research groups running agentic RL who want every gradient one file away.

Meshy — the curator's take

For RL-infra people who want on-policy, bounded off-policy and fully async GRPO from the same services, colocation layout as a recipe knob, and stalls visible as piled-up queue columns. It's alpha (0.1.0, CUDA 12.9 + Docker) and the bundled recipes are math-reasoning GRPO on Qwen3/MiniCPM. Need a proven production stack today? verl or slime have far more mileage.