Skip to content
View original post on X: SGLangOfficial· 37/100AI score37/100

Miles runs full RL loop on Rubin GPUs with SGLang and Megatron

AISummary

Miles runs the full reinforcement learning loop on Nvidia Rubin, using SGLang for rollout, Megatron for training, and one container image.

On a single 4-GPU tray, Qwen3-30B-A3B's GSM8K reward rises from about 45% to about 95% over 50 rollouts, matching the GB300 curve.

The post also reports DeepSeek-V4-Flash end-to-end rollout and training, and Qwen3.5-35B-A3B agentic RL with 64 concurrent mini-SWE-agent sandboxes on SWE-bench Verified, where reward holds near 0.6 and median response length falls about 30%.

Post on XView on X
SGLangVerified on X
@sgl_project

Part of a thread · earlier post

Miles runs the full RL loop on Rubin (SGLang rollout, Megatron training, one container image). On one 4-GPU tray:

  • Qwen3-30B-A3B on GSM8K: ~45% → ~95% reward over 50 rollouts, matching the GB300 curve.
  • DeepSeek-V4-Flash (4-layer, FP8): rollout and training end to end.
  • Qwen3.5-35B-A3B agentic RL: 64 concurrent mini-SWE-agent sandboxes on the Vera CPU, on SWE-bench Verified. Reward holds ~0.6, median response length down ~30%.

Source: SGLang · x.comPublished · added here