Miles runs full RL loop on Rubin GPUs with SGLang and Megatron
AIMiles runs the full reinforcement learning loop on Nvidia Rubin, using SGLang for rollout, Megatron for training, and one container image. On a single 4-GPU tray, Qwen3-30B-A3B's GSM8K reward rises from about 45% to about 95% over 50 rollouts, matching the GB300 curve. The post also reports DeepSeek-V4-Flash end-to-end rollout and training, and Qwen3.5-35B-A3B agentic RL with 64 concurrent mini-SWE-agent sandboxes on SWE-bench Verified, where reward holds near 0.6 and median response length falls about 30%.








