Skip to content
View original post on X: SGLangOfficial· 52/100AI score52/100

SGLang adds Rubin optimizations that speed up Kimi K3 inference

AISummary

SGLang says it worked with NVIDIA to optimize attention, MoE, and speculative verification kernels for Kimi K3 inference on early-access Rubin hardware.

It reports up to 20% faster FP8 MLA at batch 1 with 128K context, 20% faster KDA verification with bitwise-identical output, and a 5.9% end-to-end speedup from MoE tail fusion that removes 276 kernel launches per decode step.

The post also says SGLang powers Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.

Post on XView on X
SGLangVerified on X
@sgl_project

We brought SGLang to NVIDIA Vera Rubin and accelerated Kimi K3 inference.
Working closely with @NVIDIA, we optimized attention, MoE, and speculative verification kernels on early-access Rubin hardware.

Highlights:
• Up to 20% faster FP8 MLA at batch 1 / 128K context
• 20% faster KDA verification, with bitwise-identical output
• 5.9% end-to-end inference speedup from MoE tail fusion, removing 276 kernel launches per decode step

SGLang also powers rollouts for Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.

Full results and engineering details 👉 https://www.lmsys.org/blog/2026-10-09-vera-rubin

Source: SGLang · x.comPublished · added here