SGLang v0.5.20 adds Intel XPU support and faster RL rollouts
AISummary
SGLang has released v0.5.20, bringing Intel XPU into standard releases alongside RL sampling masks that make rollouts more reliable with up to 52% faster decode.
The update also adds Unified Radix Tree SWA branching-point caching, which the project says lifts cache hit rate about 20 points and cuts TTFT by roughly one-third, plus up to 12.5× faster ROCm model loading.
New models named in the release include GLM-5.3-Flash, Qwen3.8-Flash-Next, K2 Horizon, Hy4-Preview, FastH3, and VDN-H3.
Post on XView on X
@sgl_project
A reply · the post it answers
SGLang v0.5.20 landed! Welcome @intel XPU to join standard SGLang releases 🎉 Some of our favorite updates: - RL sampling masks make rollouts more reliable, with up to 52% faster decode - Unified Radix Tree adds SWA branching-point caching: ~20pt higher cache hit rate, ~1/3 lower TTFT - DSpark now supports PD + DCP for long-context serving - SGLang Simulator brings scheduler & cache experiments to CPU - ROCm model loading is up to 12.5× faster - SGLang-Diffusion gets up to ~38% lower E2E latency New models include GLM-5.3-Flash, Qwen3.8-Flash-Next, K2 Horizon, Hy4-Preview, FastH3, VDN-H3, and more. Full release notes👇
Source: SGLang · x.comPublished · added here
