SGLang-Diffusion enables fast, scalable inference for multimodal generation. The latest work with VDN-H3 brings @MiniMax_AI H3 to faster-than-playback video generation with strong scaling across GPUs. 👏
SGLang-Diffusion runs MiniMax H3 video generation faster than playback
AISummary
RadixArk's SGLang-Diffusion, paired with VDN-H3, generates 14.4 seconds of 768p video in 9.0 seconds on 8× B200 GPUs. The 8-step denoising alone takes 6.9 seconds, which is over 2× real time, and the team reports no measured quality regression against dense 50-step MiniMax H3 across 103 test prompts.
Post on XView on X
@radixark
SGLang-Diffusion with VDN-H3 now generates 14.4s of 768p video in just 9.0s 🚀 On 8× B200, 8 step denoising takes just 6.9s, reaching over 2× real time. The 9.0s figure covers the full generation request after warmup. No measured quality regression versus dense 50-step H3 across 103 test prompts. 🧵View quoted post on X
Source: RadixArk · x.comPublished · added here

