Skip to content
View original post on X: RadixArk· 43/100AI score43/100

SGLang-Diffusion runs MiniMax H3 video generation faster than playback

AISummary

RadixArk's SGLang-Diffusion, paired with VDN-H3, generates 14.4 seconds of 768p video in 9.0 seconds on 8× B200 GPUs. The 8-step denoising alone takes 6.9 seconds, which is over 2× real time, and the team reports no measured quality regression against dense 50-step MiniMax H3 across 103 test prompts.

Post on XView on X
@radixark

SGLang-Diffusion enables fast, scalable inference for multimodal generation. The latest work with VDN-H3 brings @MiniMax_AI H3 to faster-than-playback video generation with strong scaling across GPUs. 👏

SGLang@sgl_project
SGLang-Diffusion with VDN-H3 now generates 14.4s of 768p video in just 9.0s 🚀 On 8× B200, 8 step denoising takes just 6.9s, reaching over 2× real time. The 9.0s figure covers the full generation request after warmup. No measured quality regression versus dense 50-step H3 across 103 test prompts. 🧵
View quoted post on X

Source: RadixArk · x.comPublished · added here