SGLang adds Day-0 support for Qwen3.8-Flash-Next with an NVFP4 checkpoint
Original titleSGLang brings Day-0 support for Qwen 3.8-Flash-Next, an early preview of the Qwen4 architecture!
AISummary
SGLang announced Day-0 support for Qwen3.8-Flash-Next, a 125B MoE model with 6B active parameters and 51B N-gram embeddings, in collaboration with Alibaba Qwen, NVIDIA, and AMD. The post reports 540 tok/s decode speed at BS=1 on NVIDIA B200 (TP4) with an NVFP4 checkpoint, and says N-gram host offloading saves 23.5 GiB VRAM per GPU and raises KV capacity by 78.5%.
Source: LMSYS Org · x.com