Skip to content
Read the original: LMSYS Org· Published Pick65/100AI score65/100

SGLang adds NVFP4 KV cache for longer context on Blackwell GPUs

Original title🚀New blog: Accelerating Long-Context and Agentic Inference with NVFP4 KV Cache

AISummary

LMSYS Org says NVFP4 KV cache in SGLang fits about 1.78x more context into GPU memory and speeds long-context decoding by up to 78%.

Built with Alibaba Qwen and NVIDIA for Blackwell, it stores KV at about 56% of FP8's per-token footprint, with decode throughput up 37%, 58%, and 78% at 32K, 160K, and 1M context.

The post reports near-lossless accuracy versus FP8 on GPQA-Diamond and AIME 2025 using Qwen3.5-397B-A17B, and it can be enabled with --kv-cache-dtype nvfp4.

AIWhy it matters

The post gives specific memory and throughput figures for NVFP4 KV cache in SGLang, showing how the format trades cache footprint against long-context decode speed.

Read the original x.com

Source: LMSYS Org · x.comPublished · added here