vLLM Blog· Inferact and the vLLM Team·· 2d agoPickAI score62
vLLM Speeds Up DeepSeek-V4.1-Flash Agentic Serving Through Kernel and Replay Optimizations
DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0
AI summary
Inferact and the vLLM community reported a 1.9× low-concurrency speedup and about 5.3× throughput under a 150 TPS constraint for DeepSeek-V4.1-Flash over three weeks. Gains came from SWA bounded replay with CUDA graphs, which cut TTFT by about 30%, and from integrated DeepSeek kernels such as MegaAttention, Mega-mHC, Mega-Gate, and DeepSelect. The post measures these results on the SemiAnalysis AgentX benchmark.
Why it matters
The post breaks down how SWA bounded replay and fused kernels cut prefill and decode costs, a reusable engineering pattern for long-context agentic serving.
Source: vLLM Blog · vllm.ai