vLLM v0.31.0 adds DeepSeek-V4.1-Flash support and new serving features
Original title717 commits. 307 contributors. 96 first-timers. vLLM v0.31.0 is live. 🎉
AISummary
vLLM v0.31.0 is released with 717 commits from 307 contributors, including 96 first-time contributors.
Highlights include DeepSeek-V4.1-Flash support, a vllm preload command that keeps weights in GPU memory across restarts, and Model Runner V2 with draft-model speculative decoding.
The release also adds large-scale serving, scheduling, and HiSparse fixes, with full notes linked on GitHub.
Source: vLLM · x.comPublished · added here