DeepSeek-V4.1-Flash on vLLM gains 1.9× speed and 5.3× throughput
AIThree weeks after launch, vLLM's serving of DeepSeek-V4.1-Flash runs 1.9× faster at low concurrency and delivers 5.3× the throughput at 150 TPS per user on SemiAnalysis AgentX. The post offers interactive figures explaining how these gains were achieved.







