Skip to content
View original post on X: Inferact· 46/100AI score46/100

Kimi K3 serving in vLLM is now 2.2–2.8× faster

AISummary

Inferact, with Red Hat AI, NVIDIA, and Huawei, co-led an optimization effort that makes Kimi K3 on vLLM 2.2–2.8× faster. The work spans scheduling, KDA state handling, and custom MoE kernels. The vLLM project's background post cites those throughput gains on a B300 benchmark against v0.27.1 and links a technical deep dive.

Post on XView on X
@inferact

Kimi K3 on @vllm_project is now 2.2–2.8× faster 🚀

Inferact is proud to have co-led this optimization effort with @RedHat_AI, @NVIDIAAI, and @Huawei, spanning scheduling, KDA state handling, and custom MoE kernels.

Read the technical deep dive: https://vllm.ai/blog/2026-09-13-kimi-k3-performance-optimization

vLLM@vllm_project
Kimi K3 serving in vLLM now delivers 2.2–2.8x throughput on our B300 benchmark vs v0.27.1. We break down the work across scheduling, KDA state handling, and MoE kernels, with benchmarks and commands to reproduce the results. Thanks to the vLLM community for pushing Kimi K3 performance forward! Read the deep dive: https://vllm.ai/blog/2026-09-13-kimi-k3-performance-optimization
View quoted post on X

Source: Inferact · x.comPublished · added here