Skip to content
View original post on X: Inferact· 44/100AI score44/100

vLLM maintainers show TPUv7 megakernels beat GB200 NVL72 on Kimi K3

AISummary

Inferact says vLLM maintainers used megakernel optimization to reach 700 tokens per second per user on TPUv7 running Kimi K3. SemiAnalysis, which shared the work, reports this is 56% better performance than Nvidia's GB200 NVL72. Inferact links a full technical breakdown of the TPU megakernel work on its blog.

Post on XView on X
@inferact

Thanks for the shoutout @SemiAnalysis_ ! Full breakdown linked here:
https://inferact.ai/blog/tpu-megakernels

SemiAnalysis@SemiAnalysis_
ALERT ALERT ALERT 🚨 🚨 🚨 VLLM MAINTAINERS HAVE JUST SHOWN THAT TPUv7 CAN GET 700 tok/s/user,  56% BETTER PERFORMANCE THAN NVIDIA GB200 NVL72 THROUGH MEGAKERNEL OPTIMIZATION ON KIMI K3. As we said awhile ago, the TPU externalization of software is full steam ahead. This is ultra important to follow the progress of this.
View quoted post on X

Source: Inferact · x.comPublished · added here