Thanks for the shoutout @SemiAnalysis_ ! Full breakdown linked here:
https://inferact.ai/blog/tpu-megakernels
vLLM maintainers show TPUv7 megakernels beat GB200 NVL72 on Kimi K3
AISummary
Inferact says vLLM maintainers used megakernel optimization to reach 700 tokens per second per user on TPUv7 running Kimi K3. SemiAnalysis, which shared the work, reports this is 56% better performance than Nvidia's GB200 NVL72. Inferact links a full technical breakdown of the TPU megakernel work on its blog.
Post on XView on X
@inferact
ALERT ALERT ALERT 🚨 🚨 🚨 VLLM MAINTAINERS HAVE JUST SHOWN THAT TPUv7 CAN GET 700 tok/s/user, 56% BETTER PERFORMANCE THAN NVIDIA GB200 NVL72 THROUGH MEGAKERNEL OPTIMIZATION ON KIMI K3. As we said awhile ago, the TPU externalization of software is full steam ahead. This is ultra important to follow the progress of this.
Source: Inferact · x.comPublished · added here
