Read the original: PyTorch Blog· Han Xu, Jacky Zhou, Jackie (Jiaqi) Xu, Hongtao Yu, Peng Chen (Dev Infra), Darren Liu, Dev (Devashish) Shankar, Max Leung, Nick Riasanovsky, Hao Yan, Manman Ren, Yuanwei (Kevin) Fang·Published· 7d agoAI score38
TLX-Optimized Jagged Flash Attention Beats FA4 on Blackwell B200 for Meta GEM
Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell
AISummary
Meta's Jagged Flash Attention kernel, built with TLX on NVIDIA Blackwell B200, outperforms FlashAttention-4 (May 2026 version) on GEM's jagged shapes by about 13% on the forward pass and about 50% on the backward pass. The TLX attention kernel is roughly 3.2K lines of Triton-level code, about 3× shorter than FA4's ~10K-line CuteDSL kernels. The benchmarks use bfloat16 on B200.
Source: PyTorch Blog · pytorch.org