Skip to content
Read the original: PyTorch Blog· Han Xu, Jacky Zhou, Jackie (Jiaqi) Xu, Hongtao Yu, Peng Chen (Dev Infra), Darren Liu, Dev (Devashish) Shankar, Max Leung, Nick Riasanovsky, Hao Yan, Manman Ren, Yuanwei (Kevin) Fang·Published· 7d agoAI score38

TLX-Optimized Jagged Flash Attention Beats FA4 on Blackwell B200 for Meta GEM

Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell

AISummary

Meta's Jagged Flash Attention kernel, built with TLX on NVIDIA Blackwell B200, outperforms FlashAttention-4 (May 2026 version) on GEM's jagged shapes by about 13% on the forward pass and about 50% on the backward pass. The TLX attention kernel is roughly 3.2K lines of Triton-level code, about 3× shorter than FA4's ~10K-line CuteDSL kernels. The benchmarks use bfloat16 on B200.

Read the original pytorch.org

Source: PyTorch Blog · pytorch.org