Skip to content
View original post on X: Tri Dao· 62/100AI score62/100

FlashAttention-4 paper: attention on Blackwell GPUs nears matmul speed

AISummary

The FlashAttention-4 paper is out, reporting that attention on Blackwell GPUs now runs at roughly matmul speed, reaching about 1600 TFLOPs.

The forward pass is bottlenecked by exponential computation and the backward pass by shared memory bandwidth, and the redesign uses polynomial exponential emulation, a new online softmax that avoids 90% of softmax rescaling, and 2CTA MMA instructions that let two thread blocks share operands to cut shared memory traffic.

Post on XView on X
@tri_dao

The FA4 paper is finally out after a year of work. On Blackwell GPUs, attention now goes about as fast as matmul even though the bottlenecks are so different! Tensor cores are now crazy fast that attn fwd is bottlenecked by exponential, and attn bwd is bottlenecked by shared memory bandwidth.

Some fun stuff in the redesigned algorithm to overcome these bottlenecks: exponential emulation with polynomials, new online softmax to avoid 90% of softmax rescaling, 2CTA MMA instructions that allow two thread blocks to share operands to reduce smem traffic.

Ted Zadouri@tedzadouri
Asymmetric hardware scaling is here. Blackwell tensor cores are now so fast, exp2 and shared memory are the wall. FlashAttention-4 changes the algorithm & pipeline so that softmax & SMEM bandwidth no longer dictate speed. Attn reaches ~1600 TFLOPs, pretty much at matmul speed! joint work w/ Markus Hoehnerbach, Jay Shah(@ultraproduct), Timmy Liu, Vijay Thakkar (@__tensorcore__ ), Tri Dao (@tri_dao) 1/
View quoted post on X

Source: Tri Dao · x.comPublished · added here