Google Details Sparse Attention Speedup for Video Diffusion on TPUs
Original titleAccelerating Spatio-Temporal Attention for Video Diffusion on TPUs
AISummary
Google Developers Blog describes how Sparse VideoGen (SVG) routes video diffusion attention heads into spatial or temporal sparse masks and implements them as custom JAX and Pallas Splash Attention kernels on TPU v6e.
In isolated single-chip tests with 75.6K tokens and 10 heads, the sparse variants retain about 38.87% of query-key pairs. The article argues that theoretical sparsity must be converted into hardware tile skipping to yield real speedups.
Source: Google Developers Blog · developers.googleblog.comPublished · added here