Skip to content
Read the original: Prime Intellect Blog· Published Pick70/100AI score70/100

Prime Intellect releases Prime Flash MoE kernels for faster Blackwell inference

Original titlePrime Flash MoE - Faster MoE Kernels optimized for Blackwell

AISummary

Prime Intellect has released Prime Flash MoE, a set of Blackwell-optimized CUDA kernels for mixture-of-experts feed-forward layers.

The kernels are up to 2.4× faster than PyTorch grouped GEMM and deliver about 2.3× speedup across the 4k–128k token range, and are integrated into its prime-rl framework.

Two pipelines are offered: a fused single-kernel path for small problem sizes and a split three-kernel path for larger ones, supporting both bf16 and MXFP8.

AIWhy it matters

The post explains how fusing MoE expert computation on Blackwell hardware avoids intermediate memory traffic, with benchmarks showing where fused and split pipelines each win.

Read the original primeintellect.ai

Source: Prime Intellect Blog · primeintellect.aiPublished · added here