Skip to content
Read the original: PyTorch Blog· Published 47/100AI score47/100

Helion Linear Backend Boosts vLLM Hopper GPU Inference Throughput Over CUTLASS and DeepGEMM

Original titleBuilding a High-Performance and Portable vLLM Linear Backend with Helion

AISummary

The vLLM team integrated Helion, a PyTorch-native kernel DSL, into vLLM's linear backend, using per-shape autotuning to select among Standard GEMM, Split-K, and Swap-AB variants.

On NVIDIA Hopper GPUs, the Helion backend outperformed the default CUTLASS and DeepGEMM backends across the evaluated models, with more than 10% throughput gains for some workloads. The work focuses on FP8 and INT8 quantized GEMM.

Read the original pytorch.org

Source: PyTorch Blog · pytorch.orgPublished · added here