Skip to content
Read the original: Prime Intellect· Published 34/100AI score34/100

vLLM's block-major KV layout halves NVLink transfer time

Original titleKV transfer: reducing copy overhead on NVLink

AISummary

vLLM changed its KV cache layout to block-major BLHNC, cutting transfer descriptors about 10x and halving mean KV transfer time on NVLink. The original slowdown came from fragmented KV layout that split one 200K-token request into 32K tiny copies, making NVLink slower than InfiniBand.

Read the original x.com

Source: Prime Intellect · x.comPublished · added here