vLLM's block-major KV layout halves NVLink transfer time
Original titleKV transfer: reducing copy overhead on NVLink
AISummary
vLLM changed its KV cache layout to block-major BLHNC, cutting transfer descriptors about 10x and halving mean KV transfer time on NVLink. The original slowdown came from fragmented KV layout that split one 200K-token request into 32K tiny copies, making NVLink slower than InfiniBand.
Source: Prime Intellect · x.comPublished · added here