Skip to content
View original post on X: Prime Intellect· 38/100AI score38/100

Prime Intellect stores MLA KV cache in NVFP4 for more cached tokens

AISummary

Prime Intellect compresses the MLA latent KV cache to NVFP4, reducing each row from 576 to 352 bytes. This fits about 50% more cached tokens per decoder compared with FP8. Its native sparse-MLA kernel unpacks the format on-chip, and the company is contributing that kernel to FlashInfer as an experimental operation.

Post on XView on X
@PrimeIntellect

A reply · the post it answers

Decode: NVFP4 KV compression

TP4 had the lowest inter-token latency, but every rank holds a full KV copy. Storing the MLA latent in NVFP4 (576 to 352 bytes/row) fits ~50% more cached tokens per decoder compared to FP8.

Our native sparse-MLA kernel unpacks it on-chip. We are contributing the native kernel to FlashInfer as an experimental operation.

Source: Prime Intellect · x.comPublished · added here