Decode: NVFP4 KV compression
TP4 had the lowest inter-token latency, but every rank holds a full KV copy. Storing the MLA latent in NVFP4 (576 to 352 bytes/row) fits ~50% more cached tokens per decoder compared to FP8.
Our native sparse-MLA kernel unpacks it on-chip. We are contributing the native kernel to FlashInfer as an experimental operation.
