How to choose which layers to run at NVFP4 quantization precision
AIBaseten explains how to decide which layers of a model can run in 4-bit NVFP4 without losing needed information. The post compares architecture-based heuristics, isolated-layer sensitivity scoring, and SaturationQuant, which accounts for other quantized layers. It also covers calibration with representative data and block-level scales of 16 values.
Why it matters: The post explains how to choose which layers run at NVFP4 precision using heuristics, sensitivity scoring, and saturation-aware scoring, with clear calibration steps.