NVFP4
Saves to local browser storage. Followed topics appear on the homepage and refresh on each visit.Also known as nvidia nvfp4·nvfp4 quantization·nvfp4 inference·nvfp4 training·nvfp4 block scaling
Latest from across the web
External coverage we have crawled and indexed for this topic.
People also ask
Common questions on NVFP4, surfaced from across the indexed web.
What is quantization-aware distillation?
QAD uses the original full-precision model (teacher) to teach the quantized model (student). First, create a quantized model by running PTQ on the full-precision model. Then distill the frozen BF16 model into the quantized model using a KL divergence loss comparing the teacher’s and student’s logits. Figure 1 shows the two-stage QAD process used to build the Nemotron 3.5 Lightning NVFP4 checkpoint. The full-precision BF16 model serves as the frozen teacher and is also the starting point for Stage 1, a PTQ pass that quantizes weights to W4A16 to produce the quantized student. In Stage 2, the st
Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer | NVIDIA Technical Blog