How to Optimize Transformer-Based Models for Low-Precision Training | NVIDIA Technical Blog
… Use prequantized results to understand whether quantization overhead is the bottleneck, or to compare raw tensor core throughput across precisions independent of the quantization implementation. …
