Accelerate PyTorch INT8 Inference with New “X86” Quantization Backend...
By Authors Jiong Gong, Weiwen Xia Introduction INT8 quantization is one of the key features in PyTorch* for speeding up deep learning inference. By reducing the precision of weights and activations in…