Open Internet by MindsNet
Unexpected INT8 quantization accuracy
A deep learning model shows better inference accuracy with INT8 quantization than FP16, contradicting expectations. The model was exported via ONNX and used post-training quantization for INT8. The cause of this phenomenon is unclear.
Computing & Technology, Computer Science, Machine Learning