Open Internet by MindsNetOptimizing Transformer model size & inference beyond FP16 + ONNX (pruning/graph opt didn’t help much) [P]