Open Internet by MindsNetRewriting model inference with CUDA kernels: the bottleneck was not just GEMM [P]