Loading Open Internet
    Rewriting model inference with CUDA kernels: the bottleneck was not just GEMM [P]