Open Internet by MindsNet
Inefficient cuBLAS Performance on RTX GPUs
cuBLAS exhibits a performance bug on RTX GPUs, achieving only ~40% of available compute for batched FP32 workloads. This inefficiency affects various RTX non-Pro GPUs, including the RTX 5090.
Computing & Technology, Computer Science, Machine Learning