Open Internet by MindsNet
Optimizing Real-time Small-Batch Inference
Real-time small-batch inference workloads face significant bottlenecks due to runtime overhead, including fragmented small kernels, norm/residual/activation boundaries, and quantization/dequantization overhead. These bottlenecks are exacerbated in robotics, VLA, and world models where batch size is typically 1.
Computing & Technology, Computer Science, Machine Learning