In plain words: Fast 8-bit matrix units on modern chips can be split to do full-precision matrix multiplication. On a Grace Hopper chip it ran single-precision math 3 times faster than the usual full-precision code, and more than twice as fast as earlier emulation tricks.
Abstract · High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines
Recent architectures integrate high-performance and power-efficient matrix engines. These engines demonstrate remarkable performance in low-precision matrix multiplication, which is crucial in deep learning. Several techniques have been proposed to emulate single- and double-precision general matrix-matrix multiplication (SGEMM and DGEMM, respectively) by leveraging such low-precision matrix engines. In this study, we present emulation methods that significantly outperforms conventional approaches. On a GH200 Grace Hopper Superchip, the proposed DGEMM emulation achieves a 1.4x speedup and a 43% improvement in power efficiency compared to native DGEMM for sufficiently large problems. The proposed SGEMM emulation achieves a 3.0x speedup and a 154% improvement in power efficiency compared to native SGEMM for sufficiently large problems. Furthermore, compared to conventional emulation methods, the proposed emulation achieves more than 2x higher performance and superior power efficiency.
Yuki Uchino, Katsuhisa Ozaki, Toshiyuki Imamura
arXiv:2508.03984 · cs.DC · submitted Aug 6, 2025 · updated Aug 8, 2025
abstract · pdf · html · 8 pages, 9 figures