about
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking (arxiv.org)
2 points by matt_d on Jan 24, 2025 | hide | past | pdf | discuss on HN

In plain words: Tiny speed tests probe each new part of NVIDIA's Hopper graphics chip—its memory, matrix units, and data-moving hardware—to measure how fast each really is. The new 8-bit math mode ran nearly twice as fast as the usual 16-bit mode.

Abstract · Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis

This study presents a comprehensive multi-level analysis of the NVIDIA Hopper GPU architecture, focusing on its performance characteristics and novel features. We benchmark Hopper's memory subsystem, highlighting improvements in the L2 partitioned cache and global memory access compared to Ampere and Ada Lovelace. The evaluation of Hopper's fourth-generation tensor cores reveals the benefits of FP8 precision and asynchronous wgmma instructions for matrix operations. Additionally, we investigate the performance of DPX instructions for dynamic programming, distributed shared memory (DSM) for inter-SM communication, and the Tensor Memory Accelerator (TMA) for asynchronous data movement. Through multi-level evaluation, we discover that the Hopper architecture demonstrates significant acceleration potential in real-world applications. For instance, the asynchronous programming model supported by TMA achieves a 1.5x speedup in matrix multiplication, FP8 delivers nearly double the performance of FP16, and DPX instructions accelerate a computational biology algorithm by at least 4.75x. Our findings provide actionable insights for optimizing compute-intensive workloads, from AI training to bioinformatics, on Hopper GPUs.

Weile Luo, Ruibo Fan, Zeyu Li, Dayou Du, Hongyuan Liu, Qiang Wang, Xiaowen Chu
arXiv:2501.12084 · cs.DC, cs.AR, cs.PF · submitted Jan 21, 2025 · updated Sep 4, 2025
abstract · pdf · html · arXiv admin note: substantial text overlap with arXiv:2402.13499

add comment on HN