about
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 (arxiv.org)
2 points by rbanffy 152 days ago | hide | past | pdf | discuss on HN

In plain words: They compared 2.5 years of error logs from 1,056 A100 and H100 GPUs to see which parts break and how often. On the newer H100s, the time between memory errors was 3.2 times shorter than on A100s, even though their other hardware broke less.

Abstract · Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs

This study characterizes GPU resilience in Delta, a large-scale AI system that consists of 1,056 A100 and H100 GPUs, with over 1,300 petaflops of peak throughput. We used 2.5 years of operational data (11.7 million GPU hours) on GPU errors. Our major findings include: (i) H100 GPU memory resilience is worse than A100 GPU memory, with 3.2x lower per-GPU MTBE for memory errors, (ii) The GPU memory error-recovery mechanisms on H100 GPUs are insufficient to handle the increased memory capacity, (iii) H100 GPUs demonstrate significantly improved GPU hardware resilience over A100 GPUs with respect to critical hardware components, (iv) GPU errors on both A100 and H100 GPUs frequently result in job failures due to the lack of robust recovery mechanisms at the application level, and (v) We project the impact of GPU node availability on larger-scales and find that significant overprovisioning of 5% is necessary to handle GPU failures.

Shengkun Cui, Archit Patke, Hung Nguyen, Aditya Ranjan, Ziheng Chen, Phuong Cao, Gregory Bauer, Brett Bode, Catello Di Martino, Saurabh Jha, Chandra Narayanaswami, Daby Sow, et al.
arXiv:2503.11901 · cs.DC, cs.AI · submitted Mar 14, 2025 · updated Dec 10, 2025
abstract · pdf · html

add comment on HN