about
A Learned Performance Model for the Tensor Processing Unit (arxiv.org)
2 points by matt_d on Aug 16, 2020 | hide | past | pdf | discuss on HN

In plain words: Instead of hand-written formulas for chip speed, it learns to predict runtimes from programs run on AI accelerator chips. It beats a hand-tuned cost model at picking chunk sizes and merging steps, and helps a search tool find faster programs when chip time is scarce.

Abstract · A Learned Performance Model for Tensor Processing Units

Accurate hardware performance models are critical to efficient code generation. They can be used by compilers to make heuristic decisions, by superoptimizers as a minimization objective, or by autotuners to find an optimal configuration for a specific program. However, they are difficult to develop because contemporary processors are complex, and the recent proliferation of deep learning accelerators has increased the development burden. We demonstrate a method of learning performance models from a corpus of tensor computation graph programs for Tensor Processing Unit (TPU) instances. We show that our learned model outperforms a heavily-optimized analytical performance model on two tasks -- tile-size selection and operator fusion -- and that it helps an autotuner discover faster programs in a setting where access to TPUs is limited or expensive.

Samuel J. Kaufman, Phitchaya Mangpo Phothilimthana, Yanqi Zhou, Charith Mendis, Sudip Roy, Amit Sabne, Mike Burrows
arXiv:2008.01040 · cs.PF, cs.LG · submitted Aug 3, 2020 · updated Mar 18, 2021
abstract · pdf · html · A version will appear in the Proceedings of the 4th MLSys Conference, San Jose, CA, USA, 2021

add comment on HN