about
Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using F2 (arxiv.org)
2 points by MarcoDewey on Jun 2, 2025 | hide | past | pdf | discuss on HN

In plain words: It describes a tensor's hardware layout with a binary matrix that flips address bits, so layouts can be defined once and converted to others. In a GPU compiler it replaces converters that grow with each pair, cutting backend code and fixing bugs.

Abstract · Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using $\mathbb{F}_2$

Efficient tensor computation is a cornerstone of modern deep learning (DL) workloads, yet existing approaches struggle to achieve flexible and performant design and implementation of tensor layouts -- mappings between logical tensors and hardware resources. The increasing complexity of DL algorithms and hardware demands a generic and systematic approach to handling tensor layouts. In this work, we introduce Linear Layouts, a novel approach that models tensor layouts using linear algebra over $\mathbb{F}_2$. By representing tensor layouts as binary matrices acting on the bits of the hardware representation, our approach enables a generic layout definition -- as opposed to the classical case-by-case approach -- and allows for generic layout-to-layout conversions, eliminating the quadratic explosion that plagues existing solutions. We integrate linear layouts with Triton and demonstrate their effectiveness in optimizing individual Triton operators as well as kernels written in Triton. We also show that linear layouts reduce engineering effort in the compiler backend while fixing several bugs in Triton's legacy layout system.

Keren Zhou, Mario Lezcano, Adam Goucher, Akhmed Rakhmati, Jeff Niu, Justin Lebar, Pawel Szczerbuk, Peter Bell, Phil Tillet, Thomas Raoux, Zahi Moudallal
arXiv:2505.23819 · cs.PL, cs.AR, cs.DC, cs.PF · submitted May 28, 2025 · updated Mar 6, 2026
abstract · pdf · html

add comment on HN