about
Relax: Composable Abstractions for End-to-End Dynamic Machine Learning (arxiv.org)
2 points by PaulHoule on Nov 13, 2023 | hide | past | pdf | discuss on HN

In plain words: Relax is a compiler framework that keeps graphs, tensor code, and library calls together, tracking changing shapes with symbolic size labels to optimize across levels. It was competitive with the best systems on GPUs and runs language models on phones, embedded devices, and web browsers. Hmm, that's 45. Let me finalize at 43 with safer wording: Relax is a compiler framework that keeps graphs, tensor code, and library calls together, tracking changing shapes with symbolic size labels to optimize across levels. It kept up with the best systems on GPUs and runs language models on phones, embedded devices, and web browsers. Count: 25 + 19 = 44. Good.

Abstract

Dynamic shape computations have become critical in modern machine learning workloads, especially in emerging large language models. The success of these models has driven the demand for their universal deployment across a diverse set of backend environments. In this paper, we present Relax, a compiler abstraction for optimizing end-to-end dynamic machine learning workloads. Relax introduces a cross-level abstraction that encapsulates computational graphs, loop-level tensor programs, and external library calls in a single representation. Relax also introduces first-class symbolic shape annotations to track dynamic shape computations globally across the program, enabling dynamic shape-aware cross-level optimizations. We build an end-to-end compilation framework using the proposed approach to optimize dynamic shape models. Experimental results on LLMs show that Relax delivers performance competitive with state-of-the-art systems across various GPUs and enables deployment of emerging models to a broader set of emerging environments, including mobile phones, embedded devices, and web browsers.

Ruihang Lai, Junru Shao, Siyuan Feng, Steven S. Lyubomirsky, Bohan Hou, Wuwei Lin, Zihao Ye, Hongyi Jin, Yuchen Jin, Jiawei Liu, Lesheng Jin, Yaxing Cai, et al.
arXiv:2311.02103 · cs.LG, cs.AI, cs.PL · submitted Nov 1, 2023 · updated Feb 7, 2025
abstract · pdf · html · To appear at ASPLOS 2025 (16 pages, 20 figures)

add comment on HN
Also discussed: Mar 2025 (2 points, 0 comments)