In plain words: Deep learning compilers turn models written in one framework into fast code for many chips. Unlike earlier surveys that just list tools, this one dissects their shared design: models pass through several internal code stages, optimized on the way in and out.
Abstract
The difficulty of deploying various deep learning (DL) models on diverse DL hardware has boosted the research and development of DL compilers in the community. Several DL compilers have been proposed from both industry and academia such as Tensorflow XLA and TVM. Similarly, the DL compilers take the DL models described in different DL frameworks as input, and then generate optimized codes for diverse DL hardware as output. However, none of the existing survey has analyzed the unique design architecture of the DL compilers comprehensively. In this paper, we perform a comprehensive survey of existing DL compilers by dissecting the commonly adopted design in details, with emphasis on the DL oriented multi-level IRs, and frontend/backend optimizations. Specifically, we provide a comprehensive comparison among existing DL compilers from various aspects. In addition, we present detailed analysis on the design of multi-level IRs and illustrate the commonly adopted optimization techniques. Finally, several insights are highlighted as the potential research directions of DL compiler. This is the first survey paper focusing on the design architecture of DL compilers, which we hope can pave the road for future research towards DL compiler.
Mingzhen Li, Yi Liu, Xiaoyan Liu, Qingxiao Sun, Xin You, Hailong Yang, Zhongzhi Luan, Lin Gan, Guangwen Yang, Depei Qian
arXiv:2002.03794 · cs.DC, cs.LG, cs.PF · submitted Feb 6, 2020 · updated Aug 28, 2020
abstract · pdf · html
From what little I've been able to understand, there are a few different ways of handling this, but none are all that great: Onnx.js is somewhat stale and missing many PyTorch operators, Tensorflow.js works great but requires either porting to TF or doing multiple conversions (PT -> Onnx -> TF), and then there is TVM which sent me down this rabbit-hole of deep learning compilation. It seems possible to use TVM's IR to compile PyTorch (along with many other frameworks) into Wasm or WebGPU. Haven't tried yet, but this domain is fascinating!