about
The Distributed Tensor Algebra Compiler (2022) (arxiv.org)
40 points by yeesian on Jun 16, 2023 | hide | past | pdf | 6 comments on HN

In plain words: A compiler lets programmers separately describe how tensor data is laid out and how the math is split across CPUs and GPUs, then generates code. Its matrix multiply matches hand-tuned code on 256 nodes and beats other systems 1.8x to 3.7x on larger tensor operations.

Abstract · DISTAL: The Distributed Tensor Algebra Compiler

We introduce DISTAL, a compiler for dense tensor algebra that targets modern distributed and heterogeneous systems. DISTAL lets users independently describe how tensors and computation map onto target machines through separate format and scheduling languages. The combination of choices for data and computation distribution creates a large design space that includes many algorithms from both the past (e.g., Cannon's algorithm) and the present (e.g., COSMA). DISTAL compiles a tensor algebra domain specific language to a distributed task-based runtime system and supports nodes with multi-core CPUs and multiple GPUs. Code generated by DISTAL is competitive with optimized codes for matrix multiply on 256 nodes of the Lassen supercomputer and outperforms existing systems by between 1.8x to 3.7x (with a 45.7x outlier) on higher order tensor operations.

Rohan Yadav, Alex Aiken, Fredrik Kjolstad
arXiv:2203.08069 · cs.PL, cs.DC · submitted Mar 15, 2022 · updated Mar 17, 2022
abstract · pdf · html

add comment on HN

Would anyone be interested in discussing this paper together, especially through the lens of "how can we schedule algebraic expressions and checkpoint computed progress across a heterogenous pool of consumer/donated machines?"

I'm an infrastructure engineer, mostly focused on databases, data pipelines, ML infra for the past 10-15 years. Even when designing homogenous compute clusters, I had to dig in and understand compiler-level implementations in MLIR and LLVM. I'm not a compiler expert by any measure, but know just enough to be dangerous and curious about (safely) scheduling computations across a pool of volunteers machines. Seems especially important to chew on now, with training of foundational LLM weights costing 7-9 figures.

(This is more of a link-dump than a paper discussion --)

For the line of inquiry w.r.t tensor compilers and MLIR/LLVM (linalg, polyhedral, [sparse_]tensor, etc), I personally found the following really helpful: https://news.ycombinator.com/item?id=25545373 (links to a survey), https://github.com/merrymercy/awesome-tensor-compilers

I also have an interest in the community more widely associated with pandas/dataframes-like languages (e.g. modin/dask/ray/polars/ibis) with substrait/calcite/arrow their choice of IR. Some links: https://github.com/modin-project/modin, https://github.com/dask/dask/issues/8980, https://news.ycombinator.com/item?id=16510610, https://news.ycombinator.com/item?id=35521785

I broadly classify them as such since the former has a stronger disposition towards linear/tensor-algebra, while the latter towards relational algebra, and it isn't yet clear (to me) how well innovations in one carry over to the other (if they do), and hence I'm also curious to hear more about proposals for a unified language across linalg and relational alg (e.g. https://news.ycombinator.com/item?id=36349015).

I'm particularly interested in pandas precisely because it seems to be right at the intersection of both forms of algebra (and draws a strong reaction from people who are familiar/comfortable with one community and not the other). See e.g. https://datapythonista.me/blog/pandas-20-and-the-arrow-revol... and https://wesmckinney.com/blog/apache-arrow-pandas-internals/

I would love to be a part of this if that's ok :). I love this mainly as someone interested in compilers and infra
I am the author of this paper -- kind of crazy to see it posted here on its own! I'll hang around to try and answer questions
This kind of stuff is the future, but it would be much better if they targeted custom MLIR dialects instead of creating yet another incompatible compiler stack.
I agree! Much of this work was done as part of the overarching TACO project (https://github.com/tensor-compiler/taco), in an attempt to distribute sparse tensor computations (https://rohany.github.io/publications/sc2022-spdistal.pdf). MLIR recently (~mid 2022) began implementing the ideas from TACO into a "sparse tensor" dialect, so perhaps some of these ideas could make it into there. I'm working with MLIR these days, and if I could re-do the project now I would probably utilize and targetb the MLIR linalg infrastructure!