In plain words: A small language turns descriptions of how machines share training or predictions into code on a parallel toolkit to test chips and layouts. It ran on x86, ARM, and RISC-V, which usual tools struggle with, and includes the first public PyTorch build for RISC-V.
Abstract
Decentralised Machine Learning (DML) enables collaborative machine learning without centralised input data. Federated Learning (FL) and Edge Inference are examples of DML. While tools for DML (especially FL) are starting to flourish, many are not flexible and portable enough to experiment with novel processors (e.g., RISC-V), non-fully connected network topologies, and asynchronous collaboration schemes. We overcome these limitations via a domain-specific language allowing us to map DML schemes to an underlying middleware, i.e. the FastFlow parallel programming library. We experiment with it by generating different working DML schemes on x86-64 and ARM platforms and an emerging RISC-V one. We characterise the performance and energy efficiency of the presented schemes and systems. As a byproduct, we introduce a RISC-V porting of the PyTorch framework, the first publicly available to our knowledge.
Gianluca Mittone, Nicolò Tonci, Robert Birke, Iacopo Colonnelli, Doriana Medić, Andrea Bartolini, Roberto Esposito, Emanuele Parisi, Francesco Beneventi, Mirko Polato, Massimo Torquati, Luca Benini, et al.
arXiv:2302.07946 · cs.DC, cs.LG · submitted Feb 15, 2023 · updated Oct 18, 2023
abstract · pdf · html · This paper is the accepted version of ACM copyrighted material presented at the CF'23 conference in Bologna, Italy