In plain words: A community review gathers ways to run machine learning inside live experiments, so data is analyzed fast enough to steer the science as it happens. It maps examples across fields, tricks for fast, efficient algorithms and hardware to run them, plus shared challenges.
Abstract
In this community review report, we discuss applications and techniques for fast machine learning (ML) in science -- the concept of integrating power ML methods into the real-time experimental data processing loop to accelerate scientific discovery. The material for the report builds on two workshops held by the Fast ML for Science community and covers three main areas: applications for fast ML across a number of scientific domains; techniques for training and implementing performant and resource-efficient ML algorithms; and computing architectures, platforms, and technologies for deploying these algorithms. We also present overlapping challenges across the multiple scientific domains where common solutions can be found. This community report is intended to give plenty of examples and inspiration for scientific discovery through integrated and accelerated ML solutions. This is followed by a high-level overview and organization of technical advances, including an abundance of pointers to source material, which can enable these breakthroughs.
Allison McCarn Deiana, Nhan Tran, Joshua Agar, Michaela Blott, Giuseppe Di Guglielmo, Javier Duarte, Philip Harris, Scott Hauck, Mia Liu, Mark S. Neubauer, Jennifer Ngadiuba, Seda Ogrenci-Memik, et al.
arXiv:2110.13041 · cs.LG, cs.AR, physics.data-an, physics.ins-det · submitted Oct 25, 2021
abstract · pdf · html · 66 pages, 13 figures, 5 tables
"[...] the concept of integrating power ML methods into the real-time experimental data processing loop to accelerate scientific discovery"
data scientists study the science of what the business does (laundry delivery, manufacturing TVs, tracking patient health), and the point of science is insight and understanding from data to build a theory of how it all works
what this article highlights is that ML can be an exceptional tool for discovery. this is in stark contrast to how ML is usually deployed, which is some big analytics or product effort. the obvious big reason for that is the infra is expensive, the know-how is lacking, and the data sucks. well, that all is quickly changing and we're gonna see folks weaving in ML to bolster their workflows in a much bigger way
great to see academic scientists leading the charge here too. they stand to gain a lot from that perspective