In plain words: A tool that picks the best neural network design to build a fast stand-in for slow scientific simulations, even with little training data. Across many fields, the stand-ins ran up to 2 billion times faster than the originals and flag how uncertain they are.
Abstract · Building high accuracy emulators for scientific simulations with deep neural architecture search
Computer simulations are invaluable tools for scientific discovery. However, accurate simulations are often slow to execute, which limits their applicability to extensive parameter exploration, large-scale data analysis, and uncertainty quantification. A promising route to accelerate simulations by building fast emulators with machine learning requires large training datasets, which can be prohibitively expensive to obtain with slow simulations. Here we present a method based on neural architecture search to build accurate emulators even with a limited number of training data. The method successfully accelerates simulations by up to 2 billion times in 10 scientific cases including astrophysics, climate science, biogeochemistry, high energy density physics, fusion energy, and seismology, using the same super-architecture, algorithm, and hyperparameters. Our approach also inherently provides emulator uncertainty estimation, adding further confidence in their use. We anticipate this work will accelerate research involving expensive simulations, allow more extensive parameters exploration, and enable new, previously unfeasible computational discovery.
M. F. Kasim, D. Watson-Parris, L. Deaconu, S. Oliver, P. Hatfield, D. H. Froula, G. Gregori, M. Jarvis, S. Khatiwala, J. Korenaga, J. Topp-Mugglestone, E. Viezzer, et al.
arXiv:2001.08055 · stat.ML, cs.LG, physics.ao-ph, physics.comp-ph, physics.plasm-ph · submitted Jan 17, 2020 · updated Oct 8, 2020
abstract · pdf · html
It's not that it's not a useful method (it is). It's that it misrepresents the utility of the general "build an exact simulation, then train a regressor on it to make fast approximations" approach. It's very useful in certain situations (repeated calculations on similar parameters) and completely useless in others.
The key issue is that most of the slow models you'd want to use this on are highly non-linear. In certain regions of the parameter space, very small changes in input result in very large changes in output. This is fine, so long as you know where all of these regions are and can capture them in your training data. That's easier for some problems than others. Even assuming you do know how/where to collect dense training data, this approach is only useful within the bounds of the training data you collect. It's relatively uncommon (but not super rare) that you want to repeatedly run a complex simulation within the same parameter space. This method is great when you do want/need to do that, and useless otherwise.
You have to understand that what you've trained is little more than a look up table. Anything that claims it can actually learn highly non-linear and irregular behavior well outside of the training dataset's bounds is snake oil.
Understand where this general class of technique is useful and where it isn't and ignore overblown claims. The entire abstract here is overblown hogwash. The actual paper is relatively interesting, but I really wish folks would drop the absurd advertising language and focus on what distinguishes this from the hundreds of very similar studies/methods on this topic.