about
Overcoming Multi-Model Forgetting (Neural Architecture Search) (arxiv.org)
2 points by Wronskia on Nov 22, 2019 | hide | past | pdf | discuss on HN

In plain words: When networks that share some adjustable numbers are trained one after another, later training wipes out what earlier ones learned. This fix slows changes to shared numbers that mattered to earlier networks, keeping their accuracy and finding better designs for language and vision tasks.

Abstract · Overcoming Multi-Model Forgetting

We identify a phenomenon, which we refer to as multi-model forgetting, that occurs when sequentially training multiple deep networks with partially-shared parameters; the performance of previously-trained models degrades as one optimizes a subsequent one, due to the overwriting of shared parameters. To overcome this, we introduce a statistically-justified weight plasticity loss that regularizes the learning of a model's shared parameters according to their importance for the previous models, and demonstrate its effectiveness when training two models sequentially and for neural architecture search. Adding weight plasticity in neural architecture search preserves the best models to the end of the search and yields improved results in both natural language processing and computer vision tasks.

Yassine Benyahia, Kaicheng Yu, Kamil Bennani-Smires, Martin Jaggi, Anthony Davison, Mathieu Salzmann, Claudiu Musat
arXiv:1902.08232 · cs.LG, stat.ML · submitted Feb 21, 2019 · updated Mar 2, 2019
abstract · pdf · html

add comment on HN