In plain words: Instead of judging automated network-design tools by the final network's accuracy, this study scores the search step by comparing its picks with randomly chosen networks. The best tools did no better than random, and reusing one set of weights across candidates scrambled their rankings.
Abstract
Neural Architecture Search (NAS) aims to facilitate the design of deep networks for new tasks. Existing techniques rely on two stages: searching over the architecture space and validating the best architecture. NAS algorithms are currently compared solely based on their results on the downstream task. While intuitive, this fails to explicitly evaluate the effectiveness of their search strategies. In this paper, we propose to evaluate the NAS search phase. To this end, we compare the quality of the solutions obtained by NAS search policies with that of random architecture selection. We find that: (i) On average, the state-of-the-art NAS algorithms perform similarly to the random policy; (ii) the widely-used weight sharing strategy degrades the ranking of the NAS candidates to the point of not reflecting their true performance, thus reducing the effectiveness of the search process. We believe that our evaluation framework will be key to designing NAS strategies that consistently discover architectures superior to random ones.
Kaicheng Yu, Christian Sciuto, Martin Jaggi, Claudiu Musat, Mathieu Salzmann
arXiv:1902.08142 · cs.LG, stat.ML · submitted Feb 21, 2019 · updated Nov 22, 2019
abstract · pdf · html · We find that random policy in NAS works amazingly well and propose an evaluation framework to have a fair comparison. Adding additional results on standard CNN search space used for weight sharing and NASBench-101. 8 pages