about
Why M Heads Are Better Than One: Training a Diverse Ensemble of Deep Networks (arxiv.org)
2 points by beedotkiran on Dec 29, 2016 | hide | past | pdf | discuss on HN

In plain words: Instead of training several networks separately and averaging their answers, this builds them to share parts and trains them together with a loss that pushes them to disagree. Such diverse teams answered more questions correctly on at least one member than classically built ensembles.

Abstract · Why M Heads are Better than One: Training a Diverse Ensemble of Deep Networks

Convolutional Neural Networks have achieved state-of-the-art performance on a wide range of tasks. Most benchmarks are led by ensembles of these powerful learners, but ensembling is typically treated as a post-hoc procedure implemented by averaging independently trained models with model variation induced by bagging or random initialization. In this paper, we rigorously treat ensembling as a first-class problem to explicitly address the question: what are the best strategies to create an ensemble? We first compare a large number of ensembling strategies, and then propose and evaluate novel strategies, such as parameter sharing (through a new family of models we call TreeNets) as well as training under ensemble-aware and diversity-encouraging losses. We demonstrate that TreeNets can improve ensemble performance and that diverse ensembles can be trained end-to-end under a unified loss, achieving significantly higher "oracle" accuracies than classical ensembles.

Stefan Lee, Senthil Purushwalkam, Michael Cogswell, David Crandall, Dhruv Batra
arXiv:1511.06314 · cs.CV, cs.LG, cs.NE · submitted Nov 19, 2015
abstract · pdf · html

add comment on HN