about
Bam Born-Again Multi-Task Networks for Natural Language Understanding (arxiv.org)
1 point by sel1 on Jul 11, 2019 | hide | past | pdf | discuss on HN

In plain words: Separate models trained on one task each teach a shared multi-task model, then the teaching slowly fades so it learns from the real answers. This beat both the single-task teachers and ordinary multi-task training on a language understanding test suite.

Abstract · BAM! Born-Again Multi-Task Networks for Natural Language Understanding

It can be challenging to train multi-task neural networks that outperform or even match their single-task counterparts. To help address this, we propose using knowledge distillation where single-task models teach a multi-task model. We enhance this training with teacher annealing, a novel method that gradually transitions the model from distillation to supervised learning, helping the multi-task model surpass its single-task teachers. We evaluate our approach by multi-task fine-tuning BERT on the GLUE benchmark. Our method consistently improves over standard single-task and multi-task training.

Kevin Clark, Minh-Thang Luong, Urvashi Khandelwal, Christopher D. Manning, Quoc V. Le
arXiv:1907.04829 · cs.CL · submitted Jul 10, 2019
abstract · pdf · html · ACL 2019

add comment on HN