about
To Tune or Not to Tune? Adapting Pretrained Representations to Diverse Tasks (arxiv.org)
2 points by iron0013 on Mar 19, 2019 | hide | past | pdf | discuss on HN

In plain words: They compared two ways to reuse a trained language model: freeze its weights and feed its output to a new classifier, or adjust all the weights on the new task. Which wins depends on how closely the new task matches the original training task.

Abstract

While most previous work has focused on different pretraining objectives and architectures for transfer learning, we ask how to best adapt the pretrained model to a given target task. We focus on the two most common forms of adaptation, feature extraction (where the pretrained weights are frozen), and directly fine-tuning the pretrained model. Our empirical results across diverse NLP tasks with two state-of-the-art models show that the relative performance of fine-tuning vs. feature extraction depends on the similarity of the pretraining and target tasks. We explore possible explanations for this finding and provide a set of adaptation guidelines for the NLP practitioner.

Matthew E. Peters, Sebastian Ruder, Noah A. Smith
arXiv:1903.05987 · cs.CL, cs.LG · submitted Mar 14, 2019 · updated Jun 11, 2019
abstract · pdf · html · Proceedings of the 4th Workshop on Representation Learning for NLP

add comment on HN