about
BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding (arxiv.org)
78 points by liviosoares on Oct 12, 2018 | hide | past | pdf | 5 comments on HN

In plain words: It learns from unlabeled text by reading each word with the words before and after it, then needs one added layer to adapt to a task. It beat the previous best on eleven language tasks, scoring 80.5% on a broad test, 7.7 points higher.

Abstract · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models, BERT is designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers. As a result, the pre-trained BERT model can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks, such as question answering and language inference, without substantial task-specific architecture modifications. BERT is conceptually simple and empirically powerful. It obtains new state-of-the-art results on eleven natural language processing tasks, including pushing the GLUE score to 80.5% (7.7% point absolute improvement), MultiNLI accuracy to 86.7% (4.6% absolute improvement), SQuAD v1.1 question answering Test F1 to 93.2 (1.5 point absolute improvement) and SQuAD v2.0 Test F1 to 83.1 (5.1 point absolute improvement).

Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova
arXiv:1810.04805 · cs.CL · submitted Oct 11, 2018 · updated May 24, 2019
abstract · pdf · html

add comment on HN
Also discussed: Apr 2023 (3 points, 0 comments)

Fantastic!

> The code and pre-trained model will be available at https://goo.gl/language/bert. Will be released before the end of October 2018.

"We demonstrate the importance of bidirectional pre-training for language representations". can some one help me understand what bidirectional and pre-trained means?
* bidirectional - build representations of the current word by looking into both the future and the past

* pre-trained - train on lots of language modelling data (e.g. billions of words of wikipedia) and then train on the task you really care about but starting from the parameters learnt from the language modelling task.

Is this comparable to fast.ai's ULMfit, which also promises an unsupervised pretrained model, that can be tuned to best state-of-the-art NLP tasks?
The big picture is similar. But ULMfit uses amd-lstm for the language modeling, bert uses masked LM instead. Bert has some other tricks like sentence prediction as well.