In plain words: A trick that fills in blanks step by step lets a fill-in-the-blank language model answer open-ended tasks with no new training, something usually done by next-word-prediction models. The two kinds beat each other on different task types, so each has its own strengths.
Abstract · BERTs are Generative In-Context Learners
While in-context learning is commonly associated with causal language models, such as GPT, we demonstrate that this capability also 'emerges' in masked language models. Through an embarrassingly simple inference technique, we enable an existing masked model, DeBERTa, to perform generative tasks without additional training or architectural changes. Our evaluation reveals that the masked and causal language models behave very differently, as they clearly outperform each other on different categories of tasks. These complementary strengths suggest that the field's focus on causal models for in-context learning may be limiting - both architectures can develop these capabilities, but with distinct advantages; pointing toward promising hybrid approaches that combine the strengths of both objectives.
David Samuel
arXiv:2406.04823 · cs.CL, cs.AI · submitted Jun 7, 2024 · updated Oct 31, 2024
abstract · pdf · html · 26 pages, NeurIPS 2024
We found Google's T5 models which were released in 2019, pre-GPT-3, were "secretly" capable of in-context learning with a simple inference technique.
Given they use a bidirectional MLM (Masked Language Modeling) objective, it wasn't obvious how to do it, but MLM objectives are known to produce better language representations than causal (next token prediction) objectives. We were able to outperform much larger sized GPT-3 models or get very close to their performance with far smaller T5 models.