about
Pretraining Data Mixtures Enable Narrow Model Selection Capabil. In Transformers (arxiv.org)
5 points by YeGoblynQueenne on Nov 6, 2023 | hide | past | pdf | 2 comments on HN

In plain words: Transformers were trained on simple input-output pairs from several task families, then asked to figure out new tasks from examples alone. They performed near-optimally when the task type was well covered in their training data, but failed on even simple tasks outside it.

Abstract · Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models

Transformer models, notably large language models (LLMs), have the remarkable ability to perform in-context learning (ICL) -- to perform new tasks when prompted with unseen input-output examples without any explicit model training. In this work, we study how effectively transformers can bridge between their pretraining data mixture, comprised of multiple distinct task families, to identify and learn new tasks in-context which are both inside and outside the pretraining distribution. Building on previous work, we investigate this question in a controlled setting, where we study transformer models trained on sequences of $(x, f(x))$ pairs rather than natural language. Our empirical results show transformers demonstrate near-optimal unsupervised model selection capabilities, in their ability to first in-context identify different task families and in-context learn within them when the task families are well-represented in their pretraining data. However when presented with tasks or functions which are out-of-domain of their pretraining data, we demonstrate various failure modes of transformers and degradation of their generalization for even simple extrapolation tasks. Together our results highlight that the impressive ICL abilities of high-capacity sequence models may be more closely tied to the coverage of their pretraining data mixtures than inductive biases that create fundamental generalization capabilities.

Steve Yadlowsky, Lyric Doshi, Nilesh Tripuraneni
arXiv:2311.00871 · cs.LG, cs.CL, stat.ML · submitted Nov 1, 2023
abstract · pdf · html

add comment on HN
Also discussed: Nov 2023 (65 points, 109 comments) · Nov 2023 (3 points, 1 comment)

I’m tired of folks conflating computation (elements transformed using operators towards an outcome, probabilistically or otherwise) and “intelligence,” of which mere optimization and operation on like elements is not dispositive.
Interesting paper. My confidence on the following idea isn't high, but I have an idea I don't think a lot of people agree with me on. The authors pretty convincingly show that transformers don't generalize well. I agree, but I would add that I don't think humans generalize well either, so I don't think showing that transformers don't generalize well precludes them from ultimately being "intelligent" (in some way). In fact, the inability to generalize to new situations is actually a core problem in learning for humans.

For example, think of a college freshmen who knows Algebra and begins Calculus. Even though they have all the fundamental "mental tools" available at their disposal to figure out how to prove that an infinite series converges (or not), I would guess that very few students can actually connect the concepts in the right way to see that. They have to be given examples and be shown how to use the knowledge they gained in algebra and how it applies to infinite series. Is their inability to generalize well evidence that they aren't intelligent?

Perhaps what we're really saying, and the real problem for us, is that these systems aren't intelligent enough to be useful to us. Just like we might say that a very smart student, like a von Neumann, would likely be able to figure out how to solve an infinite series without having seen it before. I think it's reasonable to say "we want transformers to generalize better than the average human."

To that end, I think we'll have to imitate the way that a very smart human solves new problems. For example, the students who are capable of solving an infinite series without having seen the concept before, might use a creative approach of trying different things to see if they can discover the pattern. So, my hunch would be that we will need to provide transformers with bolted-on sub-routines like "generate several hypotheses and test them to see which pattern fits the data best" before we can expect them to generalize well.

Tl;dr: Transformers don't generalize well, but I don't think humans generalize well either, so I don't think that fact precludes their ability to be intelligent in the limit. I think we'll have to imitate how humans generalize by giving the transformers additional "mental tools" before they can generalize like very intelligent humans.