about
In-Context Learning Creates Task Vectors (arxiv.org)
3 points by famouswaffles on Oct 26, 2023 | hide | past | pdf | discuss on HN

In plain words: When a language model learns from examples in a prompt, it squeezes them into a single "task vector" that steers its answers. Tests across models and tasks showed its behavior can be rebuilt from just that vector and the question, not the example set.

Abstract

In-context learning (ICL) in Large Language Models (LLMs) has emerged as a powerful new learning paradigm. However, its underlying mechanism is still not well understood. In particular, it is challenging to map it to the "standard" machine learning framework, where one uses a training set $S$ to find a best-fitting function $f(x)$ in some hypothesis class. Here we make progress on this problem by showing that the functions learned by ICL often have a very simple structure: they correspond to the transformer LLM whose only inputs are the query $x$ and a single "task vector" calculated from the training set. Thus, ICL can be seen as compressing $S$ into a single task vector $\boldsymbolθ(S)$ and then using this task vector to modulate the transformer to produce the output. We support the above claim via comprehensive experiments across a range of models and tasks.

Roee Hendel, Mor Geva, Amir Globerson
arXiv:2310.15916 · cs.CL · submitted Oct 24, 2023
abstract · pdf · html · Accepted at Findings of EMNLP 2023

add comment on HN