about
Learning without training: The implicit dynamics of in-context learning (arxiv.org)
1 point by simonpure on Jul 25, 2025 | hide | past | pdf | discuss on HN

In plain words: A transformer's attention layer and the layer after it work together to quietly rewrite that later layer's weights from the prompt's examples. The math shows a run with examples matches one without them, where the weights carry a small update encoding the context.

Abstract

One of the most striking features of Large Language Models (LLMs) is their ability to learn in-context. Namely at inference time an LLM is able to learn new patterns without any additional weight update when these patterns are presented in the form of examples in the prompt, even if these patterns were not seen during training. The mechanisms through which this can happen are still largely unknown. In this work, we show that the stacking of a self-attention layer with an MLP allows the transformer block to implicitly modify the weights of the MLP layer according to the context. We argue through theoretical analysis and experimentation that this simple mechanism may help explain why LLMs demonstrate capabilities of in-context learning, beyond what is captured during training. Specifically, we show that a standard forward pass with context is mathematically equivalent to a forward pass without context but with the MLP weights updated by a minimal low-rank update representing the context.

Benoit Dherin, Michael Munn, Hanna Mazzawi, Michael Wunder, Javier Gonzalvo
arXiv:2507.16003 · cs.CL, cs.LG · submitted Jul 21, 2025 · updated Jun 2, 2026
abstract · pdf · html

add comment on HN
Also discussed: Jul 2025 (2 points, 0 comments) · Jul 2025 (17 points, 1 comment) · Jul 2025 (1 point, 0 comments)