about
Just-in-time and distributed task representations in language models (arxiv.org)
1 point by PaulHoule on Sep 25, 2025 | hide | past | pdf | discuss on HN

In plain words: They tracked when a language model builds a summary of a task from examples, by pulling out numbers that hand it to another copy. Task identity stays readable throughout, but the transferable summary appears only at certain moments, gathering nearby evidence rather than building steadily.

Abstract

Many of language models' impressive capabilities originate from their in-context learning: based on instructions or examples, they can infer and perform new tasks without weight updates. In this work, we investigate when representations for new tasks are formed in language models, and how these representations change over the course of context. We study two different task representations: those that are ''transferrable'' -- vector representations that can transfer task contexts to another model instance, even without the full prompt -- and simpler representations of high-level task categories. We show that transferrable task representations evolve in non-monotonic and sporadic ways, while task identity representations persist throughout the context. Specifically, transferrable task representations exhibit a two-fold locality. They successfully condense evidence when more examples are provided in the context. But this evidence accrual process exhibits strong temporal locality along the sequence dimension, coming online only at certain tokens -- despite task identity being reliably decodable throughout the context. In some cases, transferrable task representations also show semantic locality, capturing a small task ''scope'' such as an independent subtask. Language models thus represent new tasks on the fly through both an inert, sustained sensitivity to the task and an active, just-in-time representation to support inference.

Yuxuan Li, Declan Campbell, Stephanie C. Y. Chan, Andrew Kyle Lampinen
arXiv:2509.04466 · cs.CL, cs.AI · submitted Aug 28, 2025 · updated Dec 1, 2025
abstract · pdf · html

add comment on HN