In plain words: A simple linear translation lets one language model's final internal states be read by another model's output layer, without retraining either. Accuracy mostly survives, and text generation works when the two share a tokenizer and the source model is larger.
Abstract · Ventriloquist LLMs: Linear Alignment of Late-Stage Representations
Independently trained language models often learn compatible late-stage representations, despite differences in training objectives, architectures, and data modalities. We ask how far this compatibility extends: can a simple affine map let one model's hidden states be read directly by another model's output head? In this work, we learn affine transformations between the final hidden states of independent models and evaluate them across embedding classification, out-of-distribution detection, and autoregressive text generation. Across model pairs, we find that downstream performance is largely preserved under alignment, with linearly mapped source representations retaining both the decision boundaries and the confidence structure of a target model's classifier. Additionally, we show for the first time that linear alignment sometimes enables text generation across independently trained models, decoding one model's hidden states through another's frozen output head without fine-tuning. This capability is far from universal, and characterizing when it holds is a contribution of our work. We find that success is governed by two factors: tokenizer overlap, which correlates strongly with generation quality, and source-model scale, below which quality degrades sharply. Transfer is also asymmetric, with strong-to-weak mappings substantially outperforming weak-to-strong ones, indicating that source representational capacity, rather than the target head, is the limiting factor. Finally, we characterize where in the network this compatibility arises. A layer-wise analysis localizes this compatibility to the final one or two layers, indicating a late-stage, output-space phenomenon rather than evidence of shared reasoning.
Matt Gorbett, Suman Jana
arXiv:2603.18908 · cs.AI · submitted Mar 19, 2026 · updated Sep 27, 2026
abstract · pdf · html
DeepSeek-R1 and Qwen2.5-72B have cleanly separable routing layers (ablating the refusal direction recovers accurate outputs), but Qwen3-8B doesn't - it confabulates, suggesting knowledge and suppression are jointly encoded. Whether a linear alignment method holds up may depend heavily on which of those architectural regimes you're in.