In plain words: The U-shaped memory dip comes from the decoder's basic shape: causal masking boosts the prompt's start, residual links pin the end, and the middle gets almost no influence before any training. Untrained models already show it, and normal training does not remove it.
Abstract · Lost in the Middle at Birth: An Exact Theory of Transformer Position Bias
The ``Lost in the Middle'' phenomenon -- a U-shaped performance curve where LLMs retrieve well from the beginning and end of a context but fail in the middle -- is widely attributed to learned Softmax artifacts or the distance-decay of positional encodings like RoPE. This paper makes a single, precise claim: \emph{the U-shape is already present at initialization, before any training or positional encoding takes effect.} It is an inherent geometric property of the causal decoder with residual connections. We model multi-layer causal attention as iterated powers of the Cesàro matrix and derive the exact closed-form influence density in the continuous limit. Causal masking forces a logarithmic divergence of gradient influence at the start of the prompt (the Primacy Tail), while residual connections create an isolated $\mathcal{O}(1)$ anchor at the final token (the Recency Delta). Between these extremes lies a factorial dead zone of order $\mathcal{O}(1/(H{-}1)!)$, where $H$ is the network depth, making middle-context retrieval and training structurally hostile. We validate empirically that untrained Qwen2 and GPT-2 architectures exhibit this U-shape at Step~0, and that it is identical with or without RoPE. Comparing initialized and pretrained networks, we show that standard training does not overcome the topological valley, confirming that the U-shape persists as an architectural baseline under standard pretraining objectives. We do not claim that this bias is insurmountable, nor that interventions such as RoPE modifications are useless. We establish what the baseline is and where it comes from, so that future efforts to overcome it can be precisely targeted.
Borun D Chowdhury
arXiv:2603.10123 · cs.LG, cs.AI, cs.CL · submitted Mar 10, 2026
abstract · pdf · html · 11 pages, 7 figures
In this paper, I show that LitM is actually present at initialization. By deriving an exact theory using the Jacobian Norm, I demonstrate that the characteristic U-shaped attention curve is a structural property of the Transformer architecture itself.
Key findings:
This suggests that "fixing" long-context retrieval might require rethinking the initialization or the softmax-attention geometry itself, rather than just scaling up training data.I’m the author of the paper and would love to hear the community’s thoughts on whether this structural bias can ever truly be overcome within the standard Transformer paradigm.