about
Lost in the Middle at Birth: An Exact Theory of Transformer Context Bias (arxiv.org)
2 points by borundev 206 days ago | hide | past | pdf | 2 comments on HN

In plain words: The U-shaped memory dip comes from the decoder's basic shape: causal masking boosts the prompt's start, residual links pin the end, and the middle gets almost no influence before any training. Untrained models already show it, and normal training does not remove it.

Abstract · Lost in the Middle at Birth: An Exact Theory of Transformer Position Bias

The ``Lost in the Middle'' phenomenon -- a U-shaped performance curve where LLMs retrieve well from the beginning and end of a context but fail in the middle -- is widely attributed to learned Softmax artifacts or the distance-decay of positional encodings like RoPE. This paper makes a single, precise claim: \emph{the U-shape is already present at initialization, before any training or positional encoding takes effect.} It is an inherent geometric property of the causal decoder with residual connections. We model multi-layer causal attention as iterated powers of the Cesàro matrix and derive the exact closed-form influence density in the continuous limit. Causal masking forces a logarithmic divergence of gradient influence at the start of the prompt (the Primacy Tail), while residual connections create an isolated $\mathcal{O}(1)$ anchor at the final token (the Recency Delta). Between these extremes lies a factorial dead zone of order $\mathcal{O}(1/(H{-}1)!)$, where $H$ is the network depth, making middle-context retrieval and training structurally hostile. We validate empirically that untrained Qwen2 and GPT-2 architectures exhibit this U-shape at Step~0, and that it is identical with or without RoPE. Comparing initialized and pretrained networks, we show that standard training does not overcome the topological valley, confirming that the U-shape persists as an architectural baseline under standard pretraining objectives. We do not claim that this bias is insurmountable, nor that interventions such as RoPE modifications are useless. We establish what the baseline is and where it comes from, so that future efforts to overcome it can be precisely targeted.

Borun D Chowdhury
arXiv:2603.10123 · cs.LG, cs.AI, cs.CL · submitted Mar 10, 2026
abstract · pdf · html · 11 pages, 7 figures

add comment on HN

While the "Lost in the Middle" (LitM) phenomenon is well-documented empirically, it is usually attributed to training data distribution or the lack of long-range dependencies in common datasets.

In this paper, I show that LitM is actually present at initialization. By deriving an exact theory using the Jacobian Norm, I demonstrate that the characteristic U-shaped attention curve is a structural property of the Transformer architecture itself.

Key findings:

    Architectural Determinism: Even with random weights, the model is "born" prioritizing the start and end of sequences.

    Jacobian Norm Analysis: I use the Jacobian to measure how sensitive the output is to input tokens at different positions, showing a clear macroscopic bias.

    Pretraining vs. Initialization: I compare Qwen-2.5B at both stages to show that while training adds "content detectors" (local spikes), it does not remove the underlying global U-shape.
This suggests that "fixing" long-context retrieval might require rethinking the initialization or the softmax-attention geometry itself, rather than just scaling up training data.

I’m the author of the paper and would love to hear the community’s thoughts on whether this structural bias can ever truly be overcome within the standard Transformer paradigm.

I recommend asking a friend who's a better writer and mathematician than Claude Code to help you reorganize the paper so that there are no gaps in the argumentation and incorrect statements like "For a purely causal transformer without residuals, the gradient routed from the final token L to an earlier token j after H layers is given by the bottom row of the exponential Cesàro Matrix M^H" are replaced with mathematically correct descriptions.

Also have them check your experiments, because the description doesn't inspire confidence your (Claude's) implementation isn't flawed in ways that invalidate your results. In particular, "our experimental code utilizes a highly efficient one-pass scalar-probe surrogate" sounds fishy.