We derive a mathematical theory showing how local layers create this recency bias, and how it propagates through the residual stream, norms, and MLPs until it reaches the NoPE attention logits.
We then validate this theory on both randomly initialized and trained networks.
