Transformer Attention Weights: What They Show and What They Do Not
Transformer attention maps are visually compelling. A token appears to assign most of its attention to another token, so it is tempting to conclude that the second token caused the model’s prediction. That conclusion is stronger than the data supports. An attention weight has a precise local meaning: inside one attention operation, it controls how strongly a query mixes information from available value vectors. A complete transformer prediction, however, also depends on value vectors, residual connections, feed-forward layers, normalization, later layers, and often many attention heads. A large weight is therefore evidence about one routing operation, not a complete causal explanation.