ALiBi Adds Distance Penalties Directly to Attention Logits
Attention with Linear Biases (ALiBi) changes a causal attention score before softmax by adding a penalty whose magnitude grows with token distance. Position is therefore represented in the score path rather than by adding a positional vector to each token representation. For a query at position i attending to a key at position j, one head can be written schematically as: score(i, j) = q_i · k_j / sqrt(d_k) - m_h * (i - j) for causal positions j <= i. The positive slope m_h is specific to attention head h. The causal mask still prevents access to future positions; the linear term changes the relative preference among positions that remain visible.