Skip to content

Archive

Positional Encoding

2 articles
Artificial Intelligence 24 Sep 2026 5 min read

ALiBi Adds Distance Penalties Directly to Attention Logits

Attention with Linear Biases (ALiBi) changes a causal attention score before softmax by adding a penalty whose magnitude grows with token distance. Position is therefore represented in the score path rather than by adding a positional vector to each token representation. For a query at position i attending to a key at position j, one head can be written schematically as: score(i, j) = q_i · k_j / sqrt(d_k) - m_h * (i - j) for causal positions j <= i. The positive slope m_h is specific to attention head h. The causal mask still prevents access to future positions; the linear term changes the relative preference among positions that remain visible.

Artificial Intelligence 03 Sep 2026 10 min read

Positional Information in Transformer Models

Self-attention can compare every token with other tokens in a context, but the comparison alone does not tell the model where those tokens occur. A sentence is not just a collection of words: changing their order can change the meaning. Transformer models therefore need a way to represent positional information. This mechanism lets the network distinguish, for example, the first occurrence of a token from a later occurrence and reason about relationships such as “the previous token” or “far earlier in the document.”