Skip to content

Archive

Position Encoding

4 articles
Artificial Intelligence 24 Sep 2026 5 min read

Rotary Position Embedding Converts Absolute Indices into Relative Attention Phases

Self-attention can compare token content without assigning an order to token positions. Rotary Position Embedding (RoPE) inserts position into that comparison by rotating paired coordinates of queries and keys. Each token receives an absolute rotation angle, yet the query-key inner product reduces those two absolute angles to their difference. That algebraic cancellation is the central mechanism. RoPE does not add a position vector to the hidden state. It changes the orientation of query and key components before their dot product is evaluated.

Artificial Intelligence 24 Sep 2026 4 min read

RoPE Encodes Relative Offsets Through Rotated Query-Key Phases

Rotary Position Embedding (RoPE) applies position-dependent rotations to query and key coordinates before their attention dot product. The resulting score carries relative position through the phase difference between those rotations rather than through an additive position vector attached to the token representation. This distinction is structural. RoPE starts from absolute indices for each rotation, yet the query-key inner product can be written in terms of the offset between their positions.

Artificial Intelligence 24 Sep 2026 5 min read

ALiBi Adds Distance-Proportional Bias to Attention Scores

ALiBi changes an attention score before softmax rather than adding a positional vector to the token representation. For a causal transformer, a key farther behind the current query receives a larger negative offset. The offset is linear in token distance and uses a slope associated with the attention head. A simplified score for head h can be written as: score_h(i, j) = q_i k_j^T / sqrt(d) - m_h * (i - j) for an allowed causal pair with j <= i and positive slope m_h. The causal mask still decides which future positions are inaccessible. ALiBi changes the relative scores among positions that remain eligible.

Artificial Intelligence 14 Sep 2026 6 min read

Encode Token Distance with Rotary Position Embeddings

Transformer attention has no intrinsic notion that one token sits three positions before another. Rotary position embeddings, usually called RoPE, inject position into attention by rotating pairs of query and key coordinates before their dot product is computed. The mechanism is easy to reduce to a helper function, yet several details determine its actual behavior: queries and keys must use compatible rotations, each coordinate pair has its own angular frequency, offsets emerge through the dot product, and changing the position scale changes the geometry seen by attention.