Skip to content

Archive

RoPE

8 articles
Artificial Intelligence 24 Sep 2026 5 min read

Rotary Position Embedding Converts Absolute Indices into Relative Attention Phases

Self-attention can compare token content without assigning an order to token positions. Rotary Position Embedding (RoPE) inserts position into that comparison by rotating paired coordinates of queries and keys. Each token receives an absolute rotation angle, yet the query-key inner product reduces those two absolute angles to their difference. That algebraic cancellation is the central mechanism. RoPE does not add a position vector to the hidden state. It changes the orientation of query and key components before their dot product is evaluated.

Artificial Intelligence 24 Sep 2026 5 min read

RoPE Position Interpolation Compresses Indices Before Rotation

Rotary position embeddings apply a position-dependent rotation to query and key components before their dot product is evaluated. If a model was trained with positions up to a reference length, simply sending much larger position indices into the same rotation rule can place attention in angular regimes that were not represented in that training range. Position interpolation changes the coordinate supplied to RoPE. For an extension factor s > 1, a simplified form maps

Artificial Intelligence 24 Sep 2026 4 min read

RoPE Encodes Relative Offsets Through Rotated Query-Key Phases

Rotary Position Embedding (RoPE) applies position-dependent rotations to query and key coordinates before their attention dot product. The resulting score carries relative position through the phase difference between those rotations rather than through an additive position vector attached to the token representation. This distinction is structural. RoPE starts from absolute indices for each rotation, yet the query-key inner product can be written in terms of the offset between their positions.

Artificial Intelligence 24 Sep 2026 6 min read

Position Interpolation Compresses RoPE Indices into the Original Context Range

A RoPE-based Transformer associates token positions with rotations whose angles depend on the position index and per-dimension frequencies. Feeding a sequence beyond the context range used during training pushes those rotations to position indices the model did not encounter in that regime. Position Interpolation changes that boundary by scaling the extended indices back into the original range before the rotary transformation is applied. For an original context limit L and a target context L' > L, a simplified linear mapping is:

Artificial Intelligence 23 Sep 2026 5 min read

RoPE Position Shifts Preserve Relative Phase but Change Absolute Rotation

Rotary position embeddings, commonly called RoPE, inject position into attention by rotating pairs of query and key channels with angles determined by token position. The rotation happens before the query-key dot product. As a result, position is not represented by adding a standalone vector to each token representation. This distinction matters during inference. A token’s cached key already contains the rotation associated with its assigned position. Reusing that key under a different sequence coordinate without an equivalent transformation changes the attention computation, even when the token content is identical.

Artificial Intelligence 23 Sep 2026 4 min read

RoPE Position Interpolation Compresses Relative Angles Across Longer Contexts

Rotary position embeddings encode position by rotating query and key components with angles derived from token indices. When a model is asked to operate beyond the position range used during its original training, those angles can enter a regime the model did not encounter. Position interpolation addresses that boundary by mapping a longer sequence into a smaller position-coordinate range before the rotary angles are computed. The mechanism does not add extra tokens to the model’s architectural state. It changes the coordinates supplied to the positional transform. That distinction matters because a longer accepted input length and reliable behavior at that length are separate properties.

Artificial Intelligence 23 Sep 2026 4 min read

RoPE KV Caches Preserve Token Position Across Incremental Decoding

In a decoder using rotary position embeddings, a key written to the KV cache already carries the rotation associated with its token position. Incremental decoding can reuse that key directly. Applying the current token position to the cached key again changes the attention geometry. This makes position state part of the cache contract even when the cache API appears to store only tensors. Rotation is applied before a key enters attention For one two-dimensional component pair, RoPE applies a position-dependent rotation. Writing (R_m) for the rotation at position (m), a query and key become

Artificial Intelligence 12 Sep 2026 9 min read

Extend RoPE Context Windows with Position Interpolation

Extend RoPE Context Windows with Position Interpolation A RoPE-based language model trained on sequences up to a fixed length can behave poorly when inference suddenly asks it to process much larger position indices. The tokens are valid, but the positional pattern can move far outside the range used during training. Position interpolation changes that geometry. Instead of sending larger position indices directly into rotary position embeddings, it compresses a longer sequence into the positional range the model already uses. With suitable adaptation, this can extend the usable context window without changing the transformer architecture.