Skip to content

Archive

Context Window

3 articles
Artificial Intelligence 24 Sep 2026 5 min read

RoPE Position Interpolation Compresses Indices Before Rotation

Rotary position embeddings apply a position-dependent rotation to query and key components before their dot product is evaluated. If a model was trained with positions up to a reference length, simply sending much larger position indices into the same rotation rule can place attention in angular regimes that were not represented in that training range. Position interpolation changes the coordinate supplied to RoPE. For an extension factor s > 1, a simplified form maps

Artificial Intelligence 23 Sep 2026 4 min read

RoPE Position Interpolation Compresses Relative Angles Across Longer Contexts

Rotary position embeddings encode position by rotating query and key components with angles derived from token indices. When a model is asked to operate beyond the position range used during its original training, those angles can enter a regime the model did not encounter. Position interpolation addresses that boundary by mapping a longer sequence into a smaller position-coordinate range before the rotary angles are computed. The mechanism does not add extra tokens to the model’s architectural state. It changes the coordinates supplied to the positional transform. That distinction matters because a longer accepted input length and reliable behavior at that length are separate properties.

Artificial Intelligence 12 Sep 2026 9 min read

Extend RoPE Context Windows with Position Interpolation

Extend RoPE Context Windows with Position Interpolation A RoPE-based language model trained on sequences up to a fixed length can behave poorly when inference suddenly asks it to process much larger position indices. The tokens are valid, but the positional pattern can move far outside the range used during training. Position interpolation changes that geometry. Instead of sending larger position indices directly into rotary position embeddings, it compresses a longer sequence into the positional range the model already uses. With suitable adaptation, this can extend the usable context window without changing the transformer architecture.