Skip to content

Archive

Attention

44 articles
Artificial Intelligence 24 Sep 2026 5 min read

Sliding-Window Attention Bounds Each Query to a Local Token Horizon

A causal attention layer does not always expose every earlier token to every query. With a sliding window of width w, the query at position i can be restricted to recent positions rather than the full prefix. The attention graph becomes local: old tokens fall outside the direct edge set even though they remain part of the sequence. That boundary changes computation, memory traffic, and information paths at the same time. It is not merely an optimized implementation of full attention. Once the mask removes distant key-value pairs, the layer implements a different dependency pattern.

Artificial Intelligence 24 Sep 2026 5 min read

Sliding-Window Attention Bounds Direct Token Access

A causal transformer does not have to expose every earlier token to every query position. With sliding-window attention, position i can attend only to a bounded range of preceding keys. Tokens outside that range are absent from that layer’s attention operation, even if they still belong to the model input. That boundary changes more than the attention matrix shape. It separates direct access from information that can reach a position only after being propagated through intermediate hidden states.

Artificial Intelligence 24 Sep 2026 5 min read

Rotary Position Embedding Converts Absolute Indices into Relative Attention Phases

Self-attention can compare token content without assigning an order to token positions. Rotary Position Embedding (RoPE) inserts position into that comparison by rotating paired coordinates of queries and keys. Each token receives an absolute rotation angle, yet the query-key inner product reduces those two absolute angles to their difference. That algebraic cancellation is the central mechanism. RoPE does not add a position vector to the hidden state. It changes the orientation of query and key components before their dot product is evaluated.

Artificial Intelligence 24 Sep 2026 4 min read

RoPE Encodes Relative Offsets Through Rotated Query-Key Phases

Rotary Position Embedding (RoPE) applies position-dependent rotations to query and key coordinates before their attention dot product. The resulting score carries relative position through the phase difference between those rotations rather than through an additive position vector attached to the token representation. This distinction is structural. RoPE starts from absolute indices for each rotation, yet the query-key inner product can be written in terms of the offset between their positions.

Artificial Intelligence 24 Sep 2026 4 min read

QK Normalization Bounds Attention Logit Scale Before Softmax

QK normalization inserts normalization on query and key vectors before the attention dot product. The operation changes the geometry of the score calculation: vector magnitude no longer enters the dot product in the same unrestricted form, while directional alignment remains part of the score. For one query vector q and key vector k, ordinary scaled dot-product attention forms a score such as: s = dot(q, k) / sqrt(D) A QK-normalized variant first applies the model’s specified normalization functions:

Artificial Intelligence 24 Sep 2026 6 min read

Packed Sequences Need Boundary-Aware Attention Masks

Sequence packing reduces padding by placing several variable-length examples into one token buffer. The storage layout may look like one long sequence, but the examples are still semantically independent. A standard causal mask does not preserve that independence by itself. For a decoder-only Transformer, causal masking blocks attention to future positions. It does not normally block attention to earlier positions that belong to another packed example. If boundaries are ignored, tokens in a later example can attend to keys and values from an earlier one. The model then receives context that the data pipeline intended to keep separate.

Artificial Intelligence 24 Sep 2026 5 min read

Multi-Head Latent Attention Compresses KV State Before Head-Specific Expansion

Multi-Head Latent Attention changes the state retained across autoregressive decoding. Instead of requiring a full key tensor and value tensor for every cached token and attention head, the architecture can retain a lower-dimensional latent representation and derive head-specific key/value information through projection structure. That distinction matters at the cache boundary. Ordinary multi-head attention commonly stores already projected per-head keys and values. MLA moves part of that representation behind a compact bottleneck, so persistent state and compute no longer have the same shape.

Artificial Intelligence 24 Sep 2026 4 min read

KV-Cache Quantization Reduces Cache Bytes with Separate Accuracy and Kernel Costs

Autoregressive decoding retains key and value tensors from earlier tokens, so KV-cache memory grows with retained context. Quantizing those tensors changes a direct term in that memory footprint: fewer bits are stored for each cached element. The trade is not free capacity. Reduced precision adds representation error and requires a concrete scaling, storage, and kernel strategy. Cache precision is separate from weight precision Model weights and KV state have different lifetimes. Weights are persistent across requests, while KV tensors are generated from each request and grow as its sequence advances. A model can therefore use one numerical format for weights and another for its cache.

Artificial Intelligence 24 Sep 2026 4 min read

KV Cache Quantization Perturbs Attention Through Stored Keys and Values

During autoregressive inference, previously computed keys and values are reused from the KV cache instead of being recomputed for every new token. Quantizing that cache changes more than its byte representation. The stored approximation becomes an input to later attention operations, so its error can alter both attention scores and the vectors combined by those scores. This boundary differs from quantizing model weights. A weight tensor is reused across requests, while KV state is generated from the current sequence and grows with its cached length. Its numeric range can also vary across layers, heads, positions, and requests.

Artificial Intelligence 24 Sep 2026 6 min read

Grouped-Query Attention Shares Key-Value Heads Across Query Groups

Multi-head attention does not require every query head to own a distinct key head and value head. Grouped-query attention, usually abbreviated GQA, partitions query heads into groups and assigns one key-value head to each group. Query projections remain separate, but several query heads read from the same projected keys and values. That distinction changes parameter shapes and cached state without collapsing the query heads into one attention computation. Each query head still produces its own attention scores because its query vector is different.

Artificial Intelligence 24 Sep 2026 5 min read

FlashAttention Tiles Exact Softmax Without Materializing the Score Matrix

Standard scaled dot-product attention forms a score matrix whose two long axes are sequence positions. For one attention head, S = QK^T / sqrt(d) P = softmax(S) O = PV the mathematical definition is compact, but a direct GPU implementation can write the large intermediate matrices S and P to high-bandwidth memory before reading them again. FlashAttention changes that dataflow. It processes blocks of queries, keys, and values in on-chip memory and carries enough row-wise softmax state to combine score tiles exactly.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Stabilize Windowed KV-Cache Streaming

A decoder can keep its KV cache bounded by discarding states that fall outside a fixed recent-token window. The memory bound is attractive, but naive eviction can sharply degrade generation even when the removed tokens carry little obvious semantic value. A small set of key-value states from the beginning of the sequence can change that behavior. This effect is associated with attention sinks: initial positions that receive substantial attention mass even when their token content is not important to the current prediction. StreamingLLM reported that preserving those initial states together with a rolling window of recent states recovers stable language-model behavior that ordinary window eviction can lose.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Stabilize Streaming Cache Eviction

A bounded KV cache can keep inference memory from growing with an unending token stream, but deleting every old token in strict arrival order can disrupt attention more than the missing content alone suggests. In some autoregressive transformers, early positions receive substantial attention even when their token semantics are not directly relevant to the current prediction. These positions act as attention sinks. The serving consequence is specific: a sliding cache that preserves a small initial prefix alongside the most recent tokens can behave differently from a same-sized cache containing only the newest tokens. The mechanism concerns attention state and positional handling, not retrieval of forgotten text.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Absorb Probability Mass Without Carrying Matching Content

A causal attention head can assign noticeable probability mass to an early token even when that token is not a strong semantic match for the current query. The token then behaves as an attention sink: its key remains a convenient destination for probability that the head does not direct toward content-bearing positions. This behavior matters during autoregressive serving because a KV-cache policy can preserve recent tokens yet still alter model behavior sharply if it removes sink positions. Recency alone does not describe the functional role of every cached key.

Artificial Intelligence 24 Sep 2026 5 min read

ALiBi Adds Distance-Proportional Bias to Attention Scores

ALiBi changes an attention score before softmax rather than adding a positional vector to the token representation. For a causal transformer, a key farther behind the current query receives a larger negative offset. The offset is linear in token distance and uses a slope associated with the attention head. A simplified score for head h can be written as: score_h(i, j) = q_i k_j^T / sqrt(d) - m_h * (i - j) for an allowed causal pair with j <= i and positive slope m_h. The causal mask still decides which future positions are inaccessible. ALiBi changes the relative scores among positions that remain eligible.

Artificial Intelligence 24 Sep 2026 5 min read

ALiBi Adds Distance Penalties Directly to Attention Logits

Attention with Linear Biases (ALiBi) changes a causal attention score before softmax by adding a penalty whose magnitude grows with token distance. Position is therefore represented in the score path rather than by adding a positional vector to each token representation. For a query at position i attending to a key at position j, one head can be written schematically as: score(i, j) = q_i · k_j / sqrt(d_k) - m_h * (i - j) for causal positions j <= i. The positive slope m_h is specific to attention head h. The causal mask still prevents access to future positions; the linear term changes the relative preference among positions that remain visible.

Artificial Intelligence 23 Sep 2026 6 min read

Sliding-Window Attention Bounds Active KV State by Token Distance

Sliding-window attention imposes a finite token-distance boundary on causal attention. At position (t), a query can address only a recent interval of key and value positions rather than the full prefix. Once a cached position falls permanently outside that interval, later queries governed by the same local rule cannot address it. That boundary changes the state required for autoregressive decoding. Full causal attention keeps usable KV state growing with sequence length. A fixed local window can keep the active KV span bounded, provided the runtime evicts or overwrites entries that have become unreachable.

Artificial Intelligence 23 Sep 2026 5 min read

RoPE Position Shifts Preserve Relative Phase but Change Absolute Rotation

Rotary position embeddings, commonly called RoPE, inject position into attention by rotating pairs of query and key channels with angles determined by token position. The rotation happens before the query-key dot product. As a result, position is not represented by adding a standalone vector to each token representation. This distinction matters during inference. A token’s cached key already contains the rotation associated with its assigned position. Reusing that key under a different sequence coordinate without an equivalent transformation changes the attention computation, even when the token content is identical.

Artificial Intelligence 23 Sep 2026 5 min read

Grouped-Query Attention Shares KV Heads Across Query Groups

Grouped-query attention (GQA) changes a specific structural ratio inside an attention layer: the number of query heads can exceed the number of key and value heads. Several query heads then consume the same projected key and value head. The attention calculation remains head-specific on the query side, while KV state is shared within each group. That asymmetry matters during autoregressive decoding because cached keys and values persist for prior tokens. Reducing the count of distinct KV heads reduces the amount of per-token KV state that must remain available to later decoding steps.

Artificial Intelligence 23 Sep 2026 4 min read

Grouped-Query Attention Shares Key-Value Heads Across Query Heads

Grouped-query attention changes a specific part of multi-head attention: several query heads use the same key and value head. The query projections remain separate, so those query heads can produce different attention weights, but they read keys and values from a shared projected representation. That distinction matters during inference. A decoder cache stores past keys and values, not past queries. Reducing the number of key-value heads can therefore reduce KV cache storage without reducing the number of query heads by the same factor.

Artificial Intelligence 23 Sep 2026 4 min read

Attention Sinks Make Naive KV Eviction Unstable

A fixed-size KV cache seems to invite a simple policy: keep the newest entries and evict the oldest. For some decoder transformers, that policy removes positions that later queries continue to assign substantial attention mass. Those early positions act as attention sinks, and dropping them can change the attention distribution far more than their semantic content suggests. The effect matters specifically at inference time when a serving system truncates cached keys and values. It is not a general claim that the first token is semantically privileged, nor that every transformer exhibits the same pattern.

Artificial Intelligence 23 Sep 2026 6 min read

Attention Logit Scaling Keeps Dot Products in a Stable Softmax Range

A dot product between a query and a key tends to grow in magnitude as their dimension grows. In scaled dot-product attention, the score is divided by the square root of the key dimension before softmax. That factor is not a cosmetic normalization. It controls the scale presented to softmax under a specific statistical assumption about the query and key components. The familiar expression is Attention(Q, K, V) = softmax(QK^T / sqrt(d_k)) V where d_k is the query-key dimension for one attention head. The scaling term affects the distribution of attention probabilities even though it does not change the ordering of logits by itself.

Artificial Intelligence 23 Sep 2026 5 min read

Attention Entropy Measures Concentration, Not Causal Importance

An attention head can place most of its probability mass on one position and still provide little evidence that this position controls the final model output. The entropy of its attention weights captures concentration, not causal influence. That distinction matters when attention maps are inspected as diagnostics. Entropy can reveal whether a head spreads mass broadly or focuses it narrowly for a given query. It cannot, by itself, establish that the highest-weight token carries the feature responsible for a downstream prediction.

Artificial Intelligence 22 Sep 2026 6 min read

Padding Masks Do Not Remove Padding Compute in Dense Attention

A batch can contain two prompts with very different token counts yet represent both with the same rectangular tensor. The shorter prompt is extended with padding so its tensor shape matches the longest sequence in the batch. An attention mask can stop those padded positions from contributing to attention probabilities, but that semantic exclusion does not imply that dense kernels skip every operation associated with the padded rows and columns.