Skip to content

Archive

Attention

44 articles
Artificial Intelligence 16 Sep 2026 5 min read

Prune Transformer Attention Heads with Measured Impact

A multi-head attention block can contain heads whose removal changes a target metric very little on a chosen evaluation set. That observation makes attention head pruning attractive: identify low-impact heads, remove their contribution, and retain the heads that matter more for the target workload. The difficult part is not setting a head output to zero. It is deciding what that intervention measures and whether the resulting model actually executes less work. A masked head, a structurally removed head, and a faster attention kernel are related ideas, but they are not the same result.

Artificial Intelligence 16 Sep 2026 5 min read

Preserve Attention Sinks in Sliding KV Caches

A sliding key-value cache seems to offer a simple bound on transformer inference memory: retain the most recent tokens and evict the oldest entries as generation continues. That policy preserves local context, but it can change attention behavior more sharply than token age alone suggests. Some early positions can attract substantial attention even when their lexical content is not directly relevant to the current token. Removing those positions can disturb the distribution that later layers receive.

Artificial Intelligence 16 Sep 2026 6 min read

Isolate Packed Sequences During Transformer Training

Padding can consume a large share of a training batch when sequence lengths vary. Sequence packing replaces some of that padding with tokens from additional examples, placing multiple independent samples inside one fixed-length token block. The arithmetic is attractive: more of each block carries data that contributes to the training objective. The packed tensor, however, no longer describes one continuous sequence. If the model treats it that way, tokens from a later example can attend to tokens from an earlier one. The optimizer then sees dependencies that were absent from the original dataset. Packing is therefore not only a batching optimization. It changes the structure presented to the attention mechanism unless example boundaries are represented explicitly.

Artificial Intelligence 15 Sep 2026 4 min read

Preserve Attention Sinks in Streaming Transformers

A bounded attention cache creates a specific failure mode in autoregressive transformers: removing every old token can disturb attention even when those tokens no longer carry useful task content. Some early positions can absorb attention mass across many later queries. If cache eviction removes them, generation quality can degrade more than their semantic value would suggest. These positions are often called attention sinks. The practical implication is narrow but useful: a streaming cache can keep a small prefix of sink positions while rotating the rest of its capacity through recent tokens.

Artificial Intelligence 15 Sep 2026 5 min read

Isolate Attention Across Packed Training Sequences

Padding can consume a large share of a training batch when sequence lengths vary. Sequence packing reduces that waste by placing multiple short samples into one token block. The arithmetic is attractive: more non-padding tokens fit into the same fixed-length tensor. Packing also changes the structure seen by attention. A standard causal mask only prevents a token from attending to later positions. It does not know that two adjacent spans came from separate samples. Without an additional boundary constraint, a token in the second span can attend to tokens from the first span.

Artificial Intelligence 14 Sep 2026 6 min read

Reduce KV Cache Memory with Grouped-Query Attention

Autoregressive transformer inference stores key and value vectors from earlier tokens so each new token does not have to recompute them. With standard multi-head attention, every attention head has its own key and value projections, so the KV cache grows with the number of key-value heads. Grouped-query attention changes that head layout. It keeps multiple query heads but lets several query heads share one key head and one value head. The result reduces cached key-value state without collapsing all query heads into a single shared projection.

Artificial Intelligence 14 Sep 2026 6 min read

Encode Token Distance with Rotary Position Embeddings

Transformer attention has no intrinsic notion that one token sits three positions before another. Rotary position embeddings, usually called RoPE, inject position into attention by rotating pairs of query and key coordinates before their dot product is computed. The mechanism is easy to reduce to a helper function, yet several details determine its actual behavior: queries and keys must use compatible rotations, each coordinate pair has its own angular frequency, offsets emerge through the dot product, and changing the position scale changes the geometry seen by attention.

Artificial Intelligence 13 Sep 2026 6 min read

Preserve Attention Sinks in Bounded KV Caches

A decoder that keeps only the newest key-value states can degrade even when its cache still contains enough recent text for the immediate task. In some transformer models, early token positions attract substantial attention across later decoding steps. Evicting those states changes the attention distribution, not just the amount of accessible history. Attention sink retention addresses that specific failure mode. A bounded cache preserves a small prefix of initial key-value states together with a moving window of recent states. Tokens between those regions can be discarded, keeping cache size bounded as generation continues.

Artificial Intelligence 12 Sep 2026 9 min read

Stream Long LLM Sessions with Attention Sinks

Stream Long LLM Sessions with Attention Sinks Long-running LLM sessions create a simple resource problem: every generated token can add keys and values to the attention cache. Keep the entire history and memory use keeps growing. Keep only the newest tokens and some transformer models degrade sharply once older cache entries disappear. Attention sinks provide a useful middle ground for compatible models. Instead of retaining the full KV cache, keep a small group of initial tokens plus a moving window of recent tokens. The cache stays bounded, yet the model can remain much more stable than with a recent-token window alone.

Artificial Intelligence 12 Sep 2026 6 min read

Pack Transformer Training Sequences Without Cross-Sample Attention

Transformer batches often waste token slots on padding when examples have uneven lengths. Sequence packing reduces that waste by placing several shorter samples into one fixed-length token block. The arithmetic is attractive, but concatenation alone changes the training problem: tokens from one sample can attend to tokens from another unless the packed representation preserves sample boundaries. A correct packing scheme therefore has two jobs. It must fill token capacity more densely, and it must keep the model’s effective computation consistent with the intended independence of the original samples.

Artificial Intelligence 12 Sep 2026 7 min read

Extend RoPE Context with Position Interpolation

A transformer that uses rotary position embeddings can accept a larger token buffer at the serving layer and still behave poorly at positions far beyond the range used during model training. The tensor shapes may be valid while the positional phases presented to attention are outside the regime the model adapted to. Position interpolation addresses that mismatch by compressing a longer sequence’s position indices into the original position interval before applying RoPE. It does not add memory to the architecture, and it does not make long-context behavior equivalent to native training at the extended length. It changes the positional coordinates supplied to attention.

Artificial Intelligence 12 Sep 2026 7 min read

Bound KV Cache Growth in Streaming LLM Inference

Autoregressive transformer inference normally keeps key and value states from earlier tokens so each new token can attend to prior context without recomputing those states. The cache grows with sequence length. For a long-running stream, that growth eventually becomes a memory constraint even when generation itself continues one token at a time. KV cache eviction puts a bound on that state by discarding selected cached positions. The memory effect is straightforward: fewer retained key-value pairs occupy less cache space. The model effect is more subtle. Once a position is removed, later attention layers cannot use its cached key and value in the ordinary attention calculation. An eviction policy therefore changes both resource use and the effective attention history.

Artificial Intelligence 08 Sep 2026 10 min read

Transformer Attention Weights: What They Show and What They Do Not

Transformer attention maps are visually compelling. A token appears to assign most of its attention to another token, so it is tempting to conclude that the second token caused the model’s prediction. That conclusion is stronger than the data supports. An attention weight has a precise local meaning: inside one attention operation, it controls how strongly a query mixes information from available value vectors. A complete transformer prediction, however, also depends on value vectors, residual connections, feed-forward layers, normalization, later layers, and often many attention heads. A large weight is therefore evidence about one routing operation, not a complete causal explanation.

Artificial Intelligence 08 Sep 2026 9 min read

Trace Token Influence with Attention Rollout

Looking at one Transformer attention matrix can answer a local question: which positions a token attends to in that layer. It does not directly tell you how much an input token can influence a representation several layers later. The reason is mixing. After one layer, a token representation already contains information gathered from other positions. The next layer attends to those mixed representations, not to untouched input tokens. Residual connections add another path that carries each representation forward. Reading only the final layer therefore skips the paths through earlier layers.

Artificial Intelligence 06 Sep 2026 10 min read

Rotary Position Embeddings in Transformers

A Transformer attention layer needs to know more than which tokens are present. Order matters: dog bites man and man bites dog contain the same words but express different relationships. Yet the dot products used by self-attention do not inherently know whether two token representations came from adjacent positions or opposite ends of a sequence. Rotary position embedding, usually shortened to RoPE, adds position information by rotating parts of the query and key vectors before their attention scores are computed. The useful consequence is subtle: each token receives a transformation based on its absolute position, while the dot product between two transformed vectors depends on their relative position.

Artificial Intelligence 05 Sep 2026 10 min read

Measure Attention Concentration with Entropy

Transformer attention is often inspected as a matrix of weights. That works for a few examples, but it becomes difficult when you need to compare many heads, layers, tokens, or model runs. A useful summary is attention entropy: a number that describes how concentrated or spread out one attention distribution is. Entropy can answer a narrow but practical question: does this query place most of its attention mass on a few available positions, or distribute that mass broadly? It does not tell you whether the model is correct, whether a token caused the prediction, or whether a head is important. Used with those limits in mind, it is a compact diagnostic for attention behavior.

Artificial Intelligence 05 Sep 2026 9 min read

Mask Padding Tokens in Transformer Attention

Transformer batches often contain sequences with different lengths. To store them in one rectangular tensor, shorter sequences are usually extended with padding tokens. Padding solves a shape problem, but it creates a modeling problem: those extra positions are not part of the original input. If attention treats padding like ordinary content, real tokens can assign probability to positions that carry no useful information. The result may be wasted attention, representations that depend on how much padding was added, and training behavior that differs unnecessarily across batches.

Artificial Intelligence 05 Sep 2026 10 min read

Attention Sinks for Stable Streaming LLM Inference

Autoregressive language models normally reuse the keys and values of earlier tokens while generating the next token. This KV cache avoids recomputing the entire prefix at every decoding step, but its memory use grows with the cached sequence. A long-running chat, agent, or stream can therefore accumulate more cached state than a serving system wants to keep. A tempting fix is a sliding window: retain only the most recent tokens and evict everything older. For models trained with ordinary dense attention, however, abruptly dropping all early tokens can damage generation quality even when those old tokens do not appear semantically important.

Artificial Intelligence 04 Sep 2026 9 min read

Use Padding Masks for Variable-Length Transformer Batches

Transformer inputs rarely have identical lengths. One sentence may contain 8 tokens while another contains 30, yet efficient training and inference usually process multiple sequences in rectangular tensors. The usual solution is to add padding tokens to shorter sequences until their shapes match. Padding solves the shape problem but creates a semantic one: the added positions are not real input. If the model treats them like ordinary tokens, they can influence attention, pooling, and training loss. A padding mask tells the computation which positions are valid and which exist only to make the batch rectangular.

Artificial Intelligence 03 Sep 2026 6 min read

Self-Attention in Transformer Models

Transformers can process relationships between tokens without stepping through a sequence one token at a time. The mechanism that makes this possible is self-attention: each token builds a weighted view of other tokens in the same context. The formula is compact. The sections below show what the calculation does, how masking sets the information boundary, and where the computational cost comes from. Start with token representations Before attention runs, each input token is represented by a vector. Let the matrix X contain those token representations. A transformer layer applies learned projections to produce three matrices: