Skip to content

Archive

Transformers

97 articles
Artificial Intelligence 16 Sep 2026 5 min read

Prune Transformer Attention Heads with Measured Impact

A multi-head attention block can contain heads whose removal changes a target metric very little on a chosen evaluation set. That observation makes attention head pruning attractive: identify low-impact heads, remove their contribution, and retain the heads that matter more for the target workload. The difficult part is not setting a head output to zero. It is deciding what that intervention measures and whether the resulting model actually executes less work. A masked head, a structurally removed head, and a faster attention kernel are related ideas, but they are not the same result.

Artificial Intelligence 16 Sep 2026 6 min read

Preserve Attention Sinks in Streaming KV Caches

A bounded KV cache seems to invite a simple eviction rule: keep the newest tokens and discard the oldest ones. For some transformer language models, that rule can degrade generation even when the discarded prefix carries little obvious semantic value. A small set of early positions may attract substantial attention across later decoding steps. These positions are commonly called attention sinks. This behavior matters for streaming inference because cache eviction changes the attention computation itself. A fixed-size cache that preserves a few sink positions plus a recent window can behave differently from a cache containing only the same number of recent positions.

Artificial Intelligence 16 Sep 2026 5 min read

Preserve Attention Sinks in Sliding KV Caches

A sliding key-value cache seems to offer a simple bound on transformer inference memory: retain the most recent tokens and evict the oldest entries as generation continues. That policy preserves local context, but it can change attention behavior more sharply than token age alone suggests. Some early positions can attract substantial attention even when their lexical content is not directly relevant to the current token. Removing those positions can disturb the distribution that later layers receive.

Artificial Intelligence 16 Sep 2026 6 min read

Isolate Packed Sequences During Transformer Training

Padding can consume a large share of a training batch when sequence lengths vary. Sequence packing replaces some of that padding with tokens from additional examples, placing multiple independent samples inside one fixed-length token block. The arithmetic is attractive: more of each block carries data that contributes to the training objective. The packed tensor, however, no longer describes one continuous sequence. If the model treats it that way, tokens from a later example can attend to tokens from an earlier one. The optimizer then sees dependencies that were absent from the original dataset. Packing is therefore not only a batching optimization. It changes the structure presented to the attention mechanism unless example boundaries are represented explicitly.

Artificial Intelligence 16 Sep 2026 6 min read

Extend RoPE Context with Positional Interpolation

A transformer using rotary position embeddings can accept tensors longer than the sequence length used during training, yet accepting the shape does not establish that its position signal remains usable at those distances. Rotary angles at unseen positions can place attention computations outside the positional regime the model encountered during optimization. Positional interpolation changes the input to rotary position embeddings rather than merely raising a sequence-length limit. For a target context longer than the original training context, position indices are compressed so the extended sequence maps into the earlier positional range. The model then needs to adapt to denser positional spacing instead of extrapolating directly to larger indices.

Artificial Intelligence 16 Sep 2026 6 min read

Bucket Sequence Lengths to Reduce Padding Waste

A padded batch is shaped by its longest sequence, not its average sequence. If one batch contains token counts of 120, 124, 131, and 900, every sequence may be represented at length 900. Most positions in the first three rows then carry padding rather than input tokens. Length bucketing changes batch composition instead of changing the model. Examples with similar token counts are placed near each other before batches are formed. The maximum length inside each batch falls closer to the lengths of its members, reducing the number of padded positions processed by operations that still use the rectangular batch shape.

Artificial Intelligence 15 Sep 2026 6 min read

Steer Transformer Activations with Residual Stream Vectors

A transformer can produce different output behavior even when its weights and input tokens stay fixed. One way to cause that change is to alter an intermediate hidden state during the forward pass. Activation steering does this deliberately by adding a vector to a selected residual-stream position or set of positions. The mechanism is simple enough to express as an intervention, but its effect is not a global model setting. The chosen direction, coefficient, layer, token positions, and decoding setup all affect the result. Treating those choices as part of the inference configuration makes the behavior easier to reason about and test.

Artificial Intelligence 15 Sep 2026 6 min read

Quantize KV Caches with Explicit Error Budgets

Autoregressive transformer inference retains key and value tensors from earlier tokens so each new token can attend to prior context without recomputing those projections. As context length and concurrent sequence count rise, this KV cache can become a substantial part of accelerator memory. KV cache quantization stores those tensors at reduced precision and reconstructs approximations when attention consumes them. The memory arithmetic is attractive, but the resulting error is not a generic model-weight perturbation. Quantized keys affect attention scores before the softmax, while quantized values affect the weighted sum after attention probabilities have been formed.

Artificial Intelligence 15 Sep 2026 4 min read

Preserve Attention Sinks in Streaming Transformers

A bounded attention cache creates a specific failure mode in autoregressive transformers: removing every old token can disturb attention even when those tokens no longer carry useful task content. Some early positions can absorb attention mass across many later queries. If cache eviction removes them, generation quality can degrade more than their semantic value would suggest. These positions are often called attention sinks. The practical implication is narrow but useful: a streaming cache can keep a small prefix of sink positions while rotating the rest of its capacity through recent tokens.

Artificial Intelligence 15 Sep 2026 5 min read

Isolate Attention Across Packed Training Sequences

Padding can consume a large share of a training batch when sequence lengths vary. Sequence packing reduces that waste by placing multiple short samples into one token block. The arithmetic is attractive: more non-padding tokens fit into the same fixed-length tensor. Packing also changes the structure seen by attention. A standard causal mask only prevents a token from attending to later positions. It does not know that two adjacent spans came from separate samples. Without an additional boundary constraint, a token in the second span can attend to tokens from the first span.

Artificial Intelligence 14 Sep 2026 5 min read

Prevent Cross-Example Attention in Packed Sequences

Short training examples can waste much of a fixed-length transformer batch on padding. Sequence packing reduces that waste by placing several examples into one token buffer, but concatenation alone changes the computation. A causal mask prevents a token from attending to future positions; it does not prevent that token from attending to an earlier, unrelated example. The distinction matters whenever packed examples are intended to remain independent. The token buffer may be contiguous for storage and compute while attention, position handling, and loss accounting still need explicit example boundaries.

Artificial Intelligence 14 Sep 2026 5 min read

Normalize Hidden States with RMSNorm

A hidden-state vector can grow or shrink in magnitude as it passes through a neural network. RMSNorm controls that scale by dividing the vector by its root mean square magnitude, then applying a trainable gain. Unlike LayerNorm, it does not subtract the vector mean before rescaling. That missing centering operation is the defining distinction. RMSNorm constrains scale while leaving a uniform shift across coordinates present in the normalized representation. RMSNorm uses the second raw moment For a hidden vector x with d coordinates, its root mean square is:

Artificial Intelligence 14 Sep 2026 6 min read

Encode Token Distance with Rotary Position Embeddings

Transformer attention has no intrinsic notion that one token sits three positions before another. Rotary position embeddings, usually called RoPE, inject position into attention by rotating pairs of query and key coordinates before their dot product is computed. The mechanism is easy to reduce to a helper function, yet several details determine its actual behavior: queries and keys must use compatible rotations, each coordinate pair has its own angular frequency, offsets emerge through the dot product, and changing the position scale changes the geometry seen by attention.

Artificial Intelligence 14 Sep 2026 6 min read

Compare RMSNorm and Layer Normalization

Normalization layers can look interchangeable when their outputs have similar shapes, but their invariances are not the same. RMSNorm rescales an activation vector using its root mean square without first subtracting the vector mean. Layer normalization centers the vector and then rescales it using its variance. That missing centering operation is the central distinction. It changes which transformations of an activation vector disappear under normalization and which remain visible to the rest of the network.

Artificial Intelligence 14 Sep 2026 6 min read

Bound Extreme Logits with Soft Capping

A transformer can produce logits whose magnitudes grow far beyond the range needed to express a strong preference. Large attention scores can make a softmax distribution extremely concentrated, while large output logits can make token probabilities nearly one-hot. A hard clamp can bound those values, but it introduces a flat region with an abrupt derivative change at the threshold. Logit soft capping uses a smooth saturating function instead. One common form is:

Artificial Intelligence 13 Sep 2026 8 min read

Preserve Attention Sinks in Sliding-Window LLM Inference

A sliding-window KV cache seems mechanically simple: keep the most recent tokens, evict older key-value entries, and continue decoding within a fixed memory budget. The complication is that some transformer models place substantial attention mass on a small set of early positions even when those positions carry little direct semantic relevance to the current token. Those positions are often called attention sinks. If a cache policy removes them while preserving only the newest tokens, the attention distribution seen during decoding can change abruptly. A bounded cache can therefore behave differently from full-context inference even when the evicted text appears unrelated to the current request.

Artificial Intelligence 13 Sep 2026 6 min read

Inspect Intermediate Transformer Predictions with Logit Lens

A transformer produces its next-token distribution only after the final block, yet every block updates the residual state that eventually feeds that prediction. Logit lens examines those intermediate states by mapping them through the model’s output path into vocabulary logits. The result is a sequence of provisional token distributions across model depth. The method is attractive because it reuses components already present in the model. Its output also needs careful interpretation. An intermediate residual state was not necessarily optimized to behave like a final residual state, so a readable token ranking is a diagnostic projection rather than a direct transcript of internal computation.

Artificial Intelligence 13 Sep 2026 6 min read

Compare RMSNorm and LayerNorm in Transformers

LayerNorm and RMSNorm can occupy the same structural position in a transformer while applying different operations to the residual stream. LayerNorm subtracts the feature mean before scaling by a measure of spread. RMSNorm skips the centering operation and scales directly from the root mean square of the features. That small algebraic difference changes which transformations of an activation vector are removed by normalization. It also means that replacing one operation with the other is not, in general, a function-preserving edit to an existing model.

Artificial Intelligence 13 Sep 2026 8 min read

Balance Sparse Mixture-of-Experts Routing Under Capacity Limits

A sparse mixture-of-experts layer does not send every token through every parameter block. A router scores the available experts, selects a small subset for each token, and dispatches token representations only to those selected experts. That conditional computation is the main attraction of sparse MoE designs, but it also creates a resource-allocation problem inside the model. The router can prefer the same experts for many tokens. Hardware, meanwhile, has finite buffers and communication capacity. A routing policy that looks reasonable from token scores alone can therefore create overloaded experts, idle experts, uneven communication, or discarded assignments.

Artificial Intelligence 12 Sep 2026 11 min read

Understand Sparse Mixture-of-Experts Routing

Understand Sparse Mixture-of-Experts Routing A larger neural network can represent more functions, but using every parameter for every input makes each forward pass expensive. Sparse mixture-of-experts (MoE) models take a different approach: keep many parameter groups available, then activate only a small subset for each token. That sounds like a simple efficiency trick, but routing changes much more than arithmetic cost. It affects training stability, accelerator communication, memory requirements, batching, and the meaning of a model’s total parameter count.

Artificial Intelligence 12 Sep 2026 8 min read

Test Transformer Circuits with Activation Patching

A transformer can expose a clear internal pattern without that pattern being responsible for the output under inspection. Activation patching addresses this gap by changing an internal state and measuring the downstream effect. Instead of asking whether a feature is visible at a layer, it asks whether replacing a selected state changes a defined model behavior. The method is simple in form but sensitive to experimental design. A patch has meaning only relative to the paired inputs, the patched location, the replacement value, and the output metric. Changing any of those can change the causal question being tested.

Artificial Intelligence 12 Sep 2026 6 min read

Pack Transformer Training Sequences Without Cross-Sample Attention

Transformer batches often waste token slots on padding when examples have uneven lengths. Sequence packing reduces that waste by placing several shorter samples into one fixed-length token block. The arithmetic is attractive, but concatenation alone changes the training problem: tokens from one sample can attend to tokens from another unless the packed representation preserves sample boundaries. A correct packing scheme therefore has two jobs. It must fill token capacity more densely, and it must keep the model’s effective computation consistent with the intended independence of the original samples.

Artificial Intelligence 12 Sep 2026 7 min read

Inspect Transformer Layer Predictions with the Logit Lens

A decoder-only transformer normally exposes token logits only after its final block and output normalization. The residual stream inside earlier blocks has the same model-width shape, which makes another operation possible: take an intermediate state, apply the model’s output-side normalization when required, and project that state through the output matrix. The resulting vocabulary scores form the logit lens. They provide a token-space view of an internal representation before the remaining transformer blocks have processed it. That view is useful for inspecting how candidate tokens change across depth, but it is not a record of tokens that the model has secretly selected in advance.

Artificial Intelligence 12 Sep 2026 7 min read

Extend RoPE Context with Position Interpolation

A transformer that uses rotary position embeddings can accept a larger token buffer at the serving layer and still behave poorly at positions far beyond the range used during model training. The tensor shapes may be valid while the positional phases presented to attention are outside the regime the model adapted to. Position interpolation addresses that mismatch by compressing a longer sequence’s position indices into the original position interval before applying RoPE. It does not add memory to the architecture, and it does not make long-context behavior equivalent to native training at the extended length. It changes the positional coordinates supplied to attention.