Skip to content

Archive

Transformers

97 articles
Artificial Intelligence 24 Sep 2026 6 min read

Tanh Logit Soft Capping Bounds Extreme Scores Before Softmax

A softmax can accept logits of any finite magnitude, but large score gaps make its output increasingly concentrated. Tanh logit soft capping inserts a bounded nonlinear transform before softmax so that no transformed logit exceeds a configured magnitude. For a positive cap c, a common form is: softcap(z; c) = c * tanh(z / c) The operation does not clip at a hard threshold. It behaves almost linearly near zero and gradually compresses larger magnitudes as they approach -c or c.

Artificial Intelligence 24 Sep 2026 5 min read

SwiGLU Gates Transformer Feed-Forward Channels with a Second Projection

SwiGLU splits a transformer feed-forward input into two projected paths, applies SiLU to one path, then multiplies the two results element by element. The second projection is not an auxiliary statistic: its values directly gate the activated path before the output projection. For hidden state x, a common structural form is: g = SiLU(x W_gate) u = x W_up h = g * u y = h W_down Bias terms, projection orientation, intermediate width, and parameter names vary across architectures. The defining boundary is the element-wise product between a nonlinear projected branch and another projected branch.

Artificial Intelligence 24 Sep 2026 5 min read

Speculative Decoding Verifies Draft Tokens in Parallel

Autoregressive generation normally commits one token before the next target-model step can be evaluated. Speculative decoding changes that execution pattern. A cheaper proposal distribution produces a short candidate block, while the target model evaluates the proposed positions together. An acceptance rule then determines how much of that block can become output. The mechanism is not ordinary batching. Tokens inside the proposal remain autoregressive, and the target model still defines the intended output distribution when an exact speculative-sampling algorithm is used. The useful change is that one expensive verification pass can account for several output positions when enough proposals are accepted.

Artificial Intelligence 24 Sep 2026 5 min read

Sliding-Window Attention Bounds Each Query to a Local Token Horizon

A causal attention layer does not always expose every earlier token to every query. With a sliding window of width w, the query at position i can be restricted to recent positions rather than the full prefix. The attention graph becomes local: old tokens fall outside the direct edge set even though they remain part of the sequence. That boundary changes computation, memory traffic, and information paths at the same time. It is not merely an optimized implementation of full attention. Once the mask removes distant key-value pairs, the layer implements a different dependency pattern.

Artificial Intelligence 24 Sep 2026 5 min read

Sliding-Window Attention Bounds Direct Token Access

A causal transformer does not have to expose every earlier token to every query position. With sliding-window attention, position i can attend only to a bounded range of preceding keys. Tokens outside that range are absent from that layer’s attention operation, even if they still belong to the model input. That boundary changes more than the attention matrix shape. It separates direct access from information that can reach a position only after being propagated through intermediate hidden states.

Artificial Intelligence 24 Sep 2026 5 min read

Rotary Position Embedding Converts Absolute Indices into Relative Attention Phases

Self-attention can compare token content without assigning an order to token positions. Rotary Position Embedding (RoPE) inserts position into that comparison by rotating paired coordinates of queries and keys. Each token receives an absolute rotation angle, yet the query-key inner product reduces those two absolute angles to their difference. That algebraic cancellation is the central mechanism. RoPE does not add a position vector to the hidden state. It changes the orientation of query and key components before their dot product is evaluated.

Artificial Intelligence 24 Sep 2026 5 min read

RoPE Position Interpolation Compresses Indices Before Rotation

Rotary position embeddings apply a position-dependent rotation to query and key components before their dot product is evaluated. If a model was trained with positions up to a reference length, simply sending much larger position indices into the same rotation rule can place attention in angular regimes that were not represented in that training range. Position interpolation changes the coordinate supplied to RoPE. For an extension factor s > 1, a simplified form maps

Artificial Intelligence 24 Sep 2026 4 min read

RoPE Encodes Relative Offsets Through Rotated Query-Key Phases

Rotary Position Embedding (RoPE) applies position-dependent rotations to query and key coordinates before their attention dot product. The resulting score carries relative position through the phase difference between those rotations rather than through an additive position vector attached to the token representation. This distinction is structural. RoPE starts from absolute indices for each rotation, yet the query-key inner product can be written in terms of the offset between their positions.

Artificial Intelligence 24 Sep 2026 5 min read

RMSNorm Scales Hidden States Without Mean Centering

RMSNorm normalizes a hidden-state vector without subtracting its coordinate mean. That single omission separates it from LayerNorm at the mathematical interface: RMSNorm controls scale through a root-mean-square statistic, while any common offset across coordinates remains part of the transformed representation. For a vector x with width D, a common RMSNorm form is: rms(x) = sqrt(mean(x_i^2) + eps) y_i = g_i * x_i / rms(x) Here g_i is a trainable per-coordinate scale and eps is a small positive term defined by the model implementation. Exact parameterization and numeric details belong to the checkpoint and runtime contract.

Artificial Intelligence 24 Sep 2026 5 min read

RMSNorm Rescales Hidden States Without Mean Centering

RMSNorm rescales a vector from its root mean square without first subtracting the vector mean. That missing centering operation is the defining difference from LayerNorm: both can control vector scale, but only LayerNorm explicitly shifts the normalized coordinates around a zero sample mean. For a hidden vector x with width d, a common RMSNorm form is: rms = sqrt((1/d) * sum(x_i^2) + epsilon) y_i = gain_i * x_i / rms The exact placement of epsilon, numeric precision used for the reduction, and presence of extra affine terms depend on the implementation. The structural operation remains division by an RMS statistic rather than division by a standard deviation computed after mean subtraction.

Artificial Intelligence 24 Sep 2026 4 min read

QK Normalization Bounds Attention Logit Scale Before Softmax

QK normalization inserts normalization on query and key vectors before the attention dot product. The operation changes the geometry of the score calculation: vector magnitude no longer enters the dot product in the same unrestricted form, while directional alignment remains part of the score. For one query vector q and key vector k, ordinary scaled dot-product attention forms a score such as: s = dot(q, k) / sqrt(D) A QK-normalized variant first applies the model’s specified normalization functions:

Artificial Intelligence 24 Sep 2026 5 min read

Prefix Caching Reuses KV State for Identical Token Prefixes

Autoregressive serving often receives requests that start with the same long token sequence. A system prompt, tool schema, fixed document header, or other repeated context can cause the model to compute the same prefix attention state again for each request. Prefix caching targets that repeated prefill work by retaining compatible key-value state and attaching later requests to it. The reuse boundary is exact tokenized state, not semantic similarity. Two prompts that express the same idea but tokenize differently do not produce an interchangeable prefix cache entry.

Artificial Intelligence 24 Sep 2026 6 min read

Packed Sequences Need Boundary-Aware Attention Masks

Sequence packing reduces padding by placing several variable-length examples into one token buffer. The storage layout may look like one long sequence, but the examples are still semantically independent. A standard causal mask does not preserve that independence by itself. For a decoder-only Transformer, causal masking blocks attention to future positions. It does not normally block attention to earlier positions that belong to another packed example. If boundaries are ignored, tokens in a later example can attend to keys and values from an earlier one. The model then receives context that the data pipeline intended to keep separate.

Artificial Intelligence 24 Sep 2026 5 min read

Multi-Token Prediction Adds Parallel Future-Token Losses to a Shared Model Trunk

A next-token language model normally applies one predictive objective at position t: the hidden representation at that position is used to score token x_(t+1). Multi-token prediction changes that training boundary. One shared model trunk produces the representation, while multiple output heads predict several subsequent tokens from that position. The mechanism adds supervision at multiple future offsets without requiring a separate transformer trunk for every offset. It is therefore a change to the training objective and prediction heads, not a claim that ordinary autoregressive generation can emit several unchecked tokens as one exact step.

Artificial Intelligence 24 Sep 2026 5 min read

Multi-Head Latent Attention Compresses KV State Before Head-Specific Expansion

Multi-Head Latent Attention changes the state retained across autoregressive decoding. Instead of requiring a full key tensor and value tensor for every cached token and attention head, the architecture can retain a lower-dimensional latent representation and derive head-specific key/value information through projection structure. That distinction matters at the cache boundary. Ordinary multi-head attention commonly stores already projected per-head keys and values. MLA moves part of that representation behind a compact bottleneck, so persistent state and compute no longer have the same shape.

Artificial Intelligence 24 Sep 2026 5 min read

MoE Expert Capacity Bounds Token Routing

A sparse Mixture-of-Experts layer can contain many expert networks while activating only a small subset for each token. That conditional computation depends on a router, but router scores alone do not determine the executed graph. In implementations with bounded expert batches, each expert also has a finite number of token slots. This creates a second boundary after expert selection: a token can prefer an expert that has no remaining capacity. The handling of that overflow is an implementation and architecture choice with direct consequences for training and serving.

Artificial Intelligence 24 Sep 2026 6 min read

KV Cache Quantization Trades Precision for Serving Memory

Autoregressive decoding keeps past attention keys and values so each new token can reuse earlier projections instead of recomputing them. That KV cache grows with sequence length, layer count, batch size, and the number and width of cached key/value heads. Reducing its numeric precision can cut the bytes occupied by those cached tensors, but it also changes the values consumed by later attention operations. This makes KV cache quantization different from compressing data that is only stored and restored losslessly. The quantized cache remains on the inference path. Each subsequent query can interact with approximated keys and values, so the relevant question is not only how many bytes are saved, but where quantization error enters attention and how the serving implementation contains it.

Artificial Intelligence 24 Sep 2026 6 min read

Grouped-Query Attention Shrinks the KV Cache by Sharing Key-Value Heads

Grouped-query attention changes the number of key and value heads that must be stored during autoregressive decoding. Instead of giving every query head its own key-value pair, several query heads address the same key-value head. The query side can retain many heads while the persistent KV state uses fewer independent projections. This is an architectural change to attention, not a cache compression codec. The smaller cache follows from producing fewer distinct key and value heads per token.

Artificial Intelligence 24 Sep 2026 6 min read

Grouped-Query Attention Shares Key-Value Heads Across Query Groups

Multi-head attention does not require every query head to own a distinct key head and value head. Grouped-query attention, usually abbreviated GQA, partitions query heads into groups and assigns one key-value head to each group. Query projections remain separate, but several query heads read from the same projected keys and values. That distinction changes parameter shapes and cached state without collapsing the query heads into one attention computation. Each query head still produces its own attention scores because its query vector is different.

Artificial Intelligence 24 Sep 2026 6 min read

Expert Capacity Turns Uneven MoE Routing into Token Overflow

A sparse Mixture-of-Experts layer can route many tokens toward the same expert even when every expert has identical nominal capacity. The router makes token-dependent choices, while distributed execution commonly allocates bounded token slots per expert. When those two mechanisms disagree, an expert can receive more assignments than its execution buffer admits. That boundary is not an inherent property of every MoE architecture. It is a property of capacity-constrained routing designs, including the routing formulation described for Switch Transformers. In such systems, expert capacity converts an uneven routing distribution into an operational event: some assignments fit, while assignments beyond the capacity limit require an explicit overflow policy.

Artificial Intelligence 24 Sep 2026 5 min read

Contrastive Search Penalizes Hidden-State Repetition During Decoding

Autoregressive generation exposes a distribution over the next token, but a decoder still has to choose which candidate to append. Contrastive search changes that choice by combining model probability with a penalty for candidates whose resulting hidden state is too similar to hidden states already present in the generated prefix. The mechanism operates only at inference time. It does not alter model parameters or the next-token distribution itself. Instead, it changes the ranking used to select a token from a restricted candidate set.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Stabilize Windowed KV-Cache Streaming

A decoder can keep its KV cache bounded by discarding states that fall outside a fixed recent-token window. The memory bound is attractive, but naive eviction can sharply degrade generation even when the removed tokens carry little obvious semantic value. A small set of key-value states from the beginning of the sequence can change that behavior. This effect is associated with attention sinks: initial positions that receive substantial attention mass even when their token content is not important to the current prediction. StreamingLLM reported that preserving those initial states together with a rolling window of recent states recovers stable language-model behavior that ordinary window eviction can lose.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Stabilize Streaming Cache Eviction

A bounded KV cache can keep inference memory from growing with an unending token stream, but deleting every old token in strict arrival order can disrupt attention more than the missing content alone suggests. In some autoregressive transformers, early positions receive substantial attention even when their token semantics are not directly relevant to the current prediction. These positions act as attention sinks. The serving consequence is specific: a sliding cache that preserves a small initial prefix alongside the most recent tokens can behave differently from a same-sized cache containing only the newest tokens. The mechanism concerns attention state and positional handling, not retrieval of forgotten text.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Absorb Probability Mass Without Carrying Matching Content

A causal attention head can assign noticeable probability mass to an early token even when that token is not a strong semantic match for the current query. The token then behaves as an attention sink: its key remains a convenient destination for probability that the head does not direct toward content-bearing positions. This behavior matters during autoregressive serving because a KV-cache policy can preserve recent tokens yet still alter model behavior sharply if it removes sink positions. Recency alone does not describe the functional role of every cached key.