Skip to content

Archive

Transformers

97 articles
Artificial Intelligence 24 Sep 2026 5 min read

ALiBi Adds Distance-Proportional Bias to Attention Scores

ALiBi changes an attention score before softmax rather than adding a positional vector to the token representation. For a causal transformer, a key farther behind the current query receives a larger negative offset. The offset is linear in token distance and uses a slope associated with the attention head. A simplified score for head h can be written as: score_h(i, j) = q_i k_j^T / sqrt(d) - m_h * (i - j) for an allowed causal pair with j <= i and positive slope m_h. The causal mask still decides which future positions are inaccessible. ALiBi changes the relative scores among positions that remain eligible.

Artificial Intelligence 24 Sep 2026 5 min read

ALiBi Adds Distance Penalties Directly to Attention Logits

Attention with Linear Biases (ALiBi) changes a causal attention score before softmax by adding a penalty whose magnitude grows with token distance. Position is therefore represented in the score path rather than by adding a positional vector to each token representation. For a query at position i attending to a key at position j, one head can be written schematically as: score(i, j) = q_i · k_j / sqrt(d_k) - m_h * (i - j) for causal positions j <= i. The positive slope m_h is specific to attention head h. The causal mask still prevents access to future positions; the linear term changes the relative preference among positions that remain visible.

Artificial Intelligence 24 Sep 2026 5 min read

Activation Outliers Distort Low-Bit Quantization Scales

A tensor can be easy to represent in floating point and awkward to map into a small integer range. The problem becomes acute when most activation values occupy a narrow interval while a small number have much larger magnitude. A shared quantization scale must cover those extremes, so the ordinary values receive fewer representable levels. That is the practical effect of activation outliers in low-bit inference. The issue is not merely that an outlier is numerically large. Its location, persistence, and relationship to the axis over which a scale is shared determine whether it materially degrades the quantized representation.

Artificial Intelligence 23 Sep 2026 5 min read

Switch Routing Capacity Turns Expert Imbalance into Token Overflow

A Switch-style sparse layer can send many tokens toward the same expert even when every expert has identical compute capacity. The router makes token-level choices from model-produced scores; it does not inherently produce an even partition. A fixed expert capacity therefore creates a hard boundary between routing preference and the amount of expert computation admitted for a batch. In the Switch Transformer formulation, each token is routed to the expert with the highest router probability. Each expert receives a fixed token capacity derived from the token count, expert count, and a capacity factor. When assignments exceed that capacity, excess tokens overflow instead of enlarging the expert batch without bound.

Artificial Intelligence 23 Sep 2026 5 min read

Sliding-Window KV Caches Change Which Tokens Remain Addressable

Autoregressive decoding can reuse key and value tensors from earlier token positions instead of recomputing them at every step. A conventional KV cache therefore grows as generation advances. A sliding-window cache places a bound on that retained state by keeping only a recent region. The memory bound changes more than allocation size. Once an old key-value entry is evicted, a later attention operation cannot directly address that position through the cache. The resulting behavior depends on the model’s attention pattern, positional scheme, cache implementation, and any layers that use a different attention span.

Artificial Intelligence 23 Sep 2026 6 min read

Sliding-Window Attention Bounds Active KV State by Token Distance

Sliding-window attention imposes a finite token-distance boundary on causal attention. At position (t), a query can address only a recent interval of key and value positions rather than the full prefix. Once a cached position falls permanently outside that interval, later queries governed by the same local rule cannot address it. That boundary changes the state required for autoregressive decoding. Full causal attention keeps usable KV state growing with sequence length. A fixed local window can keep the active KV span bounded, provided the runtime evicts or overwrites entries that have become unreachable.

Artificial Intelligence 23 Sep 2026 5 min read

RoPE Position Shifts Preserve Relative Phase but Change Absolute Rotation

Rotary position embeddings, commonly called RoPE, inject position into attention by rotating pairs of query and key channels with angles determined by token position. The rotation happens before the query-key dot product. As a result, position is not represented by adding a standalone vector to each token representation. This distinction matters during inference. A token’s cached key already contains the rotation associated with its assigned position. Reusing that key under a different sequence coordinate without an equivalent transformation changes the attention computation, even when the token content is identical.

Artificial Intelligence 23 Sep 2026 4 min read

RoPE Position Interpolation Compresses Relative Angles Across Longer Contexts

Rotary position embeddings encode position by rotating query and key components with angles derived from token indices. When a model is asked to operate beyond the position range used during its original training, those angles can enter a regime the model did not encounter. Position interpolation addresses that boundary by mapping a longer sequence into a smaller position-coordinate range before the rotary angles are computed. The mechanism does not add extra tokens to the model’s architectural state. It changes the coordinates supplied to the positional transform. That distinction matters because a longer accepted input length and reliable behavior at that length are separate properties.

Artificial Intelligence 23 Sep 2026 4 min read

RoPE KV Caches Preserve Token Position Across Incremental Decoding

In a decoder using rotary position embeddings, a key written to the KV cache already carries the rotation associated with its token position. Incremental decoding can reuse that key directly. Applying the current token position to the cached key again changes the attention geometry. This makes position state part of the cache contract even when the cache API appears to store only tensors. Rotation is applied before a key enters attention For one two-dimensional component pair, RoPE applies a position-dependent rotation. Writing (R_m) for the rotation at position (m), a query and key become

Artificial Intelligence 23 Sep 2026 5 min read

RMSNorm Scales Activations Without Mean Centering

RMSNorm rescales a hidden vector using its root-mean-square magnitude, but it does not subtract the vector’s feature mean first. That omission is not merely a shorter expression for LayerNorm. It changes which transformations of the input disappear under normalization and which remain visible to later operations. For transformer implementations, that distinction matters at the boundary between residual state, normalization, and the next projection. The denominator comes from the second raw moment For a hidden vector (x \in \mathbb{R}^d), a common RMSNorm form is

Artificial Intelligence 23 Sep 2026 6 min read

RMSNorm Rescales Activations Without Mean Centering

RMSNorm normalizes a vector by its root mean square rather than by a centered standard deviation. That small change removes mean subtraction from the normalization step. As a result, RMSNorm and LayerNorm respond similarly to some scale changes but differently to additive shifts in the hidden state. The distinction matters in transformer implementations because normalization is part of the residual path geometry. Replacing one normalization rule with another is not merely an arithmetic shortcut; it changes which transformations of an activation vector are canceled and which remain visible to later computation.

Artificial Intelligence 23 Sep 2026 5 min read

Grouped-Query Attention Shares KV Heads Across Query Groups

Grouped-query attention (GQA) changes a specific structural ratio inside an attention layer: the number of query heads can exceed the number of key and value heads. Several query heads then consume the same projected key and value head. The attention calculation remains head-specific on the query side, while KV state is shared within each group. That asymmetry matters during autoregressive decoding because cached keys and values persist for prior tokens. Reducing the count of distinct KV heads reduces the amount of per-token KV state that must remain available to later decoding steps.

Artificial Intelligence 23 Sep 2026 4 min read

Grouped-Query Attention Shares Key-Value Heads Across Query Heads

Grouped-query attention changes a specific part of multi-head attention: several query heads use the same key and value head. The query projections remain separate, so those query heads can produce different attention weights, but they read keys and values from a shared projected representation. That distinction matters during inference. A decoder cache stores past keys and values, not past queries. Reducing the number of key-value heads can therefore reduce KV cache storage without reducing the number of query heads by the same factor.

Artificial Intelligence 23 Sep 2026 6 min read

Attention Logit Scaling Keeps Dot Products in a Stable Softmax Range

A dot product between a query and a key tends to grow in magnitude as their dimension grows. In scaled dot-product attention, the score is divided by the square root of the key dimension before softmax. That factor is not a cosmetic normalization. It controls the scale presented to softmax under a specific statistical assumption about the query and key components. The familiar expression is Attention(Q, K, V) = softmax(QK^T / sqrt(d_k)) V where d_k is the query-key dimension for one attention head. The scaling term affects the distribution of attention probabilities even though it does not change the ordering of logits by itself.

Artificial Intelligence 23 Sep 2026 5 min read

Attention Entropy Measures Concentration, Not Causal Importance

An attention head can place most of its probability mass on one position and still provide little evidence that this position controls the final model output. The entropy of its attention weights captures concentration, not causal influence. That distinction matters when attention maps are inspected as diagnostics. Entropy can reveal whether a head spreads mass broadly or focuses it narrowly for a given query. It cannot, by itself, establish that the highest-weight token carries the feature responsible for a downstream prediction.

Artificial Intelligence 22 Sep 2026 7 min read

Prefix KV Cache Reuse Depends on Exact Token History

A KV cache entry is not a reusable representation of arbitrary text that happens to look similar. For an autoregressive transformer, cached keys and values are intermediate states produced for a specific token prefix under a specific execution context. Reusing them is valid only when the new request reaches the same state boundary. That boundary is stricter than matching visible characters. Tokenization, token order, position handling, model identity, adapter state, and other inputs that affect hidden states can all determine whether a cached prefix still represents the computation required by the new request.

Artificial Intelligence 22 Sep 2026 6 min read

Padding Masks Do Not Remove Padding Compute in Dense Attention

A batch can contain two prompts with very different token counts yet represent both with the same rectangular tensor. The shorter prompt is extended with padding so its tensor shape matches the longest sequence in the batch. An attention mask can stop those padded positions from contributing to attention probabilities, but that semantic exclusion does not imply that dense kernels skip every operation associated with the padded rows and columns.

Artificial Intelligence 18 Sep 2026 5 min read

Treat the Logit Lens as a Readout, Not a Causal Trace

A transformer can expose an intermediate residual state that strongly favors a token under the model’s final vocabulary projection, then produce a different token after later blocks run. The logit lens makes that intermediate preference visible. It does not establish that the preference caused the final output. That boundary matters when developers use layer-by-layer token rankings to inspect model behavior. The logit lens is a readout: it asks what the model’s output head would report if applied to an intermediate representation. The actual forward pass asks a different question because every remaining block can transform that representation before the final readout.

Artificial Intelligence 17 Sep 2026 6 min read

Retain Attention Sinks in Bounded KV Caches

A bounded KV cache seems to invite a simple eviction rule: once the cache is full, discard the oldest key-value pair and keep the most recent tokens. For some decoder-only transformers, that rule can degrade generation sharply after the sequence moves beyond the retained window. The failure is not explained only by missing old semantic content. Early tokens can receive substantial attention even when their text carries little useful information for the current prediction.

Artificial Intelligence 17 Sep 2026 7 min read

Balance Token Routing in Sparse Mixture-of-Experts Models

A sparse mixture-of-experts layer can contain many expert networks while evaluating only a small subset for each token. The router makes that sparsity possible: it assigns scores to experts, selects a limited set, and sends each token through the selected computation paths. That selection is not only an optimization detail. If many tokens concentrate on a few experts, some devices can receive much more work than others, capacity limits can discard or redirect assignments, and experts that receive little traffic get fewer task gradients. Router balance therefore affects both computation and the function represented by the model.

Artificial Intelligence 16 Sep 2026 6 min read

Test Transformer Mechanisms with Activation Patching

A transformer can produce two different outputs from prompts that differ in one relevant detail, yet inspection of attention weights or hidden-state similarity does not establish which internal states actually matter for that difference. Activation patching addresses a narrower question by intervening on a forward pass: replace a selected activation with the corresponding activation from another run, then measure how the output changes. The result is causal with respect to that intervention. It does not automatically identify a complete circuit, a unique mechanism, or a human-readable feature. That boundary is central to using patching results correctly.

Artificial Intelligence 16 Sep 2026 5 min read

Read Intermediate Transformer States With Logit Lens

A decoder-only transformer produces its next-token distribution only after the final hidden state has passed through the model’s output normalization and vocabulary projection. Logit lens reuses that output path on states from earlier transformer blocks. The result is a sequence of vocabulary distributions that can expose how token preferences change with depth. The method is attractive because it maps internal vectors into familiar token space without fitting a separate classifier. That convenience also creates a sharp interpretive boundary: an intermediate state was not necessarily optimized to be directly decoded by the final output map. A readable token distribution is a probe of that state, not a guarantee that the model has already settled on the same prediction.

Artificial Intelligence 16 Sep 2026 6 min read

Read Intermediate Transformer Predictions with Tuned Lenses

A transformer can carry useful information about its eventual next-token distribution several blocks before the final layer. Reading that information is not as simple as applying the model’s output projection to every intermediate hidden state. The final output head is calibrated for representations at the end of the network, while residual representations can shift across depth. A tuned lens addresses that mismatch with a separate affine translator for each inspected layer. The translator maps an intermediate residual state into a representation that the frozen final normalization and output projection can decode. This produces a token distribution that can be compared across layers without assuming that every layer already uses the final representation basis.

Artificial Intelligence 16 Sep 2026 6 min read

Quantize KV Caches to Reduce Long-Context Inference Memory

Autoregressive transformer inference reuses attention keys and values from earlier tokens so each new token does not recompute the full prefix. That reuse creates the KV cache, whose memory grows with the number of cached tokens. At long context lengths or high request concurrency, the cache can become a major part of inference memory. KV cache quantization changes the representation of those stored tensors. Keys and values are written in a lower-precision format together with any scale or metadata needed for reconstruction. Attention later consumes reconstructed values or uses a kernel that handles the quantized representation directly.