Skip to content

Topic archive

Artificial Intelligence

Artificial intelligence includes machine learning, deep learning, and systems that apply models to practical tasks.

396 articles
Artificial Intelligence 23 Sep 2026 5 min read

Switch Routing Capacity Turns Expert Imbalance into Token Overflow

A Switch-style sparse layer can send many tokens toward the same expert even when every expert has identical compute capacity. The router makes token-level choices from model-produced scores; it does not inherently produce an even partition. A fixed expert capacity therefore creates a hard boundary between routing preference and the amount of expert computation admitted for a batch. In the Switch Transformer formulation, each token is routed to the expert with the highest router probability. Each expert receives a fixed token capacity derived from the token count, expert count, and a capacity factor. When assignments exceed that capacity, excess tokens overflow instead of enlarging the expert batch without bound.

Artificial Intelligence 23 Sep 2026 3 min read

Stable Softmax Subtracts the Maximum Logit Before Exponentiation

A decoder can receive logits large enough that direct exponentiation is numerically unsafe even though the intended probability distribution is ordinary. Softmax does not require exponentiating the original values. Subtracting the largest logit from every logit produces the same distribution in exact real arithmetic while moving the exponentials into a safer numeric range. This shift is a property of softmax itself, not a model-specific heuristic. A shared shift cancels during normalization For logits (z_1,\ldots,z_n), softmax assigns

Artificial Intelligence 23 Sep 2026 5 min read

Speculative Sampling Preserves Target Distribution Through Rejection Correction

A draft model can propose a token that the target model would not have sampled from the same random draw, yet speculative sampling can still preserve the target model’s distribution. The key is not that the draft model predicts the target perfectly. Distributional correctness comes from the acceptance rule and the correction applied after rejection. This separates two properties that are often grouped together. Draft quality controls how frequently proposals survive verification. The rejection-correction construction controls whether the resulting sample follows the target distribution.

Artificial Intelligence 23 Sep 2026 5 min read

Speculative Decoding Verifies Draft Tokens Before Acceptance

Autoregressive generation normally advances one token at a time: the model scores the next-token distribution, a decoding rule selects a token, and that token becomes part of the context for the next forward step. Speculative decoding changes the execution schedule. A cheaper draft process proposes several future tokens, while the target model retains authority over which proposals can enter the generated sequence. That separation is the core constraint. Draft tokens are predictions about future target-model decisions, not a replacement distribution that can be appended unchecked. The serving implementation must verify them against target-model scores and apply the acceptance rule required by the chosen speculative algorithm.

Artificial Intelligence 23 Sep 2026 5 min read

Sliding-Window KV Caches Change Which Tokens Remain Addressable

Autoregressive decoding can reuse key and value tensors from earlier token positions instead of recomputing them at every step. A conventional KV cache therefore grows as generation advances. A sliding-window cache places a bound on that retained state by keeping only a recent region. The memory bound changes more than allocation size. Once an old key-value entry is evicted, a later attention operation cannot directly address that position through the cache. The resulting behavior depends on the model’s attention pattern, positional scheme, cache implementation, and any layers that use a different attention span.

Artificial Intelligence 23 Sep 2026 6 min read

Sliding-Window Attention Bounds Active KV State by Token Distance

Sliding-window attention imposes a finite token-distance boundary on causal attention. At position (t), a query can address only a recent interval of key and value positions rather than the full prefix. Once a cached position falls permanently outside that interval, later queries governed by the same local rule cannot address it. That boundary changes the state required for autoregressive decoding. Full causal attention keeps usable KV state growing with sequence length. A fixed local window can keep the active KV span bounded, provided the runtime evicts or overwrites entries that have become unreachable.

Artificial Intelligence 23 Sep 2026 5 min read

RoPE Position Shifts Preserve Relative Phase but Change Absolute Rotation

Rotary position embeddings, commonly called RoPE, inject position into attention by rotating pairs of query and key channels with angles determined by token position. The rotation happens before the query-key dot product. As a result, position is not represented by adding a standalone vector to each token representation. This distinction matters during inference. A token’s cached key already contains the rotation associated with its assigned position. Reusing that key under a different sequence coordinate without an equivalent transformation changes the attention computation, even when the token content is identical.

Artificial Intelligence 23 Sep 2026 4 min read

RoPE Position Interpolation Compresses Relative Angles Across Longer Contexts

Rotary position embeddings encode position by rotating query and key components with angles derived from token indices. When a model is asked to operate beyond the position range used during its original training, those angles can enter a regime the model did not encounter. Position interpolation addresses that boundary by mapping a longer sequence into a smaller position-coordinate range before the rotary angles are computed. The mechanism does not add extra tokens to the model’s architectural state. It changes the coordinates supplied to the positional transform. That distinction matters because a longer accepted input length and reliable behavior at that length are separate properties.

Artificial Intelligence 23 Sep 2026 4 min read

RoPE KV Caches Preserve Token Position Across Incremental Decoding

In a decoder using rotary position embeddings, a key written to the KV cache already carries the rotation associated with its token position. Incremental decoding can reuse that key directly. Applying the current token position to the cached key again changes the attention geometry. This makes position state part of the cache contract even when the cache API appears to store only tensors. Rotation is applied before a key enters attention For one two-dimensional component pair, RoPE applies a position-dependent rotation. Writing (R_m) for the rotation at position (m), a query and key become

Artificial Intelligence 23 Sep 2026 5 min read

RMSNorm Scales Activations Without Mean Centering

RMSNorm rescales a hidden vector using its root-mean-square magnitude, but it does not subtract the vector’s feature mean first. That omission is not merely a shorter expression for LayerNorm. It changes which transformations of the input disappear under normalization and which remain visible to later operations. For transformer implementations, that distinction matters at the boundary between residual state, normalization, and the next projection. The denominator comes from the second raw moment For a hidden vector (x \in \mathbb{R}^d), a common RMSNorm form is

Artificial Intelligence 23 Sep 2026 6 min read

RMSNorm Rescales Activations Without Mean Centering

RMSNorm normalizes a vector by its root mean square rather than by a centered standard deviation. That small change removes mean subtraction from the normalization step. As a result, RMSNorm and LayerNorm respond similarly to some scale changes but differently to additive shifts in the hidden state. The distinction matters in transformer implementations because normalization is part of the residual path geometry. Replacing one normalization rule with another is not merely an arithmetic shortcut; it changes which transformations of an activation vector are canceled and which remain visible to later computation.

Artificial Intelligence 23 Sep 2026 5 min read

Repetition Penalty Rewrites Logits for Seen Token IDs

A repetition penalty can act before sampling by changing the logits of token IDs that already occur in a selected token history. The operation does not need to compare words, phrases, or rendered strings. Its unit can be the tokenizer’s integer ID, which gives the mechanism a narrower meaning than its name may suggest. That distinction matters when a decoder emits subword tokens. Two strings that appear similar to a person can map to different token sequences, while a token reused inside unrelated words can still be marked as previously seen.

Artificial Intelligence 23 Sep 2026 6 min read

Repetition Penalties Alter Token Scores Based on Prior Output

A repetition penalty changes the next-token distribution without changing the model parameters or hidden state computation. The model still produces its logits from the current context, but the decoder edits selected scores according to tokens that have already appeared. Sampling then operates on the edited scores rather than directly on the model output. That distinction matters when reproducing generation behavior. Two systems can run identical model weights on identical token IDs and still emit different continuations because their repetition rules differ in formula, token-history scope, or position in the decoding pipeline.

Artificial Intelligence 23 Sep 2026 5 min read

Prefix KV Cache Reuse Depends on Exact Context Identity

A prefix cache can remove repeated prefill work without changing the model output, but only when the cached key and value states represent the same prefix context that the new request would have produced. Matching visible text is not enough. The serving path ultimately operates on token IDs, positions, model parameters, and implementation-specific attention state. This boundary makes prefix caching different from a generic text cache. A text cache stores a result associated with an input. A KV prefix cache stores intermediate states whose validity depends on the computation that created them.

Artificial Intelligence 23 Sep 2026 5 min read

PagedAttention Maps Logical KV Blocks to Noncontiguous Physical Memory

An autoregressive request grows its KV cache as tokens arrive, but its final sequence length is not known when decoding begins. Reserving one contiguous region for the maximum possible sequence length ties memory to capacity that may never be used. PagedAttention changes that allocation boundary: a sequence is represented as logical KV blocks, while a block table maps those logical blocks to physical blocks that need not be adjacent in GPU memory.

Artificial Intelligence 23 Sep 2026 4 min read

Logit Bias Alters Token Odds Before Sampling

A token-level bias is usually applied to model scores before probabilities are normalized. That placement matters. Adding a constant to one token’s logit changes its odds relative to every token that does not receive the same constant, even though the model parameters and hidden state remain unchanged. The mechanism is simple, but its operational effect depends on the rest of the decoding pipeline. Additive bias acts on score differences For a vocabulary with logits (z_1, \ldots, z_V), softmax assigns token (i) the probability

Artificial Intelligence 23 Sep 2026 5 min read

Label Smoothing Redistributes Target Probability Before Cross-Entropy

A classifier trained with one-hot targets is asked to place all target probability on a single class. Cross-entropy does not require that target representation. Label smoothing changes the target distribution before the loss is evaluated, so the model receives a different gradient even when its logits and predicted probabilities are unchanged. This distinction matters because label smoothing is not a decoding rule and does not alter inference by itself. It changes the training objective. The resulting model parameters can differ because the optimizer follows gradients computed against softened targets.

Artificial Intelligence 23 Sep 2026 4 min read

L2 Normalization Turns Embedding Dot Products Into Cosine Similarity

Two embedding vectors can point in nearly the same direction yet have very different magnitudes. A raw dot product responds to both properties. L2 normalization removes the magnitude term, so the same dot-product operation becomes a comparison of direction. That change is not merely a numerical convenience. It changes the retrieval objective whenever vector norms carry information or vary across items. Dot product contains a magnitude term For nonzero vectors (x) and (y),

Artificial Intelligence 23 Sep 2026 5 min read

Grouped-Query Attention Shares KV Heads Across Query Groups

Grouped-query attention (GQA) changes a specific structural ratio inside an attention layer: the number of query heads can exceed the number of key and value heads. Several query heads then consume the same projected key and value head. The attention calculation remains head-specific on the query side, while KV state is shared within each group. That asymmetry matters during autoregressive decoding because cached keys and values persist for prior tokens. Reducing the count of distinct KV heads reduces the amount of per-token KV state that must remain available to later decoding steps.

Artificial Intelligence 23 Sep 2026 4 min read

Grouped-Query Attention Shares Key-Value Heads Across Query Heads

Grouped-query attention changes a specific part of multi-head attention: several query heads use the same key and value head. The query projections remain separate, so those query heads can produce different attention weights, but they read keys and values from a shared projected representation. That distinction matters during inference. A decoder cache stores past keys and values, not past queries. Reducing the number of key-value heads can therefore reduce KV cache storage without reducing the number of query heads by the same factor.

Artificial Intelligence 23 Sep 2026 5 min read

Global Gradient Clipping Caps Norm Before the Optimizer Step

A single training step can produce gradients whose combined magnitude is far larger than nearby steps. Global norm clipping changes that gradient set before the optimizer consumes it. When the measured norm exceeds a configured threshold, every selected gradient is multiplied by the same scale factor. The mechanism is simple, but its boundary matters. Clipping controls the norm of the gradients supplied to the optimizer. It does not directly impose the same bound on the eventual parameter update, especially when the optimizer keeps momentum or adaptive state.

Artificial Intelligence 23 Sep 2026 6 min read

Beam Search Length Normalization Changes Sequence Ranking

Beam search ranks partial sequences by scores accumulated across decoding steps. When that score is the sum of token log-probabilities, every additional token contributes a value that is normally non-positive. A longer candidate therefore has more opportunities to reduce its raw score, even when its continuation is locally plausible. That property is not a defect in probability theory. It follows from comparing complete sequences with different numbers of conditional factors. It becomes an implementation concern when a decoder is expected to produce useful completions rather than simply rank sequences by unmodified model probability.

Artificial Intelligence 23 Sep 2026 4 min read

Attention Sinks Make Naive KV Eviction Unstable

A fixed-size KV cache seems to invite a simple policy: keep the newest entries and evict the oldest. For some decoder transformers, that policy removes positions that later queries continue to assign substantial attention mass. Those early positions act as attention sinks, and dropping them can change the attention distribution far more than their semantic content suggests. The effect matters specifically at inference time when a serving system truncates cached keys and values. It is not a general claim that the first token is semantically privileged, nor that every transformer exhibits the same pattern.

Artificial Intelligence 23 Sep 2026 6 min read

Attention Logit Scaling Keeps Dot Products in a Stable Softmax Range

A dot product between a query and a key tends to grow in magnitude as their dimension grows. In scaled dot-product attention, the score is divided by the square root of the key dimension before softmax. That factor is not a cosmetic normalization. It controls the scale presented to softmax under a specific statistical assumption about the query and key components. The familiar expression is Attention(Q, K, V) = softmax(QK^T / sqrt(d_k)) V where d_k is the query-key dimension for one attention head. The scaling term affects the distribution of attention probabilities even though it does not change the ordering of logits by itself.