Skip to content

Topic archive

Artificial Intelligence

Artificial intelligence includes machine learning, deep learning, and systems that apply models to practical tasks.

396 articles
Artificial Intelligence 24 Sep 2026 6 min read

KV Cache Quantization Trades Precision for Serving Memory

Autoregressive decoding keeps past attention keys and values so each new token can reuse earlier projections instead of recomputing them. That KV cache grows with sequence length, layer count, batch size, and the number and width of cached key/value heads. Reducing its numeric precision can cut the bytes occupied by those cached tensors, but it also changes the values consumed by later attention operations. This makes KV cache quantization different from compressing data that is only stored and restored losslessly. The quantized cache remains on the inference path. Each subsequent query can interact with approximated keys and values, so the relevant question is not only how many bytes are saved, but where quantization error enters attention and how the serving implementation contains it.

Artificial Intelligence 24 Sep 2026 6 min read

KV Cache Quantization Reduces Stored Attention State at a Reconstruction Cost

Autoregressive decoding appends key and value tensors to a cache at every transformer layer. The cache prevents prior tokens from being projected into keys and values again, but its storage grows with sequence length. KV cache quantization changes that storage representation: older or selected cache entries are encoded with fewer bits, then reconstructed when attention consumes them. The mechanism is a memory-format trade. It does not remove tokens from the attention context and it does not change the model weights. It reduces bytes used by cached state while introducing quantization error, scale or zero-point metadata, and conversion work on the decode path.

Artificial Intelligence 24 Sep 2026 4 min read

KV Cache Quantization Perturbs Attention Through Stored Keys and Values

During autoregressive inference, previously computed keys and values are reused from the KV cache instead of being recomputed for every new token. Quantizing that cache changes more than its byte representation. The stored approximation becomes an input to later attention operations, so its error can alter both attention scores and the vectors combined by those scores. This boundary differs from quantizing model weights. A weight tensor is reused across requests, while KV state is generated from the current sequence and grows with its cached length. Its numeric range can also vary across layers, heads, positions, and requests.

Artificial Intelligence 24 Sep 2026 5 min read

Iteration-Level Scheduling Rebuilds LLM Batches Between Decode Steps

Autoregressive serving does not give every request the same completion point. One sequence may emit an end token after a few decode steps while another remains active for hundreds. If the serving engine keeps the original request batch fixed until every member finishes, completed sequences leave execution capacity stranded behind longer sequences. Iteration-level scheduling moves the scheduling boundary inward. Instead of treating an entire request as the indivisible scheduling unit, the engine returns to the scheduler after a model iteration. Finished sequences can leave, waiting work can enter, and the next iteration can run with a different set of active sequences.

Artificial Intelligence 24 Sep 2026 6 min read

Grouped-Query Attention Shrinks the KV Cache by Sharing Key-Value Heads

Grouped-query attention changes the number of key and value heads that must be stored during autoregressive decoding. Instead of giving every query head its own key-value pair, several query heads address the same key-value head. The query side can retain many heads while the persistent KV state uses fewer independent projections. This is an architectural change to attention, not a cache compression codec. The smaller cache follows from producing fewer distinct key and value heads per token.

Artificial Intelligence 24 Sep 2026 6 min read

Grouped-Query Attention Shares Key-Value Heads Across Query Groups

Multi-head attention does not require every query head to own a distinct key head and value head. Grouped-query attention, usually abbreviated GQA, partitions query heads into groups and assigns one key-value head to each group. Query projections remain separate, but several query heads read from the same projected keys and values. That distinction changes parameter shapes and cached state without collapsing the query heads into one attention computation. Each query head still produces its own attention scores because its query vector is different.

Artificial Intelligence 24 Sep 2026 4 min read

Global Gradient Norm Clipping Rescales Parameter Gradients with One Shared Factor

Global gradient norm clipping treats the current parameter gradients as one aggregate vector. If its norm exceeds a threshold, every participating gradient is multiplied by the same coefficient. The operation changes update magnitude while preserving the direction of the aggregate gradient vector, apart from finite-precision effects. One norm controls the whole gradient set Let the parameter gradients be g_1, g_2, ... g_m. For a p-norm, global clipping first forms the equivalent norm of their concatenation:

Artificial Intelligence 24 Sep 2026 5 min read

FlashAttention Tiles Exact Softmax Without Materializing the Score Matrix

Standard scaled dot-product attention forms a score matrix whose two long axes are sequence positions. For one attention head, S = QK^T / sqrt(d) P = softmax(S) O = PV the mathematical definition is compact, but a direct GPU implementation can write the large intermediate matrices S and P to high-bandwidth memory before reading them again. FlashAttention changes that dataflow. It processes blocks of queries, keys, and values in on-chip memory and carries enough row-wise softmax state to combine score tiles exactly.

Artificial Intelligence 24 Sep 2026 5 min read

Expert Parallelism Turns MoE Routing into All-to-All Communication

A sparsely activated Mixture-of-Experts layer can keep many expert parameters while sending each token to only a small subset of them. Once those experts are partitioned across devices, sparse activation does not imply local execution. A token assigned to a remote expert has to move to the device that owns that expert, and its expert output has to return to the device that continues the model computation. This movement is the defining systems cost of expert parallelism. The router makes a logical selection, but a distributed runtime must turn that selection into communication, local expert batches, expert computation, and a reverse exchange.

Artificial Intelligence 24 Sep 2026 6 min read

Expert Capacity Turns Uneven MoE Routing into Token Overflow

A sparse Mixture-of-Experts layer can route many tokens toward the same expert even when every expert has identical nominal capacity. The router makes token-dependent choices, while distributed execution commonly allocates bounded token slots per expert. When those two mechanisms disagree, an expert can receive more assignments than its execution buffer admits. That boundary is not an inherent property of every MoE architecture. It is a property of capacity-constrained routing designs, including the routing formulation described for Switch Transformers. In such systems, expert capacity converts an uneven routing distribution into an operational event: some assignments fit, while assignments beyond the capacity limit require an explicit overflow policy.

Artificial Intelligence 24 Sep 2026 5 min read

Embedding Anisotropy Compresses Cosine Similarity Ranges

Two embedding candidates can differ in semantic fit yet receive cosine scores packed into a narrow interval. The similarity function may be implemented correctly. The compression can instead come from the geometry of the embedding space: vectors may occupy preferred directions rather than spreading evenly across the available dimensions. This directional concentration is commonly described as anisotropy. Anisotropy matters to retrieval because cosine similarity measures angular alignment. When many vectors share a substantial common component, unrelated pairs can start from an elevated baseline alignment. Relevant pairs may still score higher, but the usable gap between relevant and irrelevant candidates can shrink.

Artificial Intelligence 24 Sep 2026 5 min read

Contrastive Search Penalizes Hidden-State Repetition During Decoding

Autoregressive generation exposes a distribution over the next token, but a decoder still has to choose which candidate to append. Contrastive search changes that choice by combining model probability with a penalty for candidates whose resulting hidden state is too similar to hidden states already present in the generated prefix. The mechanism operates only at inference time. It does not alter model parameters or the next-token distribution itself. Instead, it changes the ranking used to select a token from a restricted candidate set.

Artificial Intelligence 24 Sep 2026 5 min read

Continuous Batching Replaces Static Request Groups with Iteration-Level Scheduling

Autoregressive generation does not finish every request at the same iteration. One sequence may emit a stop token after a few decode steps while another remains active for hundreds more. A static batch keeps those requests coupled until the batch boundary. Continuous batching breaks that coupling by letting the active set change between model iterations. The key change is scheduling granularity. A request is no longer the indivisible scheduling unit for the entire generation. The serving system can construct an execution batch for one iteration, update request state after that iteration, remove finished sequences, and admit queued work before the next execution step.

Artificial Intelligence 24 Sep 2026 5 min read

Codebook Collapse Leaves Vector Quantizers with Unused Entries

A vector quantizer can expose thousands of codebook entries while repeatedly selecting only a small subset. The configured vocabulary then overstates the discrete capacity that is actually used. This state is often called codebook collapse: some entries receive frequent assignments, while others become inactive or nearly inactive. The failure is not a shortage of parameters by itself. It comes from the interaction between encoder outputs, the assignment rule, and the procedure used to update code vectors. Once assignments concentrate around a subset of entries, unused entries may receive little or no signal that would move them toward regions occupied by encoder outputs.

Artificial Intelligence 24 Sep 2026 6 min read

Beam Search Length Normalization Changes Which Sequence Wins

Beam search can return a short sequence even when a longer continuation contains locally plausible tokens. The effect comes from its scoring rule: token log-probabilities are usually accumulated across the sequence, and each additional token contributes another non-positive term. Length normalization changes that ranking pressure, but it does not change the model probabilities that produced the tokens. That distinction matters in generation systems. A decoding score is an inference-time objective used to compare hypotheses. It is not automatically a calibrated probability for completed outputs, and altering it can change the selected sequence without changing a single model parameter.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Stabilize Windowed KV-Cache Streaming

A decoder can keep its KV cache bounded by discarding states that fall outside a fixed recent-token window. The memory bound is attractive, but naive eviction can sharply degrade generation even when the removed tokens carry little obvious semantic value. A small set of key-value states from the beginning of the sequence can change that behavior. This effect is associated with attention sinks: initial positions that receive substantial attention mass even when their token content is not important to the current prediction. StreamingLLM reported that preserving those initial states together with a rolling window of recent states recovers stable language-model behavior that ordinary window eviction can lose.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Stabilize Streaming Cache Eviction

A bounded KV cache can keep inference memory from growing with an unending token stream, but deleting every old token in strict arrival order can disrupt attention more than the missing content alone suggests. In some autoregressive transformers, early positions receive substantial attention even when their token semantics are not directly relevant to the current prediction. These positions act as attention sinks. The serving consequence is specific: a sliding cache that preserves a small initial prefix alongside the most recent tokens can behave differently from a same-sized cache containing only the newest tokens. The mechanism concerns attention state and positional handling, not retrieval of forgotten text.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Absorb Probability Mass Without Carrying Matching Content

A causal attention head can assign noticeable probability mass to an early token even when that token is not a strong semantic match for the current query. The token then behaves as an attention sink: its key remains a convenient destination for probability that the head does not direct toward content-bearing positions. This behavior matters during autoregressive serving because a KV-cache policy can preserve recent tokens yet still alter model behavior sharply if it removes sink positions. Recency alone does not describe the functional role of every cached key.

Artificial Intelligence 24 Sep 2026 5 min read

ALiBi Adds Distance-Proportional Bias to Attention Scores

ALiBi changes an attention score before softmax rather than adding a positional vector to the token representation. For a causal transformer, a key farther behind the current query receives a larger negative offset. The offset is linear in token distance and uses a slope associated with the attention head. A simplified score for head h can be written as: score_h(i, j) = q_i k_j^T / sqrt(d) - m_h * (i - j) for an allowed causal pair with j <= i and positive slope m_h. The causal mask still decides which future positions are inaccessible. ALiBi changes the relative scores among positions that remain eligible.

Artificial Intelligence 24 Sep 2026 5 min read

ALiBi Adds Distance Penalties Directly to Attention Logits

Attention with Linear Biases (ALiBi) changes a causal attention score before softmax by adding a penalty whose magnitude grows with token distance. Position is therefore represented in the score path rather than by adding a positional vector to each token representation. For a query at position i attending to a key at position j, one head can be written schematically as: score(i, j) = q_i · k_j / sqrt(d_k) - m_h * (i - j) for causal positions j <= i. The positive slope m_h is specific to attention head h. The causal mask still prevents access to future positions; the linear term changes the relative preference among positions that remain visible.

Artificial Intelligence 24 Sep 2026 5 min read

Activation Outliers Distort Low-Bit Quantization Scales

A tensor can be easy to represent in floating point and awkward to map into a small integer range. The problem becomes acute when most activation values occupy a narrow interval while a small number have much larger magnitude. A shared quantization scale must cover those extremes, so the ordinary values receive fewer representable levels. That is the practical effect of activation outliers in low-bit inference. The issue is not merely that an outlier is numerically large. Its location, persistence, and relationship to the axis over which a scale is shared determine whether it materially degrades the quantized representation.

Artificial Intelligence 24 Sep 2026 4 min read

Activation Checkpointing Trades Saved Activations for Recomputation

Activation checkpointing changes which forward-pass tensors remain resident until backpropagation. Instead of retaining every intermediate activation required by gradient computation, a checkpointed region keeps selected boundary state and reconstructs discarded intermediates when the backward pass reaches that region. The mechanism reduces activation memory at the cost of extra computation. It does not shrink model parameters, optimizer state, or gradients, so its effect on total training memory depends on how much of the footprint comes from activations.

Artificial Intelligence 23 Sep 2026 5 min read

Temperature Scaling Changes Softmax Sharpness Without Changing Logit Order

A decoder can assign the same ranking to every token before and after temperature scaling while producing substantially different probabilities. The mechanism is simple: for positive temperature T, logits are divided by T before softmax. Division by the same positive scalar preserves order, but softmax converts the changed gaps between logits into a different probability distribution. That distinction matters in inference systems because temperature does not select tokens by itself. Its visible effect depends on what happens after the scaled softmax: direct sampling, top-k filtering, top-p filtering, greedy selection, or another decoding rule.

Artificial Intelligence 23 Sep 2026 4 min read

Temperature Scaling Changes Sampling Without Changing Logit Order

Temperature is often exposed as a single generation parameter, but its effect is narrower than a general control for output quality. For a fixed vector of finite logits, positive temperature rescales the gaps before softmax. It changes the resulting probabilities without changing which logit is larger than another. That distinction matters when a serving layer combines temperature with greedy selection, top-k filtering, top-p filtering, penalties, or implementation-specific handling of zero temperature. The same numeric setting can participate in a different decoding pipeline even though the underlying scaling operation is simple.