Skip to content

Archive

Inference

87 articles
Artificial Intelligence 24 Sep 2026 6 min read

Tanh Logit Soft Capping Bounds Extreme Scores Before Softmax

A softmax can accept logits of any finite magnitude, but large score gaps make its output increasingly concentrated. Tanh logit soft capping inserts a bounded nonlinear transform before softmax so that no transformed logit exceeds a configured magnitude. For a positive cap c, a common form is: softcap(z; c) = c * tanh(z / c) The operation does not clip at a hard threshold. It behaves almost linearly near zero and gradually compresses larger magnitudes as they approach -c or c.

Artificial Intelligence 24 Sep 2026 5 min read

Speculative Decoding Verifies Draft Tokens in Parallel

Autoregressive generation normally commits one token before the next target-model step can be evaluated. Speculative decoding changes that execution pattern. A cheaper proposal distribution produces a short candidate block, while the target model evaluates the proposed positions together. An acceptance rule then determines how much of that block can become output. The mechanism is not ordinary batching. Tokens inside the proposal remain autoregressive, and the target model still defines the intended output distribution when an exact speculative-sampling algorithm is used. The useful change is that one expensive verification pass can account for several output positions when enough proposals are accepted.

Artificial Intelligence 24 Sep 2026 6 min read

KV Cache Quantization Trades Precision for Serving Memory

Autoregressive decoding keeps past attention keys and values so each new token can reuse earlier projections instead of recomputing them. That KV cache grows with sequence length, layer count, batch size, and the number and width of cached key/value heads. Reducing its numeric precision can cut the bytes occupied by those cached tensors, but it also changes the values consumed by later attention operations. This makes KV cache quantization different from compressing data that is only stored and restored losslessly. The quantized cache remains on the inference path. Each subsequent query can interact with approximated keys and values, so the relevant question is not only how many bytes are saved, but where quantization error enters attention and how the serving implementation contains it.

Artificial Intelligence 24 Sep 2026 5 min read

Iteration-Level Scheduling Rebuilds LLM Batches Between Decode Steps

Autoregressive serving does not give every request the same completion point. One sequence may emit an end token after a few decode steps while another remains active for hundreds. If the serving engine keeps the original request batch fixed until every member finishes, completed sequences leave execution capacity stranded behind longer sequences. Iteration-level scheduling moves the scheduling boundary inward. Instead of treating an entire request as the indivisible scheduling unit, the engine returns to the scheduler after a model iteration. Finished sequences can leave, waiting work can enter, and the next iteration can run with a different set of active sequences.

Artificial Intelligence 24 Sep 2026 6 min read

Beam Search Length Normalization Changes Which Sequence Wins

Beam search can return a short sequence even when a longer continuation contains locally plausible tokens. The effect comes from its scoring rule: token log-probabilities are usually accumulated across the sequence, and each additional token contributes another non-positive term. Length normalization changes that ranking pressure, but it does not change the model probabilities that produced the tokens. That distinction matters in generation systems. A decoding score is an inference-time objective used to compare hypotheses. It is not automatically a calibrated probability for completed outputs, and altering it can change the selected sequence without changing a single model parameter.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Stabilize Windowed KV-Cache Streaming

A decoder can keep its KV cache bounded by discarding states that fall outside a fixed recent-token window. The memory bound is attractive, but naive eviction can sharply degrade generation even when the removed tokens carry little obvious semantic value. A small set of key-value states from the beginning of the sequence can change that behavior. This effect is associated with attention sinks: initial positions that receive substantial attention mass even when their token content is not important to the current prediction. StreamingLLM reported that preserving those initial states together with a rolling window of recent states recovers stable language-model behavior that ordinary window eviction can lose.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Absorb Probability Mass Without Carrying Matching Content

A causal attention head can assign noticeable probability mass to an early token even when that token is not a strong semantic match for the current query. The token then behaves as an attention sink: its key remains a convenient destination for probability that the head does not direct toward content-bearing positions. This behavior matters during autoregressive serving because a KV-cache policy can preserve recent tokens yet still alter model behavior sharply if it removes sink positions. Recency alone does not describe the functional role of every cached key.

Artificial Intelligence 23 Sep 2026 5 min read

Temperature Scaling Changes Softmax Sharpness Without Changing Logit Order

A decoder can assign the same ranking to every token before and after temperature scaling while producing substantially different probabilities. The mechanism is simple: for positive temperature T, logits are divided by T before softmax. Division by the same positive scalar preserves order, but softmax converts the changed gaps between logits into a different probability distribution. That distinction matters in inference systems because temperature does not select tokens by itself. Its visible effect depends on what happens after the scaled softmax: direct sampling, top-k filtering, top-p filtering, greedy selection, or another decoding rule.

Artificial Intelligence 23 Sep 2026 3 min read

Stable Softmax Subtracts the Maximum Logit Before Exponentiation

A decoder can receive logits large enough that direct exponentiation is numerically unsafe even though the intended probability distribution is ordinary. Softmax does not require exponentiating the original values. Subtracting the largest logit from every logit produces the same distribution in exact real arithmetic while moving the exponentials into a safer numeric range. This shift is a property of softmax itself, not a model-specific heuristic. A shared shift cancels during normalization For logits (z_1,\ldots,z_n), softmax assigns

Artificial Intelligence 23 Sep 2026 5 min read

Speculative Sampling Preserves Target Distribution Through Rejection Correction

A draft model can propose a token that the target model would not have sampled from the same random draw, yet speculative sampling can still preserve the target model’s distribution. The key is not that the draft model predicts the target perfectly. Distributional correctness comes from the acceptance rule and the correction applied after rejection. This separates two properties that are often grouped together. Draft quality controls how frequently proposals survive verification. The rejection-correction construction controls whether the resulting sample follows the target distribution.

Artificial Intelligence 23 Sep 2026 5 min read

Speculative Decoding Verifies Draft Tokens Before Acceptance

Autoregressive generation normally advances one token at a time: the model scores the next-token distribution, a decoding rule selects a token, and that token becomes part of the context for the next forward step. Speculative decoding changes the execution schedule. A cheaper draft process proposes several future tokens, while the target model retains authority over which proposals can enter the generated sequence. That separation is the core constraint. Draft tokens are predictions about future target-model decisions, not a replacement distribution that can be appended unchecked. The serving implementation must verify them against target-model scores and apply the acceptance rule required by the chosen speculative algorithm.

Artificial Intelligence 23 Sep 2026 5 min read

Sliding-Window KV Caches Change Which Tokens Remain Addressable

Autoregressive decoding can reuse key and value tensors from earlier token positions instead of recomputing them at every step. A conventional KV cache therefore grows as generation advances. A sliding-window cache places a bound on that retained state by keeping only a recent region. The memory bound changes more than allocation size. Once an old key-value entry is evicted, a later attention operation cannot directly address that position through the cache. The resulting behavior depends on the model’s attention pattern, positional scheme, cache implementation, and any layers that use a different attention span.

Artificial Intelligence 23 Sep 2026 5 min read

RoPE Position Shifts Preserve Relative Phase but Change Absolute Rotation

Rotary position embeddings, commonly called RoPE, inject position into attention by rotating pairs of query and key channels with angles determined by token position. The rotation happens before the query-key dot product. As a result, position is not represented by adding a standalone vector to each token representation. This distinction matters during inference. A token’s cached key already contains the rotation associated with its assigned position. Reusing that key under a different sequence coordinate without an equivalent transformation changes the attention computation, even when the token content is identical.

Artificial Intelligence 23 Sep 2026 5 min read

Prefix KV Cache Reuse Depends on Exact Context Identity

A prefix cache can remove repeated prefill work without changing the model output, but only when the cached key and value states represent the same prefix context that the new request would have produced. Matching visible text is not enough. The serving path ultimately operates on token IDs, positions, model parameters, and implementation-specific attention state. This boundary makes prefix caching different from a generic text cache. A text cache stores a result associated with an input. A KV prefix cache stores intermediate states whose validity depends on the computation that created them.

Artificial Intelligence 23 Sep 2026 5 min read

PagedAttention Maps Logical KV Blocks to Noncontiguous Physical Memory

An autoregressive request grows its KV cache as tokens arrive, but its final sequence length is not known when decoding begins. Reserving one contiguous region for the maximum possible sequence length ties memory to capacity that may never be used. PagedAttention changes that allocation boundary: a sequence is represented as logical KV blocks, while a block table maps those logical blocks to physical blocks that need not be adjacent in GPU memory.

Artificial Intelligence 23 Sep 2026 4 min read

Grouped-Query Attention Shares Key-Value Heads Across Query Heads

Grouped-query attention changes a specific part of multi-head attention: several query heads use the same key and value head. The query projections remain separate, so those query heads can produce different attention weights, but they read keys and values from a shared projected representation. That distinction matters during inference. A decoder cache stores past keys and values, not past queries. Reducing the number of key-value heads can therefore reduce KV cache storage without reducing the number of query heads by the same factor.

Artificial Intelligence 23 Sep 2026 6 min read

Beam Search Length Normalization Changes Sequence Ranking

Beam search ranks partial sequences by scores accumulated across decoding steps. When that score is the sum of token log-probabilities, every additional token contributes a value that is normally non-positive. A longer candidate therefore has more opportunities to reduce its raw score, even when its continuation is locally plausible. That property is not a defect in probability theory. It follows from comparing complete sequences with different numbers of conditional factors. It becomes an implementation concern when a decoder is expected to produce useful completions rather than simply rank sequences by unmodified model probability.

Artificial Intelligence 23 Sep 2026 6 min read

Attention Logit Scaling Keeps Dot Products in a Stable Softmax Range

A dot product between a query and a key tends to grow in magnitude as their dimension grows. In scaled dot-product attention, the score is divided by the square root of the key dimension before softmax. That factor is not a cosmetic normalization. It controls the scale presented to softmax under a specific statistical assumption about the query and key components. The familiar expression is Attention(Q, K, V) = softmax(QK^T / sqrt(d_k)) V where d_k is the query-key dimension for one attention head. The scaling term affects the distribution of attention probabilities even though it does not change the ordering of logits by itself.

Artificial Intelligence 23 Sep 2026 6 min read

Additive Logit Bias Changes Token Odds Before Sampling

A decoder can favor or suppress a token without changing model weights. Add a constant to that token’s logit before softmax, and its probability changes relative to the rest of the vocabulary. The operation is simple, but its effect depends on where the bias enters the decoding pipeline and on every transformation that follows it. This makes additive logit bias useful as an inference control, but not as a general semantic constraint. It changes a score used by the decoder. It does not rewrite the model’s internal representation or guarantee that a concept disappears from generated text.

Artificial Intelligence 23 Sep 2026 5 min read

Activation Outliers Can Set the Scale for an Entire Quantization Group

A quantizer with a fixed integer width has only a finite set of representable codes. When many activations share one scale, a single value with much larger magnitude can force that scale to cover a wider real-valued range. The remaining values then occupy fewer useful code intervals around the region where they are concentrated. This behavior is not a generic statement that quantization fails in the presence of large numbers. It follows from a specific coupling: values inside the same quantization group share parameters that map real numbers to integer codes.

Artificial Intelligence 23 Sep 2026 5 min read

Activation Outliers Can Dominate Per-Tensor Quantization Scale

A per-tensor quantizer maps every value in an activation tensor through one shared scale. That coupling matters when most activations occupy a narrow interval but a few values have much larger magnitude. The large values can determine the scale, while the dense central region is represented with coarser spacing than its own range would require. This is not a statement that every large activation is erroneous or removable. An outlier may carry useful model state. The issue is numerical: one scale has to cover values with very different magnitudes.

Artificial Intelligence 22 Sep 2026 5 min read

Top-p Sampling Rebuilds Its Candidate Set at Every Token

Top-p sampling does not keep a fixed shortlist of tokens throughout generation. At each decoding step, the model produces a new logit vector, that vector becomes a probability distribution, and the sampler forms a new candidate set whose cumulative probability mass reaches the configured threshold. The consequence is easy to miss in serving code: the same top_p value can admit two tokens at one step and dozens at another. The parameter controls probability mass, not candidate count.

Artificial Intelligence 22 Sep 2026 6 min read

Speculative Decoding Couples Draft Speed with Acceptance Rate

Autoregressive generation normally advances one accepted token at a time. Each new token extends the prefix, so the next target-model evaluation depends on the token selected at the preceding position. Speculative decoding changes the execution schedule: a cheaper draft process proposes several future tokens, then the target model evaluates those positions together and decides how much of the proposal can be retained. That rearrangement can reduce the number of serial target-model calls per emitted token. It does not make verification free, and a longer draft block is not automatically better. The useful operating point depends on how quickly proposals are produced, how often they survive target verification, and what the serving stack spends on rejected work.

Artificial Intelligence 22 Sep 2026 7 min read

Prefix KV Cache Reuse Depends on Exact Token History

A KV cache entry is not a reusable representation of arbitrary text that happens to look similar. For an autoregressive transformer, cached keys and values are intermediate states produced for a specific token prefix under a specific execution context. Reusing them is valid only when the new request reaches the same state boundary. That boundary is stricter than matching visible characters. Tokenization, token order, position handling, model identity, adapter state, and other inputs that affect hidden states can all determine whether a cached prefix still represents the computation required by the new request.