Skip to content

Archive

Model Serving

15 articles
Artificial Intelligence 24 Sep 2026 4 min read

Speculative Decoding Verifies Draft Tokens Without Changing the Target Distribution

Autoregressive generation normally invokes the target model once for every emitted token. Speculative decoding changes that execution pattern: a cheaper draft model proposes several tokens, and the target model evaluates the proposed block in a single verification pass. The speed opportunity comes from doing useful target-model work for multiple positions at once, not from treating draft output as authoritative. Draft tokens are proposals, not final output Let the target model define distribution p and the draft model define distribution q at a given position. The draft samples a candidate token from q. Verification then decides whether that candidate can be retained as a sample consistent with p.

Artificial Intelligence 24 Sep 2026 5 min read

Speculative Decoding Verifies Draft Tokens in Parallel

Autoregressive generation normally commits one token before the next target-model step can be evaluated. Speculative decoding changes that execution pattern. A cheaper proposal distribution produces a short candidate block, while the target model evaluates the proposed positions together. An acceptance rule then determines how much of that block can become output. The mechanism is not ordinary batching. Tokens inside the proposal remain autoregressive, and the target model still defines the intended output distribution when an exact speculative-sampling algorithm is used. The useful change is that one expensive verification pass can account for several output positions when enough proposals are accepted.

Artificial Intelligence 24 Sep 2026 5 min read

Prefix Caching Reuses KV State for Identical Token Prefixes

Autoregressive serving often receives requests that start with the same long token sequence. A system prompt, tool schema, fixed document header, or other repeated context can cause the model to compute the same prefix attention state again for each request. Prefix caching targets that repeated prefill work by retaining compatible key-value state and attaching later requests to it. The reuse boundary is exact tokenized state, not semantic similarity. Two prompts that express the same idea but tokenize differently do not produce an interchangeable prefix cache entry.

Artificial Intelligence 24 Sep 2026 6 min read

KV Cache Quantization Trades Precision for Serving Memory

Autoregressive decoding keeps past attention keys and values so each new token can reuse earlier projections instead of recomputing them. That KV cache grows with sequence length, layer count, batch size, and the number and width of cached key/value heads. Reducing its numeric precision can cut the bytes occupied by those cached tensors, but it also changes the values consumed by later attention operations. This makes KV cache quantization different from compressing data that is only stored and restored losslessly. The quantized cache remains on the inference path. Each subsequent query can interact with approximated keys and values, so the relevant question is not only how many bytes are saved, but where quantization error enters attention and how the serving implementation contains it.

Artificial Intelligence 24 Sep 2026 6 min read

KV Cache Quantization Reduces Stored Attention State at a Reconstruction Cost

Autoregressive decoding appends key and value tensors to a cache at every transformer layer. The cache prevents prior tokens from being projected into keys and values again, but its storage grows with sequence length. KV cache quantization changes that storage representation: older or selected cache entries are encoded with fewer bits, then reconstructed when attention consumes them. The mechanism is a memory-format trade. It does not remove tokens from the attention context and it does not change the model weights. It reduces bytes used by cached state while introducing quantization error, scale or zero-point metadata, and conversion work on the decode path.

Artificial Intelligence 24 Sep 2026 6 min read

Grouped-Query Attention Shrinks the KV Cache by Sharing Key-Value Heads

Grouped-query attention changes the number of key and value heads that must be stored during autoregressive decoding. Instead of giving every query head its own key-value pair, several query heads address the same key-value head. The query side can retain many heads while the persistent KV state uses fewer independent projections. This is an architectural change to attention, not a cache compression codec. The smaller cache follows from producing fewer distinct key and value heads per token.

Artificial Intelligence 24 Sep 2026 5 min read

Expert Parallelism Turns MoE Routing into All-to-All Communication

A sparsely activated Mixture-of-Experts layer can keep many expert parameters while sending each token to only a small subset of them. Once those experts are partitioned across devices, sparse activation does not imply local execution. A token assigned to a remote expert has to move to the device that owns that expert, and its expert output has to return to the device that continues the model computation. This movement is the defining systems cost of expert parallelism. The router makes a logical selection, but a distributed runtime must turn that selection into communication, local expert batches, expert computation, and a reverse exchange.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Stabilize Streaming Cache Eviction

A bounded KV cache can keep inference memory from growing with an unending token stream, but deleting every old token in strict arrival order can disrupt attention more than the missing content alone suggests. In some autoregressive transformers, early positions receive substantial attention even when their token semantics are not directly relevant to the current prediction. These positions act as attention sinks. The serving consequence is specific: a sliding cache that preserves a small initial prefix alongside the most recent tokens can behave differently from a same-sized cache containing only the newest tokens. The mechanism concerns attention state and positional handling, not retrieval of forgotten text.

Artificial Intelligence 24 Sep 2026 5 min read

Activation Outliers Distort Low-Bit Quantization Scales

A tensor can be easy to represent in floating point and awkward to map into a small integer range. The problem becomes acute when most activation values occupy a narrow interval while a small number have much larger magnitude. A shared quantization scale must cover those extremes, so the ordinary values receive fewer representable levels. That is the practical effect of activation outliers in low-bit inference. The issue is not merely that an outlier is numerically large. Its location, persistence, and relationship to the axis over which a scale is shared determine whether it materially degrades the quantized representation.

Artificial Intelligence 23 Sep 2026 5 min read

Speculative Decoding Verifies Draft Tokens Before Acceptance

Autoregressive generation normally advances one token at a time: the model scores the next-token distribution, a decoding rule selects a token, and that token becomes part of the context for the next forward step. Speculative decoding changes the execution schedule. A cheaper draft process proposes several future tokens, while the target model retains authority over which proposals can enter the generated sequence. That separation is the core constraint. Draft tokens are predictions about future target-model decisions, not a replacement distribution that can be appended unchecked. The serving implementation must verify them against target-model scores and apply the acceptance rule required by the chosen speculative algorithm.

Artificial Intelligence 23 Sep 2026 5 min read

Prefix KV Cache Reuse Depends on Exact Context Identity

A prefix cache can remove repeated prefill work without changing the model output, but only when the cached key and value states represent the same prefix context that the new request would have produced. Matching visible text is not enough. The serving path ultimately operates on token IDs, positions, model parameters, and implementation-specific attention state. This boundary makes prefix caching different from a generic text cache. A text cache stores a result associated with an input. A KV prefix cache stores intermediate states whose validity depends on the computation that created them.

Artificial Intelligence 23 Sep 2026 4 min read

Attention Sinks Make Naive KV Eviction Unstable

A fixed-size KV cache seems to invite a simple policy: keep the newest entries and evict the oldest. For some decoder transformers, that policy removes positions that later queries continue to assign substantial attention mass. Those early positions act as attention sinks, and dropping them can change the attention distribution far more than their semantic content suggests. The effect matters specifically at inference time when a serving system truncates cached keys and values. It is not a general claim that the first token is semantically privileged, nor that every transformer exhibits the same pattern.

Artificial Intelligence 23 Sep 2026 5 min read

Activation Outliers Can Dominate Per-Tensor Quantization Scale

A per-tensor quantizer maps every value in an activation tensor through one shared scale. That coupling matters when most activations occupy a narrow interval but a few values have much larger magnitude. The large values can determine the scale, while the dense central region is represented with coarser spacing than its own range would require. This is not a statement that every large activation is erroneous or removable. An outlier may carry useful model state. The issue is numerical: one scale has to cover values with very different magnitudes.

Artificial Intelligence 22 Sep 2026 6 min read

Speculative Decoding Couples Draft Speed with Acceptance Rate

Autoregressive generation normally advances one accepted token at a time. Each new token extends the prefix, so the next target-model evaluation depends on the token selected at the preceding position. Speculative decoding changes the execution schedule: a cheaper draft process proposes several future tokens, then the target model evaluates those positions together and decides how much of the proposal can be retained. That rearrangement can reduce the number of serial target-model calls per emitted token. It does not make verification free, and a longer draft block is not automatically better. The useful operating point depends on how quickly proposals are produced, how often they survive target verification, and what the serving stack spends on rejected work.

Artificial Intelligence 22 Sep 2026 7 min read

Prefix KV Cache Reuse Depends on Exact Token History

A KV cache entry is not a reusable representation of arbitrary text that happens to look similar. For an autoregressive transformer, cached keys and values are intermediate states produced for a specific token prefix under a specific execution context. Reusing them is valid only when the new request reaches the same state boundary. That boundary is stricter than matching visible characters. Tokenization, token order, position handling, model identity, adapter state, and other inputs that affect hidden states can all determine whether a cached prefix still represents the computation required by the new request.