Skip to content

Archive

LLM Inference

13 articles
Artificial Intelligence 24 Sep 2026 6 min read

Speculative Decoding Verifies Draft Tokens Without Changing the Target Sampling Distribution

Autoregressive decoding normally commits one token after each target-model evaluation. A sequence of K generated tokens therefore creates a serial dependency chain: token t+1 cannot be sampled until token t is fixed and becomes part of the prefix. Speculative decoding changes the amount of useful work obtained from a target-model call without removing that causal dependency. A cheaper draft model first proposes several continuation tokens. The target model then evaluates those candidate positions together. Tokens whose draft probabilities are compatible with the target distribution can be accepted, while the first rejected position is corrected using a residual distribution. The resulting samples follow the target model’s distribution when the acceptance and correction procedure is implemented as specified.

Artificial Intelligence 24 Sep 2026 4 min read

Speculative Decoding Verifies Draft Tokens Without Changing the Target Distribution

Autoregressive generation normally invokes the target model once for every emitted token. Speculative decoding changes that execution pattern: a cheaper draft model proposes several tokens, and the target model evaluates the proposed block in a single verification pass. The speed opportunity comes from doing useful target-model work for multiple positions at once, not from treating draft output as authoritative. Draft tokens are proposals, not final output Let the target model define distribution p and the draft model define distribution q at a given position. The draft samples a candidate token from q. Verification then decides whether that candidate can be retained as a sample consistent with p.

Artificial Intelligence 24 Sep 2026 5 min read

Sliding-Window Attention Bounds Direct Token Access

A causal transformer does not have to expose every earlier token to every query position. With sliding-window attention, position i can attend only to a bounded range of preceding keys. Tokens outside that range are absent from that layer’s attention operation, even if they still belong to the model input. That boundary changes more than the attention matrix shape. It separates direct access from information that can reach a position only after being propagated through intermediate hidden states.

Artificial Intelligence 24 Sep 2026 5 min read

Prefix KV Caching Reuses Only the Shared Token Prefix

A prefix cache hit ends at the first point where a new request can no longer reuse previously computed state. The reusable object is not a piece of source text in isolation. It is model state produced for an ordered token prefix under execution conditions that make that state compatible with the new request. For transformer inference, that state is commonly the key-value cache created during prefill. Reusing it can remove repeated computation for the shared prefix while leaving the divergent suffix to be processed normally.

Artificial Intelligence 24 Sep 2026 5 min read

PagedAttention Decouples Logical KV Sequences from Physical Cache Blocks

Autoregressive serving keeps a growing key-value state for every active sequence. If that state must occupy one contiguous physical region sized for a request, allocation becomes coupled to uncertain sequence length: reserving too much wastes capacity, while extending or relocating a growing region complicates memory management. PagedAttention changes that allocation boundary. A sequence is represented as logical KV blocks, while its physical blocks may reside at unrelated locations in the cache pool.

Artificial Intelligence 24 Sep 2026 5 min read

Multi-Head Latent Attention Compresses KV State Before Head-Specific Expansion

Multi-Head Latent Attention changes the state retained across autoregressive decoding. Instead of requiring a full key tensor and value tensor for every cached token and attention head, the architecture can retain a lower-dimensional latent representation and derive head-specific key/value information through projection structure. That distinction matters at the cache boundary. Ordinary multi-head attention commonly stores already projected per-head keys and values. MLA moves part of that representation behind a compact bottleneck, so persistent state and compute no longer have the same shape.

Artificial Intelligence 24 Sep 2026 4 min read

KV-Cache Quantization Reduces Cache Bytes with Separate Accuracy and Kernel Costs

Autoregressive decoding retains key and value tensors from earlier tokens, so KV-cache memory grows with retained context. Quantizing those tensors changes a direct term in that memory footprint: fewer bits are stored for each cached element. The trade is not free capacity. Reduced precision adds representation error and requires a concrete scaling, storage, and kernel strategy. Cache precision is separate from weight precision Model weights and KV state have different lifetimes. Weights are persistent across requests, while KV tensors are generated from each request and grow as its sequence advances. A model can therefore use one numerical format for weights and another for its cache.

Artificial Intelligence 24 Sep 2026 6 min read

KV Cache Quantization Reduces Stored Attention State at a Reconstruction Cost

Autoregressive decoding appends key and value tensors to a cache at every transformer layer. The cache prevents prior tokens from being projected into keys and values again, but its storage grows with sequence length. KV cache quantization changes that storage representation: older or selected cache entries are encoded with fewer bits, then reconstructed when attention consumes them. The mechanism is a memory-format trade. It does not remove tokens from the attention context and it does not change the model weights. It reduces bytes used by cached state while introducing quantization error, scale or zero-point metadata, and conversion work on the decode path.

Artificial Intelligence 24 Sep 2026 4 min read

KV Cache Quantization Perturbs Attention Through Stored Keys and Values

During autoregressive inference, previously computed keys and values are reused from the KV cache instead of being recomputed for every new token. Quantizing that cache changes more than its byte representation. The stored approximation becomes an input to later attention operations, so its error can alter both attention scores and the vectors combined by those scores. This boundary differs from quantizing model weights. A weight tensor is reused across requests, while KV state is generated from the current sequence and grows with its cached length. Its numeric range can also vary across layers, heads, positions, and requests.

Artificial Intelligence 24 Sep 2026 6 min read

Grouped-Query Attention Shares Key-Value Heads Across Query Groups

Multi-head attention does not require every query head to own a distinct key head and value head. Grouped-query attention, usually abbreviated GQA, partitions query heads into groups and assigns one key-value head to each group. Query projections remain separate, but several query heads read from the same projected keys and values. That distinction changes parameter shapes and cached state without collapsing the query heads into one attention computation. Each query head still produces its own attention scores because its query vector is different.

Artificial Intelligence 24 Sep 2026 5 min read

Continuous Batching Replaces Static Request Groups with Iteration-Level Scheduling

Autoregressive generation does not finish every request at the same iteration. One sequence may emit a stop token after a few decode steps while another remains active for hundreds more. A static batch keeps those requests coupled until the batch boundary. Continuous batching breaks that coupling by letting the active set change between model iterations. The key change is scheduling granularity. A request is no longer the indivisible scheduling unit for the entire generation. The serving system can construct an execution batch for one iteration, update request state after that iteration, remove finished sequences, and admit queued work before the next execution step.

Artificial Intelligence 17 Sep 2026 6 min read

Retain Attention Sinks in Bounded KV Caches

A bounded KV cache seems to invite a simple eviction rule: once the cache is full, discard the oldest key-value pair and keep the most recent tokens. For some decoder-only transformers, that rule can degrade generation sharply after the sequence moves beyond the retained window. The failure is not explained only by missing old semantic content. Early tokens can receive substantial attention even when their text carries little useful information for the current prediction.

Artificial Intelligence 15 Sep 2026 6 min read

Quantize KV Caches with Explicit Error Budgets

Autoregressive transformer inference retains key and value tensors from earlier tokens so each new token can attend to prior context without recomputing those projections. As context length and concurrent sequence count rise, this KV cache can become a substantial part of accelerator memory. KV cache quantization stores those tensors at reduced precision and reconstructs approximations when attention consumes them. The memory arithmetic is attractive, but the resulting error is not a generic model-weight perturbation. Quantized keys affect attention scores before the softmax, while quantized values affect the weighted sum after attention probabilities have been formed.