Skip to content

Archive

KV Cache

22 articles
Artificial Intelligence 24 Sep 2026 5 min read

Prefix KV Caching Reuses Only the Shared Token Prefix

A prefix cache hit ends at the first point where a new request can no longer reuse previously computed state. The reusable object is not a piece of source text in isolation. It is model state produced for an ordered token prefix under execution conditions that make that state compatible with the new request. For transformer inference, that state is commonly the key-value cache created during prefill. Reusing it can remove repeated computation for the shared prefix while leaving the divergent suffix to be processed normally.

Artificial Intelligence 24 Sep 2026 5 min read

Prefix Caching Reuses KV State for Identical Token Prefixes

Autoregressive serving often receives requests that start with the same long token sequence. A system prompt, tool schema, fixed document header, or other repeated context can cause the model to compute the same prefix attention state again for each request. Prefix caching targets that repeated prefill work by retaining compatible key-value state and attaching later requests to it. The reuse boundary is exact tokenized state, not semantic similarity. Two prompts that express the same idea but tokenize differently do not produce an interchangeable prefix cache entry.

Artificial Intelligence 24 Sep 2026 5 min read

PagedAttention Decouples Logical KV Sequences from Physical Cache Blocks

Autoregressive serving keeps a growing key-value state for every active sequence. If that state must occupy one contiguous physical region sized for a request, allocation becomes coupled to uncertain sequence length: reserving too much wastes capacity, while extending or relocating a growing region complicates memory management. PagedAttention changes that allocation boundary. A sequence is represented as logical KV blocks, while its physical blocks may reside at unrelated locations in the cache pool.

Artificial Intelligence 24 Sep 2026 5 min read

Multi-Head Latent Attention Compresses KV State Before Head-Specific Expansion

Multi-Head Latent Attention changes the state retained across autoregressive decoding. Instead of requiring a full key tensor and value tensor for every cached token and attention head, the architecture can retain a lower-dimensional latent representation and derive head-specific key/value information through projection structure. That distinction matters at the cache boundary. Ordinary multi-head attention commonly stores already projected per-head keys and values. MLA moves part of that representation behind a compact bottleneck, so persistent state and compute no longer have the same shape.

Artificial Intelligence 24 Sep 2026 4 min read

KV-Cache Quantization Reduces Cache Bytes with Separate Accuracy and Kernel Costs

Autoregressive decoding retains key and value tensors from earlier tokens, so KV-cache memory grows with retained context. Quantizing those tensors changes a direct term in that memory footprint: fewer bits are stored for each cached element. The trade is not free capacity. Reduced precision adds representation error and requires a concrete scaling, storage, and kernel strategy. Cache precision is separate from weight precision Model weights and KV state have different lifetimes. Weights are persistent across requests, while KV tensors are generated from each request and grow as its sequence advances. A model can therefore use one numerical format for weights and another for its cache.

Artificial Intelligence 24 Sep 2026 6 min read

KV Cache Quantization Reduces Stored Attention State at a Reconstruction Cost

Autoregressive decoding appends key and value tensors to a cache at every transformer layer. The cache prevents prior tokens from being projected into keys and values again, but its storage grows with sequence length. KV cache quantization changes that storage representation: older or selected cache entries are encoded with fewer bits, then reconstructed when attention consumes them. The mechanism is a memory-format trade. It does not remove tokens from the attention context and it does not change the model weights. It reduces bytes used by cached state while introducing quantization error, scale or zero-point metadata, and conversion work on the decode path.

Artificial Intelligence 24 Sep 2026 4 min read

KV Cache Quantization Perturbs Attention Through Stored Keys and Values

During autoregressive inference, previously computed keys and values are reused from the KV cache instead of being recomputed for every new token. Quantizing that cache changes more than its byte representation. The stored approximation becomes an input to later attention operations, so its error can alter both attention scores and the vectors combined by those scores. This boundary differs from quantizing model weights. A weight tensor is reused across requests, while KV state is generated from the current sequence and grows with its cached length. Its numeric range can also vary across layers, heads, positions, and requests.

Artificial Intelligence 24 Sep 2026 6 min read

Grouped-Query Attention Shrinks the KV Cache by Sharing Key-Value Heads

Grouped-query attention changes the number of key and value heads that must be stored during autoregressive decoding. Instead of giving every query head its own key-value pair, several query heads address the same key-value head. The query side can retain many heads while the persistent KV state uses fewer independent projections. This is an architectural change to attention, not a cache compression codec. The smaller cache follows from producing fewer distinct key and value heads per token.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Stabilize Windowed KV-Cache Streaming

A decoder can keep its KV cache bounded by discarding states that fall outside a fixed recent-token window. The memory bound is attractive, but naive eviction can sharply degrade generation even when the removed tokens carry little obvious semantic value. A small set of key-value states from the beginning of the sequence can change that behavior. This effect is associated with attention sinks: initial positions that receive substantial attention mass even when their token content is not important to the current prediction. StreamingLLM reported that preserving those initial states together with a rolling window of recent states recovers stable language-model behavior that ordinary window eviction can lose.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Stabilize Streaming Cache Eviction

A bounded KV cache can keep inference memory from growing with an unending token stream, but deleting every old token in strict arrival order can disrupt attention more than the missing content alone suggests. In some autoregressive transformers, early positions receive substantial attention even when their token semantics are not directly relevant to the current prediction. These positions act as attention sinks. The serving consequence is specific: a sliding cache that preserves a small initial prefix alongside the most recent tokens can behave differently from a same-sized cache containing only the newest tokens. The mechanism concerns attention state and positional handling, not retrieval of forgotten text.

Artificial Intelligence 24 Sep 2026 5 min read

Attention Sinks Absorb Probability Mass Without Carrying Matching Content

A causal attention head can assign noticeable probability mass to an early token even when that token is not a strong semantic match for the current query. The token then behaves as an attention sink: its key remains a convenient destination for probability that the head does not direct toward content-bearing positions. This behavior matters during autoregressive serving because a KV-cache policy can preserve recent tokens yet still alter model behavior sharply if it removes sink positions. Recency alone does not describe the functional role of every cached key.

Artificial Intelligence 23 Sep 2026 5 min read

Sliding-Window KV Caches Change Which Tokens Remain Addressable

Autoregressive decoding can reuse key and value tensors from earlier token positions instead of recomputing them at every step. A conventional KV cache therefore grows as generation advances. A sliding-window cache places a bound on that retained state by keeping only a recent region. The memory bound changes more than allocation size. Once an old key-value entry is evicted, a later attention operation cannot directly address that position through the cache. The resulting behavior depends on the model’s attention pattern, positional scheme, cache implementation, and any layers that use a different attention span.

Artificial Intelligence 23 Sep 2026 6 min read

Sliding-Window Attention Bounds Active KV State by Token Distance

Sliding-window attention imposes a finite token-distance boundary on causal attention. At position (t), a query can address only a recent interval of key and value positions rather than the full prefix. Once a cached position falls permanently outside that interval, later queries governed by the same local rule cannot address it. That boundary changes the state required for autoregressive decoding. Full causal attention keeps usable KV state growing with sequence length. A fixed local window can keep the active KV span bounded, provided the runtime evicts or overwrites entries that have become unreachable.

Artificial Intelligence 23 Sep 2026 4 min read

RoPE KV Caches Preserve Token Position Across Incremental Decoding

In a decoder using rotary position embeddings, a key written to the KV cache already carries the rotation associated with its token position. Incremental decoding can reuse that key directly. Applying the current token position to the cached key again changes the attention geometry. This makes position state part of the cache contract even when the cache API appears to store only tensors. Rotation is applied before a key enters attention For one two-dimensional component pair, RoPE applies a position-dependent rotation. Writing (R_m) for the rotation at position (m), a query and key become

Artificial Intelligence 23 Sep 2026 5 min read

Prefix KV Cache Reuse Depends on Exact Context Identity

A prefix cache can remove repeated prefill work without changing the model output, but only when the cached key and value states represent the same prefix context that the new request would have produced. Matching visible text is not enough. The serving path ultimately operates on token IDs, positions, model parameters, and implementation-specific attention state. This boundary makes prefix caching different from a generic text cache. A text cache stores a result associated with an input. A KV prefix cache stores intermediate states whose validity depends on the computation that created them.

Artificial Intelligence 23 Sep 2026 5 min read

PagedAttention Maps Logical KV Blocks to Noncontiguous Physical Memory

An autoregressive request grows its KV cache as tokens arrive, but its final sequence length is not known when decoding begins. Reserving one contiguous region for the maximum possible sequence length ties memory to capacity that may never be used. PagedAttention changes that allocation boundary: a sequence is represented as logical KV blocks, while a block table maps those logical blocks to physical blocks that need not be adjacent in GPU memory.

Artificial Intelligence 23 Sep 2026 4 min read

Attention Sinks Make Naive KV Eviction Unstable

A fixed-size KV cache seems to invite a simple policy: keep the newest entries and evict the oldest. For some decoder transformers, that policy removes positions that later queries continue to assign substantial attention mass. Those early positions act as attention sinks, and dropping them can change the attention distribution far more than their semantic content suggests. The effect matters specifically at inference time when a serving system truncates cached keys and values. It is not a general claim that the first token is semantically privileged, nor that every transformer exhibits the same pattern.

Artificial Intelligence 22 Sep 2026 7 min read

Prefix KV Cache Reuse Depends on Exact Token History

A KV cache entry is not a reusable representation of arbitrary text that happens to look similar. For an autoregressive transformer, cached keys and values are intermediate states produced for a specific token prefix under a specific execution context. Reusing them is valid only when the new request reaches the same state boundary. That boundary is stricter than matching visible characters. Tokenization, token order, position handling, model identity, adapter state, and other inputs that affect hidden states can all determine whether a cached prefix still represents the computation required by the new request.

Artificial Intelligence 19 Sep 2026 6 min read

Prefix Caching Reuses KV State Only Across Identical Prompt Prefixes

Autoregressive transformer serving often repeats the same prompt prefix across requests: a system message, a long document header, or a fixed tool schema may precede user-specific text. Prefix caching stores the key-value state produced by that shared prefix so a later request can resume computation from the cached boundary instead of recomputing the entire prefix. The useful boundary is narrower than “similar prompts.” Reuse depends on the exact token sequence and on model state that affects the cached activations. A one-character text edit may preserve most tokens, shift tokenization near the edit, or change every token after a formatting boundary. The cache can only reuse the portion whose effective input is still identical.

Artificial Intelligence 17 Sep 2026 6 min read

Retain Attention Sinks in Bounded KV Caches

A bounded KV cache seems to invite a simple eviction rule: once the cache is full, discard the oldest key-value pair and keep the most recent tokens. For some decoder-only transformers, that rule can degrade generation sharply after the sequence moves beyond the retained window. The failure is not explained only by missing old semantic content. Early tokens can receive substantial attention even when their text carries little useful information for the current prediction.

Artificial Intelligence 13 Sep 2026 6 min read

Reuse Shared Prefix State in LLM Inference

Autoregressive LLM serving often repeats the same initial tokens across many requests. A fixed system prompt, tool schema, or document prefix can occupy thousands of tokens before request-specific text begins. Computing attention state for that identical prefix on every request repeats prefill work that has already produced the same cached keys and values under compatible execution conditions. Prefix caching stores reusable attention state for such shared token prefixes. It changes the amount of prefill computation required for a cache hit, but it does not make arbitrary similar prompts interchangeable. The reusable unit is tied to exact model input state, not semantic resemblance.

Artificial Intelligence 13 Sep 2026 6 min read

Preserve Attention Sinks in Bounded KV Caches

A decoder that keeps only the newest key-value states can degrade even when its cache still contains enough recent text for the immediate task. In some transformer models, early token positions attract substantial attention across later decoding steps. Evicting those states changes the attention distribution, not just the amount of accessible history. Attention sink retention addresses that specific failure mode. A bounded cache preserves a small prefix of initial key-value states together with a moving window of recent states. Tokens between those regions can be discarded, keeping cache size bounded as generation continues.