Skip to content

Archive

PagedAttention

2 articles
Artificial Intelligence 24 Sep 2026 5 min read

PagedAttention Decouples Logical KV Sequences from Physical Cache Blocks

Autoregressive serving keeps a growing key-value state for every active sequence. If that state must occupy one contiguous physical region sized for a request, allocation becomes coupled to uncertain sequence length: reserving too much wastes capacity, while extending or relocating a growing region complicates memory management. PagedAttention changes that allocation boundary. A sequence is represented as logical KV blocks, while its physical blocks may reside at unrelated locations in the cache pool.

Artificial Intelligence 23 Sep 2026 5 min read

PagedAttention Maps Logical KV Blocks to Noncontiguous Physical Memory

An autoregressive request grows its KV cache as tokens arrive, but its final sequence length is not known when decoding begins. Reserving one contiguous region for the maximum possible sequence length ties memory to capacity that may never be used. PagedAttention changes that allocation boundary: a sequence is represented as logical KV blocks, while a block table maps those logical blocks to physical blocks that need not be adjacent in GPU memory.