Quantize KV Caches to Reduce Long-Context Inference Memory
Autoregressive transformer inference reuses attention keys and values from earlier tokens so each new token does not recompute the full prefix. That reuse creates the KV cache, whose memory grows with the number of cached tokens. At long context lengths or high request concurrency, the cache can become a major part of inference memory. KV cache quantization changes the representation of those stored tensors. Keys and values are written in a lower-precision format together with any scale or metadata needed for reconstruction. Attention later consumes reconstructed values or uses a kernel that handles the quantized representation directly.