KV Cache Quantization Trades Precision for Serving Memory
Autoregressive decoding keeps past attention keys and values so each new token can reuse earlier projections instead of recomputing them. That KV cache grows with sequence length, layer count, batch size, and the number and width of cached key/value heads. Reducing its numeric precision can cut the bytes occupied by those cached tensors, but it also changes the values consumed by later attention operations. This makes KV cache quantization different from compressing data that is only stored and restored losslessly. The quantized cache remains on the inference path. Each subsequent query can interact with approximated keys and values, so the relevant question is not only how many bytes are saved, but where quantization error enters attention and how the serving implementation contains it.