KV-Cache Quantization Reduces Cache Bytes with Separate Accuracy and Kernel Costs
Autoregressive decoding retains key and value tensors from earlier tokens, so KV-cache memory grows with retained context. Quantizing those tensors changes a direct term in that memory footprint: fewer bits are stored for each cached element. The trade is not free capacity. Reduced precision adds representation error and requires a concrete scaling, storage, and kernel strategy. Cache precision is separate from weight precision Model weights and KV state have different lifetimes. Weights are persistent across requests, while KV tensors are generated from each request and grow as its sequence advances. A model can therefore use one numerical format for weights and another for its cache.