Skip to content

Archive

Quantization

10 articles
Artificial Intelligence 24 Sep 2026 4 min read

KV-Cache Quantization Reduces Cache Bytes with Separate Accuracy and Kernel Costs

Autoregressive decoding retains key and value tensors from earlier tokens, so KV-cache memory grows with retained context. Quantizing those tensors changes a direct term in that memory footprint: fewer bits are stored for each cached element. The trade is not free capacity. Reduced precision adds representation error and requires a concrete scaling, storage, and kernel strategy. Cache precision is separate from weight precision Model weights and KV state have different lifetimes. Weights are persistent across requests, while KV tensors are generated from each request and grow as its sequence advances. A model can therefore use one numerical format for weights and another for its cache.

Artificial Intelligence 24 Sep 2026 6 min read

KV Cache Quantization Reduces Stored Attention State at a Reconstruction Cost

Autoregressive decoding appends key and value tensors to a cache at every transformer layer. The cache prevents prior tokens from being projected into keys and values again, but its storage grows with sequence length. KV cache quantization changes that storage representation: older or selected cache entries are encoded with fewer bits, then reconstructed when attention consumes them. The mechanism is a memory-format trade. It does not remove tokens from the attention context and it does not change the model weights. It reduces bytes used by cached state while introducing quantization error, scale or zero-point metadata, and conversion work on the decode path.

Artificial Intelligence 24 Sep 2026 4 min read

KV Cache Quantization Perturbs Attention Through Stored Keys and Values

During autoregressive inference, previously computed keys and values are reused from the KV cache instead of being recomputed for every new token. Quantizing that cache changes more than its byte representation. The stored approximation becomes an input to later attention operations, so its error can alter both attention scores and the vectors combined by those scores. This boundary differs from quantizing model weights. A weight tensor is reused across requests, while KV state is generated from the current sequence and grows with its cached length. Its numeric range can also vary across layers, heads, positions, and requests.

Artificial Intelligence 24 Sep 2026 5 min read

Activation Outliers Distort Low-Bit Quantization Scales

A tensor can be easy to represent in floating point and awkward to map into a small integer range. The problem becomes acute when most activation values occupy a narrow interval while a small number have much larger magnitude. A shared quantization scale must cover those extremes, so the ordinary values receive fewer representable levels. That is the practical effect of activation outliers in low-bit inference. The issue is not merely that an outlier is numerically large. Its location, persistence, and relationship to the axis over which a scale is shared determine whether it materially degrades the quantized representation.

Artificial Intelligence 23 Sep 2026 5 min read

Activation Outliers Can Set the Scale for an Entire Quantization Group

A quantizer with a fixed integer width has only a finite set of representable codes. When many activations share one scale, a single value with much larger magnitude can force that scale to cover a wider real-valued range. The remaining values then occupy fewer useful code intervals around the region where they are concentrated. This behavior is not a generic statement that quantization fails in the presence of large numbers. It follows from a specific coupling: values inside the same quantization group share parameters that map real numbers to integer codes.

Artificial Intelligence 23 Sep 2026 5 min read

Activation Outliers Can Dominate Per-Tensor Quantization Scale

A per-tensor quantizer maps every value in an activation tensor through one shared scale. That coupling matters when most activations occupy a narrow interval but a few values have much larger magnitude. The large values can determine the scale, while the dense central region is represented with coarser spacing than its own range would require. This is not a statement that every large activation is erroneous or removable. An outlier may carry useful model state. The issue is numerical: one scale has to cover values with very different magnitudes.

Artificial Intelligence 15 Sep 2026 6 min read

Quantize KV Caches with Explicit Error Budgets

Autoregressive transformer inference retains key and value tensors from earlier tokens so each new token can attend to prior context without recomputing those projections. As context length and concurrent sequence count rise, this KV cache can become a substantial part of accelerator memory. KV cache quantization stores those tensors at reduced precision and reconstructs approximations when attention consumes them. The memory arithmetic is attractive, but the resulting error is not a generic model-weight perturbation. Quantized keys affect attention scores before the softmax, while quantized values affect the weighted sum after attention probabilities have been formed.

Artificial Intelligence 13 Sep 2026 7 min read

Quantize KV Caches with Separate Key and Value Error Budgets

Autoregressive transformer inference keeps past key and value tensors so each new token can attend to prior positions without recomputing the full prefix. That KV cache grows with sequence length, layer count, batch size, and the number of stored key-value heads. At long contexts, its memory footprint can become a direct limit on concurrent requests or usable context length. Quantizing the KV cache reduces bytes per stored element. The resulting approximation is not equivalent to quantizing a passive data structure, however. Cached keys participate in attention score computation, while cached values are mixed according to the resulting attention weights. Error in those two tensors therefore enters the attention operation at different points.

Artificial Intelligence 06 Sep 2026 11 min read

Compress Embeddings with Scalar Quantization

Embedding systems can become expensive for a reason that has little to do with the embedding model itself: storing and scanning the vectors. A collection of millions of dense vectors can consume gigabytes even before an index adds its own data structures. Moving those vectors through memory can also become part of query latency. Scalar quantization reduces that cost by representing each embedding coordinate with fewer bits. Instead of storing every coordinate as a 32-bit floating-point value, a system might map it to an 8-bit integer and keep enough information to approximately reconstruct or compare the original value.

Artificial Intelligence 03 Sep 2026 7 min read

LLM Quantization for Efficient Inference

Large language models can require substantial memory bandwidth and compute during inference. Quantization reduces those requirements by representing some model values with fewer bits than the floating-point formats commonly used during training. The idea sounds simple: store numbers with lower precision. In practice, quantization trades memory, latency, hardware support, implementation complexity, and model quality. Choose the configuration from measurements rather than assuming that fewer bits are always better. What quantization changes A neural network contains many numerical values, especially weights. A model stored with 16-bit weights needs roughly two bytes per weight before accounting for runtime buffers and other overhead. If those weights can instead be represented with 8 or 4 bits, their raw storage requirement falls substantially.