Skip to content

Archive

Model Inference

7 articles
Artificial Intelligence 16 Sep 2026 6 min read

Quantize KV Caches to Reduce Long-Context Inference Memory

Autoregressive transformer inference reuses attention keys and values from earlier tokens so each new token does not recompute the full prefix. That reuse creates the KV cache, whose memory grows with the number of cached tokens. At long context lengths or high request concurrency, the cache can become a major part of inference memory. KV cache quantization changes the representation of those stored tensors. Keys and values are written in a lower-precision format together with any scale or metadata needed for reconstruction. Attention later consumes reconstructed values or uses a kernel that handles the quantized representation directly.

Artificial Intelligence 16 Sep 2026 6 min read

Control Repetition with Contrastive Search Decoding

Greedy decoding can keep selecting locally probable tokens even when the resulting continuation becomes repetitive. Sampling can break that pattern, but it does so by introducing randomness. Contrastive search takes a different route: it remains deterministic for fixed inputs and settings while scoring likely next-token candidates against a representation-level repetition penalty. The method combines two signals that describe different properties of a candidate. The language-model probability favors tokens that fit the current prefix. A degeneration penalty disfavors candidates whose new hidden representation is too similar to representations already present in the generated context.

Artificial Intelligence 15 Sep 2026 6 min read

Steer Transformer Activations with Residual Stream Vectors

A transformer can produce different output behavior even when its weights and input tokens stay fixed. One way to cause that change is to alter an intermediate hidden state during the forward pass. Activation steering does this deliberately by adding a vector to a selected residual-stream position or set of positions. The mechanism is simple enough to express as an intervention, but its effect is not a global model setting. The chosen direction, coefficient, layer, token positions, and decoding setup all affect the result. Treating those choices as part of the inference configuration makes the behavior easier to reason about and test.

Artificial Intelligence 08 Sep 2026 9 min read

Use Test-Time Augmentation for More Stable Predictions

A classifier can change its prediction because an object moved a few pixels, an image was cropped differently, or another harmless transformation changed the input representation. If those transformations should not change the correct answer, that sensitivity is undesirable. Test-time augmentation (TTA) addresses this problem by running the same trained model on several meaning-preserving versions of an input and combining their predictions. Instead of asking the model for one view of the evidence, TTA asks it to evaluate several valid views.

Artificial Intelligence 06 Sep 2026 10 min read

Reduce Transformer Inference with Early Exits

A transformer classifier normally spends the same number of layers on every input. A straightforward support ticket and an ambiguous one both travel through the entire network, even when an intermediate representation may already contain enough information for the easy case. Early exiting changes that fixed-compute rule. It adds prediction points inside the model and lets sufficiently confident inputs stop before the final layer. Harder inputs continue through more layers. The result is input-dependent computation: the model can reduce average work without forcing every request to use a smaller network.

Artificial Intelligence 06 Sep 2026 11 min read

Reduce Transformer Inference Cost with Early Exiting

A transformer classifier normally spends the same number of layers on every input. A clear support request and an ambiguous one both pass through the entire network, even when an intermediate representation already contains enough information to classify the easy case correctly. Early exiting changes that fixed-compute rule. It attaches prediction heads to intermediate layers and lets an input stop once a chosen exit rule considers the prediction sufficiently reliable. Easy inputs can use less computation, while harder inputs continue through deeper layers.

Artificial Intelligence 06 Sep 2026 11 min read

Length-Normalized Log Probabilities for Comparing Generated Sequences

A language model assigns a probability to each next token, but applications often need to compare complete candidate sequences. A reranker may choose among generated answers. A decoder may keep several partial hypotheses. An evaluator may compare alternative completions under the same prompt. The obvious approach is to multiply each candidate’s token probabilities, or equivalently add their log probabilities. That gives the probability the model assigns to the whole continuation. It also creates an important bias: every additional token contributes a probability no greater than 1, so longer sequences usually accumulate lower raw scores even when their individual tokens are highly plausible.