Skip to content

Archive

Large Language Models

8 articles
Artificial Intelligence 16 Sep 2026 6 min read

Quantize KV Caches to Reduce Long-Context Inference Memory

Autoregressive transformer inference reuses attention keys and values from earlier tokens so each new token does not recompute the full prefix. That reuse creates the KV cache, whose memory grows with the number of cached tokens. At long context lengths or high request concurrency, the cache can become a major part of inference memory. KV cache quantization changes the representation of those stored tensors. Keys and values are written in a lower-precision format together with any scale or metadata needed for reconstruction. Attention later consumes reconstructed values or uses a kernel that handles the quantized representation directly.

Artificial Intelligence 16 Sep 2026 6 min read

Extend RoPE Context with Positional Interpolation

A transformer using rotary position embeddings can accept tensors longer than the sequence length used during training, yet accepting the shape does not establish that its position signal remains usable at those distances. Rotary angles at unseen positions can place attention computations outside the positional regime the model encountered during optimization. Positional interpolation changes the input to rotary position embeddings rather than merely raising a sequence-length limit. For a target context longer than the original training context, position indices are compressed so the extended sequence maps into the earlier positional range. The model then needs to adapt to denser positional spacing instead of extrapolating directly to larger indices.

Artificial Intelligence 16 Sep 2026 6 min read

Control Repetition with Contrastive Search Decoding

Greedy decoding can keep selecting locally probable tokens even when the resulting continuation becomes repetitive. Sampling can break that pattern, but it does so by introducing randomness. Contrastive search takes a different route: it remains deterministic for fixed inputs and settings while scoring likely next-token candidates against a representation-level repetition penalty. The method combines two signals that describe different properties of a candidate. The language-model probability favors tokens that fit the current prefix. A degeneration penalty disfavors candidates whose new hidden representation is too similar to representations already present in the generated context.

Artificial Intelligence 12 Sep 2026 6 min read

Contrast Language Model Logits with Expert-Amateur Decoding

A language model can assign high probability to tokens that are fluent but generic. Contrastive decoding changes token selection by comparing a stronger expert model with a weaker amateur model at the same generation position. A token becomes attractive when the expert favors it more strongly than the amateur does. The comparison is not an unrestricted subtraction across the vocabulary. The original method also keeps candidate tokens inside a plausibility set defined by the expert. That constraint matters because a large expert-amateur score gap can otherwise promote a token that both models consider implausible.

Artificial Intelligence 10 Sep 2026 11 min read

Improve LLM Fine-Tuning with Rejection Sampling

Improve LLM Fine-Tuning with Rejection Sampling Suppose you can tell a good model response from a bad one, but writing thousands of ideal responses by hand is expensive. A capable language model may already produce acceptable answers some of the time. The problem is that those answers are mixed with weaker ones. Rejection sampling fine-tuning turns that observation into a data-generation loop. For each prompt, generate several candidate responses, evaluate them, keep responses that satisfy a selection rule, and use the accepted prompt-response pairs for supervised fine-tuning. The method can concentrate training on behavior you want without requiring a human to author every target from scratch.

Artificial Intelligence 06 Sep 2026 10 min read

Steer Language Models by Editing Hidden Activations

Prompting changes what a language model reads. Fine-tuning changes its parameters. There is another, more experimental way to influence generation: change the model’s internal activations while it runs. This technique is commonly called activation steering or activation engineering. A simple version measures how hidden representations differ between examples that express opposite properties, turns that difference into a steering vector, and adds a scaled version of the vector during inference. The model weights stay unchanged.

Artificial Intelligence 06 Sep 2026 9 min read

Inspect Language Model Uncertainty with Token Entropy

A language model can produce fluent text even when several continuations look similarly plausible to the model. Looking only at the selected token hides that ambiguity: a token chosen with probability 0.90 and one chosen from a nearly even 0.51 versus 0.49 split both appear as a single output token. Token entropy summarizes how spread out the model’s next-token probability distribution is. It can help developers inspect uncertain generation steps, compare decoding behavior under controlled conditions, and build diagnostic signals for evaluation. But entropy is not a probability that the model is correct, and using it as one leads to unreliable decisions.

Artificial Intelligence 05 Sep 2026 10 min read

Attention Sinks for Stable Streaming LLM Inference

Autoregressive language models normally reuse the keys and values of earlier tokens while generating the next token. This KV cache avoids recomputing the entire prefix at every decoding step, but its memory use grows with the cached sequence. A long-running chat, agent, or stream can therefore accumulate more cached state than a serving system wants to keep. A tempting fix is a sliding window: retain only the most recent tokens and evict everything older. For models trained with ordinary dense attention, however, abruptly dropping all early tokens can damage generation quality even when those old tokens do not appear semantically important.