Skip to content

Archive

Language Models

40 articles
Artificial Intelligence 24 Sep 2026 5 min read

Multi-Token Prediction Adds Parallel Future-Token Losses to a Shared Model Trunk

A next-token language model normally applies one predictive objective at position t: the hidden representation at that position is used to score token x_(t+1). Multi-token prediction changes that training boundary. One shared model trunk produces the representation, while multiple output heads predict several subsequent tokens from that position. The mechanism adds supervision at multiple future offsets without requiring a separate transformer trunk for every offset. It is therefore a change to the training objective and prediction heads, not a claim that ordinary autoregressive generation can emit several unchecked tokens as one exact step.

Artificial Intelligence 24 Sep 2026 5 min read

Contrastive Search Penalizes Hidden-State Repetition During Decoding

Autoregressive generation exposes a distribution over the next token, but a decoder still has to choose which candidate to append. Contrastive search changes that choice by combining model probability with a penalty for candidates whose resulting hidden state is too similar to hidden states already present in the generated prefix. The mechanism operates only at inference time. It does not alter model parameters or the next-token distribution itself. Instead, it changes the ranking used to select a token from a restricted candidate set.

Artificial Intelligence 23 Sep 2026 3 min read

Stable Softmax Subtracts the Maximum Logit Before Exponentiation

A decoder can receive logits large enough that direct exponentiation is numerically unsafe even though the intended probability distribution is ordinary. Softmax does not require exponentiating the original values. Subtracting the largest logit from every logit produces the same distribution in exact real arithmetic while moving the exponentials into a safer numeric range. This shift is a property of softmax itself, not a model-specific heuristic. A shared shift cancels during normalization For logits (z_1,\ldots,z_n), softmax assigns

Artificial Intelligence 23 Sep 2026 5 min read

Speculative Sampling Preserves Target Distribution Through Rejection Correction

A draft model can propose a token that the target model would not have sampled from the same random draw, yet speculative sampling can still preserve the target model’s distribution. The key is not that the draft model predicts the target perfectly. Distributional correctness comes from the acceptance rule and the correction applied after rejection. This separates two properties that are often grouped together. Draft quality controls how frequently proposals survive verification. The rejection-correction construction controls whether the resulting sample follows the target distribution.

Artificial Intelligence 23 Sep 2026 6 min read

Sliding-Window Attention Bounds Active KV State by Token Distance

Sliding-window attention imposes a finite token-distance boundary on causal attention. At position (t), a query can address only a recent interval of key and value positions rather than the full prefix. Once a cached position falls permanently outside that interval, later queries governed by the same local rule cannot address it. That boundary changes the state required for autoregressive decoding. Full causal attention keeps usable KV state growing with sequence length. A fixed local window can keep the active KV span bounded, provided the runtime evicts or overwrites entries that have become unreachable.

Artificial Intelligence 23 Sep 2026 5 min read

RMSNorm Scales Activations Without Mean Centering

RMSNorm rescales a hidden vector using its root-mean-square magnitude, but it does not subtract the vector’s feature mean first. That omission is not merely a shorter expression for LayerNorm. It changes which transformations of the input disappear under normalization and which remain visible to later operations. For transformer implementations, that distinction matters at the boundary between residual state, normalization, and the next projection. The denominator comes from the second raw moment For a hidden vector (x \in \mathbb{R}^d), a common RMSNorm form is

Artificial Intelligence 23 Sep 2026 5 min read

Repetition Penalty Rewrites Logits for Seen Token IDs

A repetition penalty can act before sampling by changing the logits of token IDs that already occur in a selected token history. The operation does not need to compare words, phrases, or rendered strings. Its unit can be the tokenizer’s integer ID, which gives the mechanism a narrower meaning than its name may suggest. That distinction matters when a decoder emits subword tokens. Two strings that appear similar to a person can map to different token sequences, while a token reused inside unrelated words can still be marked as previously seen.

Artificial Intelligence 23 Sep 2026 5 min read

PagedAttention Maps Logical KV Blocks to Noncontiguous Physical Memory

An autoregressive request grows its KV cache as tokens arrive, but its final sequence length is not known when decoding begins. Reserving one contiguous region for the maximum possible sequence length ties memory to capacity that may never be used. PagedAttention changes that allocation boundary: a sequence is represented as logical KV blocks, while a block table maps those logical blocks to physical blocks that need not be adjacent in GPU memory.

Artificial Intelligence 23 Sep 2026 4 min read

Logit Bias Alters Token Odds Before Sampling

A token-level bias is usually applied to model scores before probabilities are normalized. That placement matters. Adding a constant to one token’s logit changes its odds relative to every token that does not receive the same constant, even though the model parameters and hidden state remain unchanged. The mechanism is simple, but its operational effect depends on the rest of the decoding pipeline. Additive bias acts on score differences For a vocabulary with logits (z_1, \ldots, z_V), softmax assigns token (i) the probability

Artificial Intelligence 23 Sep 2026 5 min read

Grouped-Query Attention Shares KV Heads Across Query Groups

Grouped-query attention (GQA) changes a specific structural ratio inside an attention layer: the number of query heads can exceed the number of key and value heads. Several query heads then consume the same projected key and value head. The attention calculation remains head-specific on the query side, while KV state is shared within each group. That asymmetry matters during autoregressive decoding because cached keys and values persist for prior tokens. Reducing the count of distinct KV heads reduces the amount of per-token KV state that must remain available to later decoding steps.

Artificial Intelligence 22 Sep 2026 6 min read

Speculative Decoding Couples Draft Speed with Acceptance Rate

Autoregressive generation normally advances one accepted token at a time. Each new token extends the prefix, so the next target-model evaluation depends on the token selected at the preceding position. Speculative decoding changes the execution schedule: a cheaper draft process proposes several future tokens, then the target model evaluates those positions together and decides how much of the proposal can be retained. That rearrangement can reduce the number of serial target-model calls per emitted token. It does not make verification free, and a longer draft block is not automatically better. The useful operating point depends on how quickly proposals are produced, how often they survive target verification, and what the serving stack spends on rejected work.

Artificial Intelligence 19 Sep 2026 6 min read

Prefix Caching Reuses KV State Only Across Identical Prompt Prefixes

Autoregressive transformer serving often repeats the same prompt prefix across requests: a system message, a long document header, or a fixed tool schema may precede user-specific text. Prefix caching stores the key-value state produced by that shared prefix so a later request can resume computation from the cached boundary instead of recomputing the entire prefix. The useful boundary is narrower than “similar prompts.” Reuse depends on the exact token sequence and on model state that affects the cached activations. A one-character text edit may preserve most tokens, shift tokenization near the edit, or change every token after a formatting boundary. The cache can only reuse the portion whose effective input is still identical.

Artificial Intelligence 19 Sep 2026 8 min read

Chunk Long Prefills to Limit Decode Stalls in LLM Serving

A long prompt can occupy an accelerator for a much larger scheduling interval than a single decode iteration. When a serving engine mixes new prefills with requests that are already generating tokens, that difference can show up as irregular time between output tokens. The model has not changed; the interference comes from how two distinct inference phases share execution time. Prefill processes a prompt and builds the key-value state required by later causal attention. Decode then extends the sequence autoregressively, usually one new token per active request per iteration. Those phases place different pressure on hardware, so treating them as interchangeable scheduling units can produce avoidable stalls.

Artificial Intelligence 18 Sep 2026 5 min read

Treat the Logit Lens as a Readout, Not a Causal Trace

A transformer can expose an intermediate residual state that strongly favors a token under the model’s final vocabulary projection, then produce a different token after later blocks run. The logit lens makes that intermediate preference visible. It does not establish that the preference caused the final output. That boundary matters when developers use layer-by-layer token rankings to inspect model behavior. The logit lens is a readout: it asks what the model’s output head would report if applied to an intermediate representation. The actual forward pass asks a different question because every remaining block can transform that representation before the final readout.

Artificial Intelligence 17 Sep 2026 6 min read

Teacher Forcing Creates a Prefix Distribution Gap

Autoregressive models predict the next token from a prefix. During teacher-forced training, that prefix usually comes from the reference sequence. During generation, it contains tokens emitted by the model itself. A prediction error can therefore change the context used for every later prediction. This difference is often called exposure bias. The useful engineering detail is more specific: training and generation can present different prefix distributions to the same conditional predictor. Token-level validation on clean reference prefixes does not fully characterize behavior after the model enters a prefix that its training data rarely presented.

Artificial Intelligence 16 Sep 2026 5 min read

Read Intermediate Transformer States With Logit Lens

A decoder-only transformer produces its next-token distribution only after the final hidden state has passed through the model’s output normalization and vocabulary projection. Logit lens reuses that output path on states from earlier transformer blocks. The result is a sequence of vocabulary distributions that can expose how token preferences change with depth. The method is attractive because it maps internal vectors into familiar token space without fitting a separate classifier. That convenience also creates a sharp interpretive boundary: an intermediate state was not necessarily optimized to be directly decoded by the final output map. A readable token distribution is a probe of that state, not a guarantee that the model has already settled on the same prediction.

Artificial Intelligence 16 Sep 2026 6 min read

Read Intermediate Transformer Predictions with Tuned Lenses

A transformer can carry useful information about its eventual next-token distribution several blocks before the final layer. Reading that information is not as simple as applying the model’s output projection to every intermediate hidden state. The final output head is calibrated for representations at the end of the network, while residual representations can shift across depth. A tuned lens addresses that mismatch with a separate affine translator for each inspected layer. The translator maps an intermediate residual state into a representation that the frozen final normalization and output projection can decode. This produces a token distribution that can be compared across layers without assuming that every layer already uses the final representation basis.

Artificial Intelligence 16 Sep 2026 6 min read

Mask Padding Tokens in Language Model Loss

Variable-length text batches are commonly padded into rectangular tensors. The extra positions simplify batching, but they are not ordinary training targets. If padded target positions contribute to cross-entropy, the optimizer receives gradients for synthetic symbols that were introduced only to align tensor shapes. Preventing that signal requires a loss mask. An attention mask can stop selected positions from participating in attention, but that does not by itself remove their target terms from the objective.

Artificial Intelligence 14 Sep 2026 5 min read

Tie Input Embeddings to the Output Projection

A language model can contain two large matrices indexed by the same vocabulary: one maps token IDs into embedding vectors, while another maps hidden states into vocabulary logits. Weight tying makes those roles share parameters instead of maintaining two independent matrices. The change is compact in code, but it affects parameter counting, gradient flow, dimensional constraints, checkpoint handling, and any component that assumes the input and output weights are separate.

Artificial Intelligence 14 Sep 2026 7 min read

Reduce Repetition with Unlikelihood Training

An autoregressive language model is usually trained to increase the probability of the observed next token. That positive objective does not directly state which plausible but unwanted alternatives should receive less probability. When repetitive tokens or phrases remain locally probable, ordinary next-token training can leave generation with a strong route back into content that has already appeared. Unlikelihood training adds a negative signal for selected candidates. Instead of only rewarding the target token, the objective can also penalize tokens chosen because they represent an unwanted behavior, such as repetition within the generated prefix.

Artificial Intelligence 14 Sep 2026 6 min read

Control Beam Search Length Bias with Score Normalization

Beam search usually ranks partial sequences by accumulated token log probability. That score has a built-in dependence on sequence length: each additional token contributes another log probability that is zero or negative. As a result, raw cumulative scores can favor shorter completed sequences even when a longer candidate is preferable for the application. Length normalization changes the ranking rule rather than the model distribution. The distinction matters because decoding can produce different outputs without changing a single model parameter or next-token probability.

Artificial Intelligence 14 Sep 2026 5 min read

Contrastive Decoding with an Amateur Model

A language model can assign high probability to a token for several reasons. Some reflect context-specific structure; others reflect generic tendencies that also appear in a weaker model. Contrastive decoding separates those signals by scoring candidate tokens with two models rather than one. The larger model acts as an expert. A smaller or otherwise weaker model acts as an amateur. Generation favors tokens that the expert supports more strongly relative to the amateur, subject to a plausibility constraint from the expert distribution.

Artificial Intelligence 13 Sep 2026 6 min read

Inspect Intermediate Transformer Predictions with Logit Lens

A transformer produces its next-token distribution only after the final block, yet every block updates the residual state that eventually feeds that prediction. Logit lens examines those intermediate states by mapping them through the model’s output path into vocabulary logits. The result is a sequence of provisional token distributions across model depth. The method is attractive because it reuses components already present in the model. Its output also needs careful interpretation. An intermediate residual state was not necessarily optimized to behave like a final residual state, so a readable token ranking is a diagnostic projection rather than a direct transcript of internal computation.

Artificial Intelligence 13 Sep 2026 7 min read

Control Sequence Length Bias in Beam Search Scoring

Beam search keeps several partial sequences alive while decoding, but the score used to compare those sequences can create a systematic preference for particular lengths. With the common sum of token log probabilities, each additional token contributes a value at or below zero. A completed sequence can therefore lose score simply by continuing, even when the continuation is plausible. This is not only a property of beam width. It comes from the objective used to rank hypotheses. Changing the beam size changes how much of the search space is explored; changing the scoring rule changes which sequences the search considers preferable.