Skip to content

Archive

Decoding

23 articles
Artificial Intelligence 24 Sep 2026 5 min read

Speculative Decoding Verifies Draft Tokens in Parallel

Autoregressive generation normally commits one token before the next target-model step can be evaluated. Speculative decoding changes that execution pattern. A cheaper proposal distribution produces a short candidate block, while the target model evaluates the proposed positions together. An acceptance rule then determines how much of that block can become output. The mechanism is not ordinary batching. Tokens inside the proposal remain autoregressive, and the target model still defines the intended output distribution when an exact speculative-sampling algorithm is used. The useful change is that one expensive verification pass can account for several output positions when enough proposals are accepted.

Artificial Intelligence 24 Sep 2026 5 min read

Contrastive Search Penalizes Hidden-State Repetition During Decoding

Autoregressive generation exposes a distribution over the next token, but a decoder still has to choose which candidate to append. Contrastive search changes that choice by combining model probability with a penalty for candidates whose resulting hidden state is too similar to hidden states already present in the generated prefix. The mechanism operates only at inference time. It does not alter model parameters or the next-token distribution itself. Instead, it changes the ranking used to select a token from a restricted candidate set.

Artificial Intelligence 24 Sep 2026 6 min read

Beam Search Length Normalization Changes Which Sequence Wins

Beam search can return a short sequence even when a longer continuation contains locally plausible tokens. The effect comes from its scoring rule: token log-probabilities are usually accumulated across the sequence, and each additional token contributes another non-positive term. Length normalization changes that ranking pressure, but it does not change the model probabilities that produced the tokens. That distinction matters in generation systems. A decoding score is an inference-time objective used to compare hypotheses. It is not automatically a calibrated probability for completed outputs, and altering it can change the selected sequence without changing a single model parameter.

Artificial Intelligence 23 Sep 2026 5 min read

Temperature Scaling Changes Softmax Sharpness Without Changing Logit Order

A decoder can assign the same ranking to every token before and after temperature scaling while producing substantially different probabilities. The mechanism is simple: for positive temperature T, logits are divided by T before softmax. Division by the same positive scalar preserves order, but softmax converts the changed gaps between logits into a different probability distribution. That distinction matters in inference systems because temperature does not select tokens by itself. Its visible effect depends on what happens after the scaled softmax: direct sampling, top-k filtering, top-p filtering, greedy selection, or another decoding rule.

Artificial Intelligence 23 Sep 2026 4 min read

Temperature Scaling Changes Sampling Without Changing Logit Order

Temperature is often exposed as a single generation parameter, but its effect is narrower than a general control for output quality. For a fixed vector of finite logits, positive temperature rescales the gaps before softmax. It changes the resulting probabilities without changing which logit is larger than another. That distinction matters when a serving layer combines temperature with greedy selection, top-k filtering, top-p filtering, penalties, or implementation-specific handling of zero temperature. The same numeric setting can participate in a different decoding pipeline even though the underlying scaling operation is simple.

Artificial Intelligence 23 Sep 2026 5 min read

Speculative Decoding Verifies Draft Tokens Before Acceptance

Autoregressive generation normally advances one token at a time: the model scores the next-token distribution, a decoding rule selects a token, and that token becomes part of the context for the next forward step. Speculative decoding changes the execution schedule. A cheaper draft process proposes several future tokens, while the target model retains authority over which proposals can enter the generated sequence. That separation is the core constraint. Draft tokens are predictions about future target-model decisions, not a replacement distribution that can be appended unchecked. The serving implementation must verify them against target-model scores and apply the acceptance rule required by the chosen speculative algorithm.

Artificial Intelligence 23 Sep 2026 5 min read

Repetition Penalty Rewrites Logits for Seen Token IDs

A repetition penalty can act before sampling by changing the logits of token IDs that already occur in a selected token history. The operation does not need to compare words, phrases, or rendered strings. Its unit can be the tokenizer’s integer ID, which gives the mechanism a narrower meaning than its name may suggest. That distinction matters when a decoder emits subword tokens. Two strings that appear similar to a person can map to different token sequences, while a token reused inside unrelated words can still be marked as previously seen.

Artificial Intelligence 23 Sep 2026 6 min read

Repetition Penalties Alter Token Scores Based on Prior Output

A repetition penalty changes the next-token distribution without changing the model parameters or hidden state computation. The model still produces its logits from the current context, but the decoder edits selected scores according to tokens that have already appeared. Sampling then operates on the edited scores rather than directly on the model output. That distinction matters when reproducing generation behavior. Two systems can run identical model weights on identical token IDs and still emit different continuations because their repetition rules differ in formula, token-history scope, or position in the decoding pipeline.

Artificial Intelligence 23 Sep 2026 4 min read

Logit Bias Alters Token Odds Before Sampling

A token-level bias is usually applied to model scores before probabilities are normalized. That placement matters. Adding a constant to one token’s logit changes its odds relative to every token that does not receive the same constant, even though the model parameters and hidden state remain unchanged. The mechanism is simple, but its operational effect depends on the rest of the decoding pipeline. Additive bias acts on score differences For a vocabulary with logits (z_1, \ldots, z_V), softmax assigns token (i) the probability

Artificial Intelligence 23 Sep 2026 6 min read

Beam Search Length Normalization Changes Sequence Ranking

Beam search ranks partial sequences by scores accumulated across decoding steps. When that score is the sum of token log-probabilities, every additional token contributes a value that is normally non-positive. A longer candidate therefore has more opportunities to reduce its raw score, even when its continuation is locally plausible. That property is not a defect in probability theory. It follows from comparing complete sequences with different numbers of conditional factors. It becomes an implementation concern when a decoder is expected to produce useful completions rather than simply rank sequences by unmodified model probability.

Artificial Intelligence 23 Sep 2026 6 min read

Additive Logit Bias Changes Token Odds Before Sampling

A decoder can favor or suppress a token without changing model weights. Add a constant to that token’s logit before softmax, and its probability changes relative to the rest of the vocabulary. The operation is simple, but its effect depends on where the bias enters the decoding pipeline and on every transformation that follows it. This makes additive logit bias useful as an inference control, but not as a general semantic constraint. It changes a score used by the decoder. It does not rewrite the model’s internal representation or guarantee that a concept disappears from generated text.

Artificial Intelligence 22 Sep 2026 5 min read

Top-p Sampling Rebuilds Its Candidate Set at Every Token

Top-p sampling does not keep a fixed shortlist of tokens throughout generation. At each decoding step, the model produces a new logit vector, that vector becomes a probability distribution, and the sampler forms a new candidate set whose cumulative probability mass reaches the configured threshold. The consequence is easy to miss in serving code: the same top_p value can admit two tokens at one step and dozens at another. The parameter controls probability mass, not candidate count.

Artificial Intelligence 16 Sep 2026 6 min read

Verify Speculative Decoding Without Changing Model Output

Autoregressive generation normally asks the target model to produce one next-token distribution at a time. Speculative decoding changes that execution pattern. A cheaper draft model proposes several tokens, then the target model evaluates those candidates in a batch and decides how much of the proposal can be accepted. The useful property is not merely that two models participate. The verification rule determines whether the optimization preserves the target model’s intended decoding distribution or silently changes it.

Artificial Intelligence 15 Sep 2026 6 min read

Control Length Bias in Beam Search Scoring

Beam search compares multiple partial outputs while autoregressive generation advances token by token. A common scoring rule adds token log probabilities along each candidate sequence. That rule is mathematically consistent with sequence probability, but it also creates a structural preference that developers can miss: extending a sequence normally makes its accumulated log score smaller. This matters whenever candidates of different lengths compete. A decoder can rank a short completed sequence above a longer candidate even when the longer candidate is more useful for the application. Length normalization and length penalties modify that ranking, but they also change the objective being optimized.

Artificial Intelligence 14 Sep 2026 6 min read

Speculative Decoding Trades Draft Accuracy for Target Model Work

Autoregressive generation normally commits tokens one position at a time. Even when a large model has ample parallel compute available, each next-token decision depends on the prefix produced so far. That serial dependency makes decoding latency sensitive to the number of target-model passes. Speculative decoding changes the unit of work. A cheaper draft model proposes several candidate tokens, then the target model evaluates the proposed continuation in one pass. Accepted candidates advance generation by multiple positions without requiring one separate target pass per accepted token.

Artificial Intelligence 14 Sep 2026 6 min read

Control Beam Search Length Bias with Score Normalization

Beam search usually ranks partial sequences by accumulated token log probability. That score has a built-in dependence on sequence length: each additional token contributes another log probability that is zero or negative. As a result, raw cumulative scores can favor shorter completed sequences even when a longer candidate is preferable for the application. Length normalization changes the ranking rule rather than the model distribution. The distinction matters because decoding can produce different outputs without changing a single model parameter or next-token probability.

Artificial Intelligence 14 Sep 2026 5 min read

Contrastive Decoding with an Amateur Model

A language model can assign high probability to a token for several reasons. Some reflect context-specific structure; others reflect generic tendencies that also appear in a weaker model. Contrastive decoding separates those signals by scoring candidate tokens with two models rather than one. The larger model acts as an expert. A smaller or otherwise weaker model acts as an amateur. Generation favors tokens that the expert supports more strongly relative to the amateur, subject to a plausibility constraint from the expert distribution.

Artificial Intelligence 13 Sep 2026 6 min read

Control Token Repetition with Logit Penalties

A language model can assign high probability to a token that has already appeared several times in the generated text. If the decoder keeps selecting that token or a short pattern containing it, the output may settle into repetition even though each individual choice is plausible under the model. A repetition penalty changes this behavior at decoding time. It modifies candidate scores according to token history before the next token is selected. The model parameters stay fixed, but the effective distribution used by the decoder no longer matches the model’s unmodified next-token distribution.

Artificial Intelligence 13 Sep 2026 6 min read

Contrastive Decoding with Expert and Amateur Models

A language model can give high next-token probability to text that is fluent but generic. Contrastive decoding changes the ranking by asking for a second signal: does a weaker model also find the same candidate easy to predict? A candidate favored by the expert but not by the amateur receives stronger relative support than one both models score highly. This is an inference-time mechanism. It does not alter either model’s parameters, and it does not convert the amateur model into a verifier. The decoder combines two token distributions and then selects from the resulting scores.

Artificial Intelligence 13 Sep 2026 7 min read

Contrast Expert and Amateur Models During Decoding

A language model can assign high probability to a token for two different reasons: the token may fit the prompt particularly well, or it may simply be common under many contexts. Contrastive decoding tries to separate those effects by comparing the next-token distributions of two models. A stronger expert supplies the main distribution, while a weaker amateur supplies a signal for patterns that do not require the expert’s extra capability.

Artificial Intelligence 06 Sep 2026 9 min read

Token and Sequence Biases for LLM Decoding

Sometimes an LLM produces generally good text but makes one narrow decoding choice too often. Perhaps a domain-specific abbreviation should be preferred, a deprecated product name should be discouraged, or a particular token must not appear in generated text. Changing temperature is a poor fit for this problem because temperature affects the whole next-token distribution. Retraining a model is usually excessive when the desired change is local. Token and sequence biases provide a narrower tool: modify selected prediction scores during decoding while leaving the model parameters unchanged.

Artificial Intelligence 06 Sep 2026 11 min read

Length-Normalized Log Probabilities for Comparing Generated Sequences

A language model assigns a probability to each next token, but applications often need to compare complete candidate sequences. A reranker may choose among generated answers. A decoder may keep several partial hypotheses. An evaluator may compare alternative completions under the same prompt. The obvious approach is to multiply each candidate’s token probabilities, or equivalently add their log probabilities. That gives the probability the model assigns to the whole continuation. It also creates an important bias: every additional token contributes a probability no greater than 1, so longer sequences usually accumulate lower raw scores even when their individual tokens are highly plausible.

Artificial Intelligence 06 Sep 2026 10 min read

Control LLM Repetition with Token Penalties

Language models sometimes repeat a phrase, return to the same point, or fall into a short loop even when the prompt asks for a concise answer. A common response is to increase randomness, but temperature changes the whole next-token distribution. That can reduce repetition while also making unrelated choices less predictable. Token penalties provide a more targeted control. They adjust the scores of tokens that have already appeared, making some repeated tokens less likely before the decoder chooses the next token. This can be useful for open-ended generation, but it is not a general quality switch: repeated tokens are often exactly what correct text requires.