Skip to content

Archive

Beam Search

8 articles
Artificial Intelligence 24 Sep 2026 6 min read

Beam Search Length Normalization Changes Which Sequence Wins

Beam search can return a short sequence even when a longer continuation contains locally plausible tokens. The effect comes from its scoring rule: token log-probabilities are usually accumulated across the sequence, and each additional token contributes another non-positive term. Length normalization changes that ranking pressure, but it does not change the model probabilities that produced the tokens. That distinction matters in generation systems. A decoding score is an inference-time objective used to compare hypotheses. It is not automatically a calibrated probability for completed outputs, and altering it can change the selected sequence without changing a single model parameter.

Artificial Intelligence 23 Sep 2026 6 min read

Beam Search Length Normalization Changes Sequence Ranking

Beam search ranks partial sequences by scores accumulated across decoding steps. When that score is the sum of token log-probabilities, every additional token contributes a value that is normally non-positive. A longer candidate therefore has more opportunities to reduce its raw score, even when its continuation is locally plausible. That property is not a defect in probability theory. It follows from comparing complete sequences with different numbers of conditional factors. It becomes an implementation concern when a decoder is expected to produce useful completions rather than simply rank sequences by unmodified model probability.

Artificial Intelligence 16 Sep 2026 5 min read

Control Beam Search Length Bias with Sequence Scoring

Beam search can prefer a short completed sequence even when a longer continuation looks locally plausible at every token. The effect follows from the score being optimized. If a decoder ranks complete hypotheses by the sum of token log probabilities, every additional token contributes a value that is at most zero. Extending a sequence therefore cannot increase its raw accumulated log probability. This property is not a defect in probability theory. A sequence probability is a product of conditional probabilities, and its logarithm is their sum. The implementation concern appears when raw sequence probability is also used as the ranking objective for outputs whose lengths vary.

Artificial Intelligence 15 Sep 2026 6 min read

Control Length Bias in Beam Search Scoring

Beam search compares multiple partial outputs while autoregressive generation advances token by token. A common scoring rule adds token log probabilities along each candidate sequence. That rule is mathematically consistent with sequence probability, but it also creates a structural preference that developers can miss: extending a sequence normally makes its accumulated log score smaller. This matters whenever candidates of different lengths compete. A decoder can rank a short completed sequence above a longer candidate even when the longer candidate is more useful for the application. Length normalization and length penalties modify that ranking, but they also change the objective being optimized.

Artificial Intelligence 14 Sep 2026 6 min read

Control Beam Search Length Bias with Score Normalization

Beam search usually ranks partial sequences by accumulated token log probability. That score has a built-in dependence on sequence length: each additional token contributes another log probability that is zero or negative. As a result, raw cumulative scores can favor shorter completed sequences even when a longer candidate is preferable for the application. Length normalization changes the ranking rule rather than the model distribution. The distinction matters because decoding can produce different outputs without changing a single model parameter or next-token probability.

Artificial Intelligence 13 Sep 2026 7 min read

Control Sequence Length Bias in Beam Search Scoring

Beam search keeps several partial sequences alive while decoding, but the score used to compare those sequences can create a systematic preference for particular lengths. With the common sum of token log probabilities, each additional token contributes a value at or below zero. A completed sequence can therefore lose score simply by continuing, even when the continuation is plausible. This is not only a property of beam width. It comes from the objective used to rank hypotheses. Changing the beam size changes how much of the search space is explored; changing the scoring rule changes which sequences the search considers preferable.

Artificial Intelligence 13 Sep 2026 7 min read

Control Sequence Length Bias in Beam Search

Beam search can return a shorter sequence even when a longer candidate contains locally plausible tokens at every position. The behavior follows directly from sequence scoring: autoregressive models multiply conditional token probabilities, or equivalently add their log probabilities. Since token probabilities are at most one, each additional token contributes a non-positive log term. That arithmetic makes sequence length part of decoding. Beam width changes which candidates survive, but it does not remove the scoring effect. A decoder therefore needs a deliberate policy for comparing hypotheses of different lengths and for deciding when a completed hypothesis is good enough to stop the search.

Artificial Intelligence 09 Sep 2026 11 min read

Control Beam Search Length Bias with Length Normalization

Beam search is a common way to decode sequence models when choosing the most likely token at every step is too shortsighted. It keeps several partial candidates alive, expands them, and repeatedly retains the strongest alternatives. There is a subtle problem: the score used for a sequence usually accumulates one log-probability per generated token. Because token probabilities are at most 1, their log-probabilities are normally non-positive. Extending a sequence therefore tends to make its raw cumulative score smaller. When finished candidates of different lengths compete directly, this can create a preference for outputs that end too early.