Skip to content

Archive

Sequence Models

9 articles
Artificial Intelligence 24 Sep 2026 6 min read

Beam Search Length Normalization Changes Which Sequence Wins

Beam search can return a short sequence even when a longer continuation contains locally plausible tokens. The effect comes from its scoring rule: token log-probabilities are usually accumulated across the sequence, and each additional token contributes another non-positive term. Length normalization changes that ranking pressure, but it does not change the model probabilities that produced the tokens. That distinction matters in generation systems. A decoding score is an inference-time objective used to compare hypotheses. It is not automatically a calibrated probability for completed outputs, and altering it can change the selected sequence without changing a single model parameter.

Artificial Intelligence 17 Sep 2026 6 min read

Teacher Forcing Creates a Prefix Distribution Gap

Autoregressive models predict the next token from a prefix. During teacher-forced training, that prefix usually comes from the reference sequence. During generation, it contains tokens emitted by the model itself. A prediction error can therefore change the context used for every later prediction. This difference is often called exposure bias. The useful engineering detail is more specific: training and generation can present different prefix distributions to the same conditional predictor. Token-level validation on clean reference prefixes does not fully characterize behavior after the model enters a prefix that its training data rarely presented.

Artificial Intelligence 16 Sep 2026 5 min read

Control Beam Search Length Bias with Sequence Scoring

Beam search can prefer a short completed sequence even when a longer continuation looks locally plausible at every token. The effect follows from the score being optimized. If a decoder ranks complete hypotheses by the sum of token log probabilities, every additional token contributes a value that is at most zero. Extending a sequence therefore cannot increase its raw accumulated log probability. This property is not a defect in probability theory. A sequence probability is a product of conditional probabilities, and its logarithm is their sum. The implementation concern appears when raw sequence probability is also used as the ranking objective for outputs whose lengths vary.

Artificial Intelligence 15 Sep 2026 6 min read

Control Length Bias in Beam Search Scoring

Beam search compares multiple partial outputs while autoregressive generation advances token by token. A common scoring rule adds token log probabilities along each candidate sequence. That rule is mathematically consistent with sequence probability, but it also creates a structural preference that developers can miss: extending a sequence normally makes its accumulated log score smaller. This matters whenever candidates of different lengths compete. A decoder can rank a short completed sequence above a longer candidate even when the longer candidate is more useful for the application. Length normalization and length penalties modify that ranking, but they also change the objective being optimized.

Artificial Intelligence 14 Sep 2026 5 min read

Account for Exposure Bias in Autoregressive Generation

An autoregressive model can receive a clean prefix at every training position and still face a different input distribution during generation. Training commonly scores the next reference token while conditioning on earlier reference tokens. At inference time, the prefix contains the model’s own outputs instead. This mismatch is called exposure bias. It matters because an early generation error does more than make one token incorrect. That token becomes part of the context for later predictions, placing the model in a prefix state that may have been rare or absent during training.

Artificial Intelligence 13 Sep 2026 7 min read

Control Sequence Length Bias in Beam Search Scoring

Beam search keeps several partial sequences alive while decoding, but the score used to compare those sequences can create a systematic preference for particular lengths. With the common sum of token log probabilities, each additional token contributes a value at or below zero. A completed sequence can therefore lose score simply by continuing, even when the continuation is plausible. This is not only a property of beam width. It comes from the objective used to rank hypotheses. Changing the beam size changes how much of the search space is explored; changing the scoring rule changes which sequences the search considers preferable.

Artificial Intelligence 09 Sep 2026 12 min read

Evaluate Sequence Models Beyond Training Length

A sequence model can score well on a random test split and still fail when an input is longer than the sequences it saw during training. This matters for language, symbolic reasoning, event sequences, and other tasks where production inputs do not have one fixed length. The problem is easy to hide. If training and test examples come from the same length distribution, an aggregate metric mostly measures performance on familiar lengths. It does not tell you whether the model learned a rule that extends to longer sequences or a strategy that works only inside the observed range.

Artificial Intelligence 09 Sep 2026 11 min read

Control Beam Search Length Bias with Length Normalization

Beam search is a common way to decode sequence models when choosing the most likely token at every step is too shortsighted. It keeps several partial candidates alive, expands them, and repeatedly retains the strongest alternatives. There is a subtle problem: the score used for a sequence usually accumulates one log-probability per generated token. Because token probabilities are at most 1, their log-probabilities are normally non-positive. Extending a sequence therefore tends to make its raw cumulative score smaller. When finished candidates of different lengths compete directly, this can create a preference for outputs that end too early.

Artificial Intelligence 04 Sep 2026 10 min read

Teacher Forcing in Autoregressive Models

An autoregressive model generates a sequence one element at a time. A language model predicts the next token from the tokens before it; a sequence model might similarly predict the next symbol, event, or value from an existing prefix. That creates a practical training question: when teaching the model to predict step 5, should the input contain the correct steps 1–4 from the dataset, or the model’s own earlier predictions?