Teacher Forcing Creates a Prefix Distribution Gap
Autoregressive models predict the next token from a prefix. During teacher-forced training, that prefix usually comes from the reference sequence. During generation, it contains tokens emitted by the model itself. A prediction error can therefore change the context used for every later prediction. This difference is often called exposure bias. The useful engineering detail is more specific: training and generation can present different prefix distributions to the same conditional predictor. Token-level validation on clean reference prefixes does not fully characterize behavior after the model enters a prefix that its training data rarely presented.