Contrast Expert and Amateur Models During Decoding
A language model can assign high probability to a token for two different reasons: the token may fit the prompt particularly well, or it may simply be common under many contexts. Contrastive decoding tries to separate those effects by comparing the next-token distributions of two models. A stronger expert supplies the main distribution, while a weaker amateur supplies a signal for patterns that do not require the expert’s extra capability.