Contrastive Decoding with an Amateur Model
A language model can assign high probability to a token for several reasons. Some reflect context-specific structure; others reflect generic tendencies that also appear in a weaker model. Contrastive decoding separates those signals by scoring candidate tokens with two models rather than one. The larger model acts as an expert. A smaller or otherwise weaker model acts as an amateur. Generation favors tokens that the expert supports more strongly relative to the amateur, subject to a plausibility constraint from the expert distribution.