Skip to content

Archive

Probability

2 articles
Artificial Intelligence 22 Sep 2026 5 min read

Top-p Sampling Rebuilds Its Candidate Set at Every Token

Top-p sampling does not keep a fixed shortlist of tokens throughout generation. At each decoding step, the model produces a new logit vector, that vector becomes a probability distribution, and the sampler forms a new candidate set whose cumulative probability mass reaches the configured threshold. The consequence is easy to miss in serving code: the same top_p value can admit two tokens at one step and dozens at another. The parameter controls probability mass, not candidate count.

Artificial Intelligence 06 Sep 2026 10 min read

Compare Model Distributions with KL Divergence

AI systems often produce probability distributions rather than single answers. A classifier assigns probabilities to classes, a language model assigns probabilities to possible next tokens, and a teacher model can provide a soft target distribution for a smaller student. In all of these cases, developers need a way to ask: how different is one probability distribution from another? Kullback-Leibler divergence, usually shortened to KL divergence, is one answer. It measures how much a comparison distribution Q differs from a reference distribution P, with the differences weighted by what P considers important.