Skip to content

Archive

Sampling

6 articles
Artificial Intelligence 24 Sep 2026 6 min read

Speculative Decoding Verifies Draft Tokens Without Changing the Target Sampling Distribution

Autoregressive decoding normally commits one token after each target-model evaluation. A sequence of K generated tokens therefore creates a serial dependency chain: token t+1 cannot be sampled until token t is fixed and becomes part of the prefix. Speculative decoding changes the amount of useful work obtained from a target-model call without removing that causal dependency. A cheaper draft model first proposes several continuation tokens. The target model then evaluates those candidate positions together. Tokens whose draft probabilities are compatible with the target distribution can be accepted, while the first rejected position is corrected using a residual distribution. The resulting samples follow the target model’s distribution when the acceptance and correction procedure is implemented as specified.

Artificial Intelligence 24 Sep 2026 4 min read

Speculative Decoding Verifies Draft Tokens Without Changing the Target Distribution

Autoregressive generation normally invokes the target model once for every emitted token. Speculative decoding changes that execution pattern: a cheaper draft model proposes several tokens, and the target model evaluates the proposed block in a single verification pass. The speed opportunity comes from doing useful target-model work for multiple positions at once, not from treating draft output as authoritative. Draft tokens are proposals, not final output Let the target model define distribution p and the draft model define distribution q at a given position. The draft samples a candidate token from q. Verification then decides whether that candidate can be retained as a sample consistent with p.

Artificial Intelligence 23 Sep 2026 4 min read

Temperature Scaling Changes Sampling Without Changing Logit Order

Temperature is often exposed as a single generation parameter, but its effect is narrower than a general control for output quality. For a fixed vector of finite logits, positive temperature rescales the gaps before softmax. It changes the resulting probabilities without changing which logit is larger than another. That distinction matters when a serving layer combines temperature with greedy selection, top-k filtering, top-p filtering, penalties, or implementation-specific handling of zero temperature. The same numeric setting can participate in a different decoding pipeline even though the underlying scaling operation is simple.

Artificial Intelligence 23 Sep 2026 5 min read

Speculative Sampling Preserves Target Distribution Through Rejection Correction

A draft model can propose a token that the target model would not have sampled from the same random draw, yet speculative sampling can still preserve the target model’s distribution. The key is not that the draft model predicts the target perfectly. Distributional correctness comes from the acceptance rule and the correction applied after rejection. This separates two properties that are often grouped together. Draft quality controls how frequently proposals survive verification. The rejection-correction construction controls whether the resulting sample follows the target distribution.

Artificial Intelligence 10 Sep 2026 10 min read

Temperature in LLM Sampling

A language model can produce very different continuations from the same prompt even when its weights and context haven’t changed. One of the controls behind that variation is temperature. Temperature is often described as a creativity knob. That description is convenient but incomplete. Temperature doesn’t add ideas to a model, improve its knowledge, or directly control factual accuracy. It changes the probability distribution used to choose the next token. The practical effect depends on what the model already considers plausible at that step.

Artificial Intelligence 09 Sep 2026 9 min read

LLM Sampling: Temperature, Top-K, and Top-P

A language model does not normally produce a single inevitable next token. Given a prefix, it assigns scores to many possible tokens. A decoding algorithm then decides how to turn those scores into the next output. That last step matters. If you sample too freely, a model can drift into unlikely continuations. If you restrict sampling too aggressively, outputs can become repetitive or lose useful variation. Parameters such as temperature, top-k, and top-p control different parts of this trade-off, so treating them as interchangeable “creativity settings” leads to confusing results.