Skip to content

Archive

Speculative Decoding

5 articles
Artificial Intelligence 24 Sep 2026 6 min read

Speculative Decoding Verifies Draft Tokens Without Changing the Target Sampling Distribution

Autoregressive decoding normally commits one token after each target-model evaluation. A sequence of K generated tokens therefore creates a serial dependency chain: token t+1 cannot be sampled until token t is fixed and becomes part of the prefix. Speculative decoding changes the amount of useful work obtained from a target-model call without removing that causal dependency. A cheaper draft model first proposes several continuation tokens. The target model then evaluates those candidate positions together. Tokens whose draft probabilities are compatible with the target distribution can be accepted, while the first rejected position is corrected using a residual distribution. The resulting samples follow the target model’s distribution when the acceptance and correction procedure is implemented as specified.

Artificial Intelligence 24 Sep 2026 4 min read

Speculative Decoding Verifies Draft Tokens Without Changing the Target Distribution

Autoregressive generation normally invokes the target model once for every emitted token. Speculative decoding changes that execution pattern: a cheaper draft model proposes several tokens, and the target model evaluates the proposed block in a single verification pass. The speed opportunity comes from doing useful target-model work for multiple positions at once, not from treating draft output as authoritative. Draft tokens are proposals, not final output Let the target model define distribution p and the draft model define distribution q at a given position. The draft samples a candidate token from q. Verification then decides whether that candidate can be retained as a sample consistent with p.

Artificial Intelligence 23 Sep 2026 5 min read

Speculative Sampling Preserves Target Distribution Through Rejection Correction

A draft model can propose a token that the target model would not have sampled from the same random draw, yet speculative sampling can still preserve the target model’s distribution. The key is not that the draft model predicts the target perfectly. Distributional correctness comes from the acceptance rule and the correction applied after rejection. This separates two properties that are often grouped together. Draft quality controls how frequently proposals survive verification. The rejection-correction construction controls whether the resulting sample follows the target distribution.

Artificial Intelligence 22 Sep 2026 6 min read

Speculative Decoding Couples Draft Speed with Acceptance Rate

Autoregressive generation normally advances one accepted token at a time. Each new token extends the prefix, so the next target-model evaluation depends on the token selected at the preceding position. Speculative decoding changes the execution schedule: a cheaper draft process proposes several future tokens, then the target model evaluates those positions together and decides how much of the proposal can be retained. That rearrangement can reduce the number of serial target-model calls per emitted token. It does not make verification free, and a longer draft block is not automatically better. The useful operating point depends on how quickly proposals are produced, how often they survive target verification, and what the serving stack spends on rejected work.

Artificial Intelligence 12 Sep 2026 7 min read

Speculative Decoding Depends on Draft Acceptance

Autoregressive generation normally commits one token after each model pass, creating a serial dependency across the output sequence. Speculative decoding changes that execution pattern. A cheaper draft process proposes several future tokens, then the target model evaluates those proposals together and determines which tokens can be committed. The attraction is fewer serial target-model iterations per generated token. That does not make speculative decoding an automatic latency reduction. Its useful operating point depends on how cheaply candidates are produced, how many survive verification, and how much extra work the target model performs while checking them.