Speculative decoding won't change your model's distribution. It might still change your output.

Article automatically generated from technical news.

There's a thread on the DeepSeek-R1 model page that's been sitting unresolved since March last year, and it bothered me enough to go read the papers. Someone had tried speculative decoding and reported that the output got worse — "very low quality words for the given context, words it would never generate by itself". Someone else replied that this is impossible, because the main model verifies and corrects everything the draft model proposes. Neither budged. They're both right

Fonte originale