This research analyzes lossy verification in Speculative Decoding, which improves inference efficiency by relaxing the strict distributional matching between draft and target models. The authors examine how this relaxation can silently rewrite decoding distributions, potentially leading to unstable or severely degraded generation quality. The study explores the mechanisms, trade-offs, and failure modes associated with these efficiency gains.

Read original