The authors study why aggressive token pruning hurts vision‑language model (VLM) reasoning despite retaining some visual tokens. Using a fixed‑context Pass@K analysis they show that repeatedly sampling from the same pruned visual representation recovers many instances missed by greedy decoding, indicating that useful visual information remains accessible but is used unreliably—a phenomenon they term the representation‑utilization gap. To close this gap they propose SCOPD, a sparse‑context on‑policy self‑distillation method: a student model generates reasoning trajectories from the pruned token stream while a teacher with full‑context visual tokens provides supervision on the same on‑policy prefixes, requiring no ground‑truth labels, architectural changes, or extra inference cost. An extension, SCOPD+, adds a small visual‑budget probe to locate response positions that are most sensitive to visual information and selectively distills those steps. Experiments report that with only 10 % of the original visual tokens retained, a vanilla pruned VLM preserves 86.37 % of its full‑token performance averaged over 13 benchmarks; SCOPD improves this to 90.49 % and SCOPD+ to 92.43 %. The gains hold across various token budgets, benchmarks, and pruning operators, demonstrating that efficient VLMs depend both on preserving task‑relevant visual evidence and on learning to use it reliably.

Read original