The paper introduces a label‑free encoder pruning technique for Whisper that removes up to six encoder layers while preserving transcription performance. By targeting the encoder rather than the decoder, the method yields end‑to‑end speedups similar to those achieved by decoder‑only pruning without needing labeled data for recovery. Experiments demonstrate that a Whisper model with six fewer encoder layers retains competitive word error rates on standard ASR benchmarks. This approach fills the gap in widely adopted encoder‑size reduction strategies for pre‑trained transformer‑based speech models.

Read original