Researchers propose a methodology to quantify benchmark optimization in ASR models, particularly when audio signals underdetermine reference transcripts. The approach uses three families of behavioral probes to assess models' tendencies toward reproducing benchmark-specific patterns rather than generalizing to real-world data. This work aims to provide a systematic way to detect and measure overfitting to public ASR benchmarks.

Read original