The paper critiques the G-AP metric for benchmark contamination mitigation, arguing that aggregating performance metrics masks per-question over- and under-suppression effects. It proposes a stratified per-question probability evaluation framework and step-wise mitigation strategy to better assess and restore genuine model capabilities affected by test data leakage. The authors demonstrate that existing metrics fail to distinguish between true capability recovery and memorization suppression during decoding interventions.