This paper addresses two critical bottlenecks in Retrieval-Augmented Generation: flawed measurement of evidence utilization and suboptimal