Yike Wang et al. critique existing evaluation protocols for automatic harness evolution in LLM agents, arguing that methods using unit tests to search for harness configurations and reporting performance on public benchmarks are flawed. They highlight that harness evolution is an iterative process relying on task feedback, akin to agentic test-time scaling, and should be evaluated differently to avoid