Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize
Most agent benchmarks freeze the harness and grade the model inside it. That hides the part of the system doing much of the work, and ByteDance Seed just built a benchmark that grades
→ View original source