Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize
Article automatically generated from technical news.
Most agent benchmarks freeze the harness and grade the model inside it. That hides the part of the system doing much of the work, and ByteDance Seed just built a benchmark that grades the harness itself. They introduced HarnessDev, with SUTD, Georgia Tech, M-A-P, and TokenWave.AI: a 2-stage benchmark where a creator LLM starts from a weak seed that scores 0 on every task, builds a complete runnable harness (loop, tools, context, state, lifecycle, verification), freezes it, and
Fonte originale